Object Statistics Method, Apparatus, Device, and Storage Medium

Through iteratively refine features, the similarity calculation and fusion of image feature maps and example feature maps are used to generate refined feature maps, which solves the problem of existing counting algorithms dependence on high-quality labeled data, and achieves high accuracy and universality in multi-class object counting.

CN115035478BActive Publication Date: 2025-07-04SHANGHAI SENSETIME TECH DEV CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210764687.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-29
Publication Date
2025-07-04
Estimated Expiration
2042-06-29

AI Technical Summary

Technical Problem

Existing deep learning-based counting algorithms require a large amount of high-quality labeled data and are difficult to effectively apply in object counting tasks that lack training data. The existing algorithms are usually targeted at specific categories and cannot be migrated to other counting scenarios.

Method used

Using the iteratively refined features method, a refined feature map is generated through the similarity calculation and fusion of image feature maps and example feature maps, which are used to generate density estimation maps and perform object statistics, including steps such as feature extraction, correlation calculation, feature refinement and density regression.

Benefits of technology

It improves the accuracy and versatility of object statistics, and can effectively count multiple objects under a small amount of labeled data. It is suitable for monitoring, nature conservation and industrial manufacturing scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115035478B_ABST
    Figure CN115035478B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose an object statistical method, apparatus, device, and storage medium. Among them, the method includes: obtaining a to-be-processed image and at least one example image corresponding to the object to be statistically analyzed; extracting an image feature map corresponding to the to-be-processed image and example feature maps corresponding to each of the at least one example image; generating a refined feature map based on the image feature map and each of the example feature maps; the refined feature map retains features related to each of the example feature maps on the basis of the image feature map; converting the refined feature map into a density estimation map, and generating a statistical result for the object to be statistically analyzed based on the density estimation map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of data processing, and in particular, to an object counting method, apparatus, device, and storage medium. Background Art

[0002] In recent years, deep learning algorithms have made great progress in various fields and are also applied to many scene analysis fields. In the field of scene analysis, counting dense objects is a very important task. Currently, there are many fully automatic image-based counting methods based on deep learning, such as counting people, counting goats, counting airport luggage, etc. These methods use deep learning models to use image information to quickly and automatically count specific objects, and have been widely applied in scenarios such as counting the number of people in the monitoring field, counting animals in nature protection applications, and counting workpieces in factory industrial manufacturing. Summary of the Invention

[0003] Embodiments of the present disclosure provide an object counting method, apparatus, device, and storage medium.

[0004] In a first aspect, an object counting method is provided, including:

[0005] Obtaining at least one example image corresponding to the image to be processed and the object to be counted;

[0006] Extracting an image feature map corresponding to the image to be processed and an example feature map corresponding to each of the at least one example image;

[0007] Generating a refined feature map based on the image feature map and each example feature map; the refined feature map retains features related to each example feature map on the basis of the image feature map;

[0008] Converting the refined feature map into a density estimation map, and generating a counting result for the object to be counted based on the density estimation map.

[0009] In a second aspect, a model training method is provided, the method including:

[0010] Obtaining a sample image, a true density map corresponding to the sample image, and at least one example image corresponding to the object to be counted;

[0011] Input the sample image and the at least one example image into a statistical model to be trained to obtain a predicted density map; the statistical model is used to extract an image feature map corresponding to the sample image and an example feature map corresponding to each of the at least one example image; generate a refined feature map based on the image feature map and each example feature map, the refined feature map retaining features related to each example feature map on the basis of the image feature map; and convert the refined feature map into the predicted density map;

[0012] Based on the predicted density map and the ground truth density map, adjust the parameters of the statistical model to be trained to obtain a trained statistical model.

[0013] In a third aspect, there is provided an object statistics device, comprising:

[0014] A first acquisition module, configured to acquire an image to be processed and at least one example image corresponding to an object to be counted;

[0015] An extraction module, configured to extract an image feature map corresponding to the image to be processed and an example feature map corresponding to each of the at least one example image;

[0016] A generation module, configured to generate a refined feature map based on the image feature map and each example feature map; the refined feature map retaining features related to each example feature map on the basis of the image feature map;

[0017] A conversion module, configured to convert the refined feature map into a density estimation map and generate a statistical result for the object to be counted based on the density estimation map.

[0018] In a fourth aspect, there is provided a model training device, comprising:

[0019] A second acquisition module, configured to acquire a sample image, a ground truth density map corresponding to the sample image, and at least one example image corresponding to an object to be counted;

[0020] An input module, configured to input the sample image and the at least one example image into a statistical model to be trained to obtain a predicted density map; the statistical model is used to extract an image feature map corresponding to the sample image and an example feature map corresponding to each of the at least one example image; generate a refined feature map based on the image feature map and each example feature map, the refined feature map retaining features related to each example feature map on the basis of the image feature map; and convert the refined feature map into the predicted density map;

[0021] An adjustment module, configured to adjust parameters of the statistical model to be trained based on the predicted density map and the ground-truth density map, so as to obtain a trained statistical model.

[0022] In a fifth aspect, an object statistics device is provided, including: a memory and a processor, where the memory stores a computer program that can run on the processor, and when the processor executes the computer program, the steps in the above method are implemented.

[0023] In a sixth aspect, a computer storage medium is provided, where the computer storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the above method.

[0024] Based on the above embodiments, a corresponding refined feature map is generated through the image feature map and each of the example feature maps. Since the refined feature map retains the features related to each of the example feature maps on the basis of the image feature map, furthermore, in the process of determining the statistical result for the object to be counted based on the refined feature map, a statistical result corresponding to the object to be counted can be obtained, improving the statistical accuracy.

[0025] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the technical solutions of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 A flowchart of an object statistics method provided by an embodiment of the present disclosure;

[0027] Figure 2 A flowchart of an object statistics method provided by an embodiment of the present disclosure;

[0028] Figure 3 A flowchart of an object statistics method provided by an embodiment of the present disclosure;

[0029] Figure 4 A flowchart of an object statistics method provided by an embodiment of the present disclosure;

[0030] Figure 5 A flowchart of an object statistics method provided by an embodiment of the present disclosure;

[0031] Figure 6 A flowchart of an object statistics method provided by an embodiment of the present disclosure;

[0032] Figure 7 A flowchart of an object statistics method provided by an embodiment of the present disclosure;

[0033] Figure 8AAn optional flowchart of the model training method provided by the embodiments of the present disclosure;

[0034] Figure 8B A schematic structural diagram of a statistical model provided by the embodiments of the present disclosure;

[0035] Figure 9A An optional flowchart of the model training method provided by the embodiments of the present disclosure;

[0036] Figure 9B A schematic structural diagram of a relevance calculation network provided by the embodiments of the present disclosure;

[0037] Figure 10A An optional flowchart of the model training method provided by the embodiments of the present disclosure;

[0038] Figure 10B A schematic structural diagram of another relevance calculation network provided by the embodiments of the present disclosure;

[0039] Figure 11A An optional flowchart of the model training method provided by the embodiments of the present disclosure;

[0040] Figure 11B A schematic structural diagram of a feature refinement network provided by the embodiments of the present disclosure;

[0041] Figure 12A An optional flowchart of the model training method provided by the embodiments of the present disclosure;

[0042] Figure 12B A schematic structural diagram of another statistical model provided by the embodiments of the present disclosure;

[0043] Figure 13 A schematic system architecture diagram of the general counting system provided by the embodiments of the present disclosure;

[0044] Figure 14 A schematic diagram of the labeled data provided by the embodiments of the present disclosure;

[0045] Figure 15 A schematic processing flowchart of the network module provided by the embodiments of the present disclosure;

[0046] Figure 16 A schematic network structure diagram provided by the embodiments of the present disclosure;

[0047] Figure 17 A schematic composition structure diagram of an object counting device provided by the embodiments of the present disclosure;

[0048] Figure 18 A schematic composition structure diagram of a model training device provided by the embodiments of the present disclosure

[0049] Figure 19 Schematic diagram of the hardware entity of an object statistics device provided by an embodiment of the present disclosure. Detailed implementation manners

[0050] The technical solutions of the present disclosure will be described in detail below through embodiments in combination with the accompanying drawings. These specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0051] It should be noted that in the examples of the present disclosure, "first", "second", etc. are used to distinguish similar objects, and do not necessarily need to be used to describe the order or sequence of the targets. In addition, the technical solutions described in the embodiments of the present disclosure can be combined arbitrarily without conflict.

[0052] The exemplary applications of the electronic device provided by the embodiments of the present disclosure are described below. The electronic device provided by the embodiments of the present disclosure can be implemented as various types of user terminals (hereinafter referred to as terminals) such as notebook computers, tablet computers, desktop computers, set-top boxes, mobile devices (such as mobile phones, portable music players, personal digital assistants, dedicated messaging devices, portable game devices), etc., or can be implemented as a server.

[0053] See Figure 1 , Figure 1 is an optional flowchart of the object statistics method provided by the embodiments of the present disclosure, and will be described in combination with Figure 1 the steps shown.

[0054] S101. Obtain at least one example image corresponding to the image to be processed and the object to be counted.

[0055] Wherein, the image to be processed is an image for which the number of objects to be counted needs to be counted. In some implementation scenarios, the image to be processed includes a large number of objects to be counted, and the object statistics method provided by the embodiments of the present disclosure is used to count the objects to be counted in the image to be processed.

[0056] Wherein, the example image corresponding to the object to be counted is an image including the object to be counted, and the at least one example image includes the same object to be counted. That is to say, the object to be counted can be determined through the at least one example image. That is to say, the image to be processed may include multiple different types of objects. When the objects to be counted corresponding to the input at least one example image are different, the embodiments of the present disclosure can output the statistical results corresponding to different objects to be counted.

[0057] In some embodiments, the example image may be from the same image as the image to be processed. In this embodiment, the user may select at least one object to be counted in the image to be processed to obtain at least one detection box for the object to be counted. Through the at least one detection box corresponding to the object to be counted, the corresponding example image may be intercepted from the example image.

[0058] In some embodiments, the example image may be from a different image from the image to be processed. Exemplarily, in the case where the object to be processed in image A needs to be counted, the user may select at least one object to be counted in image B to obtain at least one detection box for the object to be counted, and then obtain an example image from image B; the user may also select at least one object to be counted in image C to obtain at least one detection box for the object to be counted, and then obtain an example image from image C; the example images from image B and image C are used as the at least one example image corresponding to the object to be counted.

[0059] It should be noted that there is at least one object to be counted in the example image. That is to say, in the example image, there may be one object to be counted, or there may be two or more objects to be counted.

[0060] S102. Extract the image feature map corresponding to the image to be processed and the example feature map corresponding to each example image in the at least one example image.

[0061] In some embodiments, through a pre-trained feature extraction network, the image features corresponding to the image to be processed and each example image may be extracted respectively to obtain the image feature map corresponding to the image to be processed and the example feature map corresponding to each example image.

[0062] Among them, the feature extraction network for extracting the image feature map may be the same network as the feature extraction network for extracting the example feature map, or may be different networks.

[0063] It should be noted that the example feature map corresponding to each example image corresponds to a first size, that is, the sizes of the example feature maps corresponding to each example image are the same; the image feature map is a second size, and the second size is larger than the first size.

[0064] In some embodiments, the feature extraction network may be a multi-scale feature extraction network. Exemplarily, taking the size of the image feature map as H×W×C as an example, the C channels may include feature information of multiple scales.

[0065] S103. Generate a refined feature map based on the image feature map and each example feature map; the refined feature map retains the features related to each example feature map on the basis of the image feature map.

[0066] In some embodiments, intermediate feature maps corresponding to the image feature map and each example feature map can be obtained respectively. The intermediate feature map retains the features related to the example feature map on the basis of the image feature map. Then, by fusing the intermediate feature maps corresponding to each example feature map, the refined feature map can be obtained.

[0067] Among them, for each example feature map, the intermediate feature map corresponding to the example feature map can be generated by calculating the similarity between the example feature map and the image feature map. For example, the image feature map may include at least one sub-region, and the similarity between the partial image feature map of each sub-region and the example feature map is calculated. Using the similarity between the partial image feature map of each sub-region and the example feature map as a weight, the partial image feature map of each sub-region of the image feature map is updated to obtain the intermediate feature map corresponding to the example feature map.

[0068] Among them, after obtaining the intermediate feature maps corresponding to each of the example feature maps, the intermediate feature maps corresponding to each of the example feature maps can be fused by fusion methods such as weighting and splicing to obtain the refined feature map.

[0069] In some embodiments, each example feature map can be fused to obtain a fused example feature map. The fused example feature map retains the feature information of the object to be counted in each example feature map. Based on the fused feature map and the image feature map, the refined feature map is obtained.

[0070] S104. Convert the refined feature map into a density estimation map, and generate a statistical result for the object to be counted based on the density estimation map.

[0071] In some embodiments, the statistical result may include at least one of the following: the statistical quantity of the object to be counted, the position information of each object to be counted in the image to be processed.

[0072] Based on the above embodiments, a corresponding refined feature map is generated through the image feature map and each of the example feature maps. Since the refined feature map retains the features related to each of the example feature maps on the basis of the image feature map, furthermore, in the process of determining the statistical result for the object to be counted based on the refined feature map, a statistical result corresponding to the object to be counted can be obtained, improving the statistical accuracy.

[0073] See Figure 2 , Figure 2 is an optional flowchart of the object counting method provided by the embodiments of the present disclosure. Based on Figure 1 , Figure 1 in, S103 can be updated to S201 to S203, which will be combined withFigure 2 The steps shown will be described.

[0074] S201. Generate a relevance feature corresponding to each of the example feature maps based on the image feature map and each of the example feature maps; the relevance feature is used to characterize the degree of relevance between the image feature map and the example feature map.

[0075] In some embodiments, each of the example feature maps can be used to perform convolutional processing on the image feature map to obtain a relevance feature corresponding to each of the example feature maps.

[0076] S202. Obtain a correlation feature map based on the relevance feature corresponding to each of the example feature maps and the corresponding example feature map.

[0077] In some embodiments, since the relevance feature corresponding to each example feature map characterizes the degree of relevance between the image feature map and the example feature map, therefore, the relevance feature corresponding to each example feature map can be used as a weight to weight the example feature map to obtain a sub-correlation feature map corresponding to each example feature map, and the correlation feature map is obtained by fusing the sub-correlation feature maps corresponding to each example feature map.

[0078] It should be noted that the size of the correlation feature map is the same as the size of the image feature map.

[0079] S203. Fuse the correlation feature map and the image feature map to obtain a refined feature map.

[0080] In some embodiments, at least one of the fusion means such as splicing, weighting, element-wise addition, etc. can be used to fuse the correlation feature map and the image feature map to obtain a refined feature map.

[0081] Through the above embodiments, since the problem of gradient disappearance in the model training process can be avoided by fusing the low-level image feature map and the high-level correlation feature map; at the same time, since the relevance feature corresponding to each example feature map is generated based on the image feature map and each example feature map, not only the relevance of the objects to be counted in different example images can be considered, but also the degree of relevance between the image feature map and the corresponding example feature map can be determined by the obtained relevance feature.

[0082] See Figure 3 , Figure 3 is an optional process schematic diagram of the object counting method provided by the embodiments of the present disclosure. Based on Figure 2 , Figure 2 S2011 in Figure 3 The steps shown will be described.

[0083] S301. Process the image feature map through a shared normalization layer to obtain an intermediate image feature map.

[0084] S302. For each example feature map, process the example feature map through the shared normalization layer to obtain an intermediate example feature map; use the intermediate example feature map as a convolution kernel and the intermediate image feature map as the object to be convolved to obtain the correlation feature corresponding to the example feature map.

[0085] Among them, each example feature map corresponds to a correlation feature.

[0086] In some embodiments, after obtaining the correlation feature corresponding to each example feature map, the method may further include S303.

[0087] S303. Perform normalization processing on the correlation feature corresponding to each example feature map.

[0088] Through the above embodiments, before using the example feature map as a convolution kernel to normalize the image feature map, first perform normalization processing on the example feature map and the image feature map respectively through the shared normalization layer, which can reduce the computational complexity in the convolution process.

[0089] See Figure 4 , Figure 4 which is an optional process schematic diagram of the object statistics method provided by the embodiments of the present disclosure. Based on Figure 3 , Figure 3 in which S303 can be updated to S401 to S402, and will be described in combination with the steps shown in Figure 4 .

[0090] S401. Perform normalization processing on the correlation feature corresponding to each example feature map from at least one dimension to obtain the correlation feature to be fused for each example feature map in each dimension.

[0091] In some embodiments, the at least one dimension may include an example dimension and a spatial dimension. Among them, based on the example dimension, the example normalization feature corresponding to the example dimension can be obtained; based on the spatial dimension, the spatial normalization feature corresponding to the spatial dimension can be obtained.

[0092] It should be noted that for the correlation feature corresponding to each example feature map, in the process of performing normalization processing on the correlation feature from one dimension, the normalization feature corresponding to each example feature map in this dimension can be obtained.

[0093] In some embodiments, the at least one dimension includes an example dimension. The normalization process of the relevance features corresponding to each example feature map can be implemented through S4011 to S4022 to obtain the relevance features to be fused for each example feature map in each dimension.

[0094] S4011. For each feature position in the relevance features corresponding to the example feature map, obtain the original relevance of the relevance features corresponding to each example feature map at the feature position; perform a normalization process on the original relevance corresponding to each example feature map to obtain the updated relevance of the relevance features corresponding to each example feature map at the feature position.

[0095] In some embodiments, the relevance features corresponding to the example feature map may include multiple feature positions. Taking the case where the size of the relevance features corresponding to the example feature map is H×W as an example, the relevance features may include H×W feature positions. That is to say, each feature position corresponds to an original relevance, and the original relevances corresponding to the H×W feature positions constitute the relevance features.

[0096] Correspondingly, since the sizes of the relevance features corresponding to each example feature map are the same, for each feature position in the relevance features, the corresponding original relevance can be found in the relevance features corresponding to each example feature map. For this feature position, the updated relevance corresponding to this feature position is obtained based on the original relevance corresponding to this feature position in the relevance features corresponding to each example feature map.

[0097] Exemplarily, taking the feature position (0, 0) as an example, in the process of determining the updated relevance corresponding to the feature position (0, 0), it is necessary to obtain the original relevance corresponding to the feature position (0, 0) in the relevance features corresponding to each example feature map. For example, the original relevance of the first example feature map at the feature position (0, 0) is C1, the original relevance of the second example feature map at the feature position (0, 0) is C2,..., and the original relevance of the Nth example feature map at the feature position (0, 0) is CN. Then the updated relevance of the first example feature map at the feature position (0, 0) is C1 / sum(C1, C2,..., CN), the updated relevance of the second example feature map at the feature position (0, 0) is C2 / sum(C1, C2,..., CN),..., and the updated relevance of the Nth example feature map at the feature position (0, 0) is CN / sum(C1, C2,..., CN).

[0098] S4012. Generate the example normalization features corresponding to each example feature map based on the updated relevance of the relevance features corresponding to each example feature map at each feature position.

[0099] In some embodiments, after obtaining the updated relevance at each feature position of the relevance feature corresponding to each example feature map based on the above method, for each example feature map, the updated relevance of all feature positions corresponding to the example feature map can be used as the example normalization feature corresponding to the example feature map.

[0100] Exemplarily, based on the above example, for each example feature map, the updated relevance corresponding to H×W feature positions can be obtained, and the updated relevance corresponding to these H×W feature positions constitutes the example normalization feature corresponding to the example feature map.

[0101] In some embodiments, the at least one dimension may include a spatial dimension, and S4013 is used to perform normalization processing on the relevance feature corresponding to each example feature map from at least one dimension, so as to obtain the normalized relevance feature corresponding to each example feature map in each dimension of the at least one dimension.

[0102] S4013: For each example feature map, obtain the original relevance of the relevance feature corresponding to the example feature map at each feature position, perform normalization processing on the original relevance at each feature position to obtain the updated relevance at each feature position; based on the updated relevance at each feature position, generate the spatial normalization feature corresponding to the example feature map.

[0103] In some embodiments, the relevance feature corresponding to the example feature map may include multiple feature positions. Taking the case where the size of the relevance feature corresponding to the example feature map is H×W as an example, the relevance feature may include H×W feature positions, that is, each feature position corresponds to an original relevance, and the original relevance corresponding to H×W feature positions constitutes the relevance feature.

[0104] Different from the above example dimension, in the process of performing normalization processing on the relevance feature corresponding to each example feature map based on the spatial dimension, the relationship between different example images (example feature maps) is not considered. That is, in the process of normalizing the relevance feature of an example feature map, only the relevance feature itself of this example feature map is considered.

[0105] For each example feature map (corresponding relevance feature), obtain the original relevance of the relevance feature corresponding to the example feature map at each feature position, where the updated relevance at each feature position is related to all the original relevance corresponding to the example feature map.

[0106] Exemplarily, taking the case where the size of the relevance feature corresponding to an exemplary feature map is H×W as an example, for each feature position in this relevance feature, that is, including a total of H×W feature positions from (0, 0) to (H - 1, W - 1), the original relevance at each feature position can be obtained first, and the maximum original relevance feature degree can be determined. For each feature position, the updated relevance of this feature position is the ratio between the original feature degree and the maximum original relevance feature degree. By analogy, the spatially normalized feature corresponding to each exemplary feature map can be obtained.

[0107] S402. Generate the normalized relevance feature corresponding to the exemplary feature map based on the normalized relevance features corresponding to each exemplary feature map corresponding to each dimension.

[0108] In some embodiments, the above-mentioned generating the normalized relevance feature corresponding to the exemplary feature map based on the normalized relevance features corresponding to each exemplary feature map corresponding to each dimension can be implemented through S4021.

[0109] S4021. For each exemplary feature map, fuse the normalized relevance features corresponding to each dimension based on element-wise multiplication to generate the normalized relevance feature corresponding to the exemplary feature map.

[0110] Among them, the sizes of the normalized relevance features corresponding to each dimension of the exemplary feature map are the same. In the case where the at least one dimension includes a spatial dimension and an exemplary dimension, the sizes of the obtained exemplary normalized feature and the spatial normalized feature are the same.

[0111] In some embodiments, the exemplary normalized feature and the spatial normalized feature can be fused based on element-wise multiplication to obtain the normalized relevance feature corresponding to the exemplary feature map.

[0112] Based on the above embodiments, after obtaining the relevance features corresponding to each example feature map, normalization processing is also performed on the relevance features, which can not only reduce the computational complexity, accelerate the model convergence speed, but also avoid gradient saturation. Moreover, the embodiments of the present disclosure also perform normalization on the relevance features corresponding to multiple example feature maps based on the example dimension, and the normalized relevance features obtained can take into account the differences between each example image (example feature map); at the same time, the embodiments of the present disclosure also perform normalization on the relevance features corresponding to each example feature map based on the spatial dimension, and the normalized relevance features obtained can reflect the degree of relevance between the image feature map at different positions and the example feature map; correspondingly, after fusing the example normalization features corresponding to the example dimension and the spatial normalization features corresponding to the spatial dimension, the processed normalized relevance features can not only take into account the differences between each example image (example feature map), but also reflect the degree of relevance between the image feature map at different positions and the example feature map.

[0113] See Figure 5 , Figure 5 is an optional flowchart of the object statistics method provided by the embodiments of the present disclosure. Based on any of the above embodiments, taking Figure 1 as an example, after S103 in Figure 1 , the method further includes S501, and S104 can be updated to S502, which will be described in conjunction with Figure 5 the steps shown.

[0114] S501. Use the refined feature map as the image feature map and return to the step of generating the refined feature map based on the image feature map and the example feature map for iterative processing.

[0115] In some embodiments, after obtaining the refined feature map based on the image to be processed and at least one example image, one iteration process is completed. That is to say, one iteration process includes: generating a refined feature map based on the image feature map and each example feature map.

[0116] Among them, after one iteration process is completed, the refined feature map obtained from the previous iteration process can be used as the image feature map corresponding to the image to be processed in the current iteration process, and a refined feature map (of the current iteration process) is generated based on this new image feature map (the refined feature map obtained from the previous iteration process) and each example feature map.

[0117] In some embodiments, a preset number of iterative processing times can be obtained. Before executing S501, it can be first determined whether the current iterative processing times has reached the preset number of iterative processing times. In the case where the preset number of iterative processing times is reached, the refined feature map (obtained from the previous iterative processing) is directly used as the final refined feature map, and a statistical result is determined based on the final refined feature map; in the case where the preset number of iterative processing times is not reached, the refined feature map obtained from the previous iterative processing is used as the image feature map corresponding to the image to be processed in the current iterative processing, and a refined feature map (for the current iterative processing) is generated based on the new image feature map (the refined feature map obtained from the previous iterative processing) and each of the example feature maps.

[0118] S502. Convert the refined feature map after iterative processing into a density estimation map, and generate a statistical result for the object to be counted based on the density estimation map.

[0119] In some embodiments, the refined feature map after iterative processing is used as the final refined feature map.

[0120] Based on the above embodiments, by using the obtained refined feature map as the image feature map of the original image to be processed and obtaining a new refined feature map with at least one example feature map, thus, through at least one iterative refinement process, the refined feature map can be continuously refined, so that the correlation between the final refined feature map and the example feature map is higher, and thus a more accurate statistical result can be obtained.

[0121] See Figure 6 , Figure 6 is an optional flowchart of the object counting method provided by the embodiments of the present disclosure. Based on any of the above embodiments, taking Figure 1 as an example, Figure 1 S104 in Figure 6 can be updated to S601 to S602, and will be described in conjunction with the steps shown in

[0122] S601. Convert the refined feature map into a density estimation map.

[0123] In some embodiments, the above conversion of the refined feature map into a density estimation map can be implemented through S6011 to S6012.

[0124] S6011. Perform iterative convolution processing on the refined feature map.

[0125] S6012. Upsample the refined feature map after iterative convolution processing to obtain a density estimation map with the same size as the image to be processed.

[0126] S602. Generate a statistical result for the object to be counted based on the density estimation map.

[0127] In some embodiments, the above-mentioned generation of the statistical result for the object to be counted based on the density estimation map can be achieved through S6021.

[0128] S6021. Sum the density estimation map to obtain the statistical quantity of the object to be counted in the image to be processed.

[0129] In some embodiments, the values at each position in the density estimation map are between 0 and 1. When there is an object to be counted at a position in the image to be processed, the larger the value at this position in the density estimation map and the closer it is to 1. Among them, if there is no object to be counted in the image to be processed, the values at all positions in the corresponding density estimation map are 0; when there is an object to be counted at (X, Y) in the image to be processed, this object to be counted can affect the value size of the affected area corresponding to (X, Y) in the density estimation map. For example, it can affect the value sizes of all positions within the affected area with (X, Y) as the center and radius r, and the value at (X, Y) is the largest, and the values of other positions in the affected area become smaller as the distance from the center increases. It should be noted that in this affected area, the sum of the values at all positions is 1, that is, an object to be counted in the image to be processed can increase the sum of the values at all positions in the density estimation map by 1. Thus, the statistical quantity of the object to be counted in the image to be processed can be obtained by summing the density estimation map.

[0130] In some embodiments, the above-mentioned generation of the statistical result for the object to be counted based on the density estimation map can be achieved through S6022.

[0131] S6022. Obtain the local peak points in the density estimation map, and screen the local peak points through a non-maximum suppression algorithm to obtain the positions of the objects to be counted; the pixel values of the local peak points are greater than the pixel values of adjacent pixel points.

[0132] In some embodiments, since the larger the value at each position in the density estimation map, the higher the probability that there is an object to be counted at this position in the image to be processed. Therefore, the preliminary positions of the objects to be counted can be determined by obtaining the local peak points in the density estimation map. Considering the situation of errors, that is, for the same object to be counted in the image to be processed, there are at least two corresponding local peak points in the density estimation map. Therefore, it is necessary to use a non-maximum suppression algorithm to screen the obtained local peak points, and then obtain the positions of the objects to be counted.

[0133] Through the above embodiments, not only can the statistical quantity of the object to be counted in the image to be processed be obtained, but also the position of each object to be counted in the image to be processed can be obtained, improving the applicability of the present disclosure.

[0134] See Figure 7 , Figure 7 , which is an optional flowchart of the model training method provided by the embodiments of the present disclosure, will be described in conjunction with Figure 7 the steps shown.

[0135] S701. Obtain a sample image, the true density map corresponding to the sample image, and at least one example image corresponding to the object to be counted.

[0136] S702. Input the sample image and the at least one example image into a statistical model to be trained to obtain a predicted density map; the statistical model is used to extract the image feature map corresponding to the sample image and the example feature map corresponding to each of the at least one example image; generate a refined feature map based on the image feature map and each example feature map, and the refined feature map retains the features related to each example feature map on the basis of the image feature map; and convert the refined feature map into the predicted density map.

[0137] S703. Adjust the parameters of the statistical model to be trained based on the predicted density map and the true density map to obtain a trained statistical model.

[0138] See Figure 8A , Figure 8A which is an optional flowchart of the model training method provided by the embodiments of the present disclosure. Based on Figure 7 , Figure 7 S702 in Figure 8A may include S801 to S804, and will be described in conjunction with Figure 8B the steps shown. Please refer to

[0139] S801. Input the sample image and the at least one example image into the feature extraction network to obtain the image feature map corresponding to the sample image and the example feature map corresponding to each of the at least one example image;

[0140] S802. Input the image feature map and each example feature map into the relevance calculation network to obtain the relevance feature corresponding to each example feature map; the relevance feature is used to characterize the degree of relevance between the image feature map and the example feature map;

[0141] S803. Input the image feature map, each of the example feature maps, and the relevance feature corresponding to each of the example feature maps into the feature refinement network to obtain the refined feature map; the refined feature map retains the features related to each of the example feature maps on the basis of the image feature map;

[0142] S804. Input the refined feature map into the density regression network to obtain the predicted density map.

[0143] See Figure 9A , Figure 9A is an optional process schematic diagram of the model training method provided by the embodiments of the present disclosure. Based on Figure 8A , Figure 8A S802 in Figure 9A may include S901 to S903, which will be described in conjunction with the steps shown in Figure 9B . Please refer to Figure 9B , which shows a schematic structural diagram of a relevance calculation network provided by the embodiments of the present disclosure.

[0144] S901. Input the image feature map and each of the example feature maps into the first normalization layer to obtain the intermediate image feature map corresponding to the image feature map and the intermediate example feature maps corresponding to each of the example feature maps;

[0145] S902. Use each of the intermediate example feature maps as a convolution kernel and use the intermediate image feature map as the object to be convolved to obtain the relevance feature corresponding to each of the example feature maps;

[0146] S903. Input the relevance feature corresponding to each of the example feature maps into the second normalization layer to obtain the normalized relevance feature corresponding to the example feature map.

[0147] See Figure 10A , Figure 10A is an optional process schematic diagram of the model training method provided by the embodiments of the present disclosure. Based on Figure 9A , Figure 9A S9003 in Figure 10A may include S1001 to S1002, which will be described in conjunction with the steps shown in Figure 10B . Please refer to Figure 10B , which shows a schematic structural diagram of another relevance calculation network provided by the embodiments of the present disclosure.

[0148] S1001. Input the relevance feature corresponding to each of the example feature maps into the sub-normalization layer corresponding to each dimension to obtain the normalized relevance feature corresponding to each example feature map corresponding to each dimension.

[0149] In some embodiments, the above-mentioned step of inputting the relevance features corresponding to each of the example feature maps into the sub-normalization layer corresponding to each dimension to obtain the normalized relevance features corresponding to each example feature map in each dimension can be implemented through S10011.

[0150] S10011: Input the relevance features corresponding to each of the example feature maps into the example sub-normalization layer to obtain the example-normalized features corresponding to each of the example feature maps.

[0151] Among them, the example sub-normalization layer is used to, for each feature position in the example feature map, obtain the original relevance of the relevance features corresponding to each of the example feature maps at the feature position, perform normalization processing on the original relevance corresponding to each of the example feature maps to obtain the updated relevance of the relevance features corresponding to each of the example feature maps at the feature position; and generate the example-normalized features corresponding to each of the example feature maps based on the updated relevance of the relevance features corresponding to each of the example feature maps at each feature position.

[0152] In some embodiments, the above-mentioned step of inputting the relevance features corresponding to each of the example feature maps into the sub-normalization layer corresponding to each dimension to obtain the normalized relevance features corresponding to each example feature map in each dimension can be implemented through S10012.

[0153] S10012: Input the relevance features corresponding to each of the example feature maps into the spatial sub-normalization layer to obtain the spatial-normalized features corresponding to each of the example feature maps.

[0154] Among them, the spatial sub-normalization layer is used to, for each of the example feature maps, obtain the original relevance of the relevance features corresponding to the example feature map at each feature position, perform normalization processing on the original relevance at each feature position to obtain the updated relevance at each feature position; and generate the spatial-normalized features corresponding to the example feature map based on the updated relevance at each feature position.

[0155] S1002: Input the relevance features corresponding to each of the example feature maps into the sub-normalization layer corresponding to each dimension to obtain the normalized relevance features corresponding to each example feature map in each dimension.

[0156] In some embodiments, the above-mentioned S1002 can be implemented through S10021.

[0157] S10021. For each of the example feature maps, input the correlation features after normalization corresponding to each dimension into the sub-fusion layer to obtain the correlation features after normalization corresponding to the example feature map.

[0158] See Figure 11A , Figure 11A is an alternative flowchart of the model training method provided by the embodiments of the present disclosure. Based on Figure 8A , Figure 8A S803 in can include S1101 to S1102, which will be described in conjunction with the steps shown in Figure 11A . Please refer to Figure 11B , which shows a schematic structural diagram of a feature refinement network provided by the embodiments of the present disclosure.

[0159] S1101. Input each of the example feature maps and the correlation features corresponding to each of the example feature maps into the correlation feature map generation layer to obtain a correlation feature map.

[0160] In some embodiments, the above-mentioned inputting each of the example feature maps and the correlation features corresponding to each of the example feature maps into the correlation feature map generation layer to obtain a correlation feature map can be implemented through S11011 to S11013.

[0161] S11011. For each of the example feature maps, input the example feature map into the feature flipping layer to obtain a flipped example feature map;

[0162] S11012. Input the flipped example feature map and the correlation features corresponding to the example feature map into the second convolutional layer to obtain a sub-correlation feature map corresponding to the example feature map;

[0163] S11013. Input the sub-correlation feature map corresponding to each of the example feature maps into the second sub-fusion layer to generate the correlation feature map.

[0164] S1102. Input the correlation feature map and the image feature map into the iterative fusion layer to obtain the refined feature map.

[0165] Among them, the iterative fusion layer is used to perform fusion processing on the correlation feature map and the image feature map to obtain the refined feature map; the fusion processing includes at least one of the following: skip connection, convolution, and normalization layer processing.

[0166] See Figure 12A , Figure 12A is an alternative flowchart of the model training method provided by the embodiments of the present disclosure. Based on any of the above embodiments, taking Figure 8A as an example,Figure 8A After S803 in Figure 12A it may further include S1201, which will be described in conjunction with the steps shown in

[0167] S1201: Take the refined feature map as the image feature map, and return the step of inputting the image feature map and each example feature map into the relevance calculation network to obtain the relevance feature corresponding to each example feature map for iterative processing.

[0168] S1202: Input the refined feature map after iterative processing into the density regression network to obtain the predicted density map.

[0169] Please refer to Figure 12B , which shows a schematic structural diagram of another statistical model provided by an embodiment of the present disclosure. It can be seen that, compared with the statistical model provided by Figure 8B After obtaining the refined feature map, the feature refinement network 813 does not directly input it into the density regression network 814. Instead, the refined feature map is used as the image feature map for iterative processing, that is, based on this new image feature map (refined feature map) and the example feature map corresponding to each example image, it passes through the relevance calculation network 812 and the feature refinement network 813 in sequence to obtain the refined feature map after iterative processing, and input the refined feature map after iterative processing into the density regression network to obtain the predicted density map.

[0170] Among them, the number of iterations can be set to 2 to 4 times.

[0171] Next, an exemplary application of the embodiments of the present application in an actual application scenario will be described.

[0172] In recent years, deep learning algorithms have made great progress in various fields and have also been applied to many scene analysis fields. In the field of scene analysis, counting dense objects is a very important task. Currently, there are many fully automatic image-based counting methods based on deep learning, such as counting people, counting goats, counting airport luggage, etc. These methods use deep learning models to use image information to quickly and automatically count specific objects, and have been widely applied in scenarios such as people counting in the monitoring field, animal counting in nature conservation applications, and workpiece counting in factory industrial manufacturing.

[0173] However, existing deep learning-based counting algorithms often require a large amount of high-quality labeled data. For object counting tasks lacking training data, it is difficult to collect data and train a good counting model. Additionally, existing counting algorithms are all specific to certain categories, such as a model for crowd counting, which can only count people and cannot be migrated to other counting scenarios. A small-sample general counting algorithm only requires a small amount of labeled data to count various objects, which is very important.

[0174] Embodiments of the present disclosure provide a general counting system based on iterative refined features. Refer to Figure 13 , Figure 13 , which is a schematic diagram of the system architecture of the general counting system 130 in the embodiments of the present disclosure. The general counting system includes a data module 131, a network module 132, a training module 133, and an inference module 134.

[0175] The data module 131 is used to obtain image data and perform data annotation and processing. The images obtained by the data module 131 may contain many (1 to 5000) same-kind objects (such as birds, people, cars, books, etc.). There may be differences between different individuals of the same-kind objects. Among them, in the specific implementation process of the data module 131, it can be divided into a training scenario and an inference scenario.

[0176] In the training scenario, after obtaining an image, it is necessary to obtain first annotation data for the target objects to be counted in the image. The first annotation data may include rectangular bounding box annotation and position point annotation of the target objects. In some implementation scenarios, the rectangular bounding box annotation may include N rectangular bounding boxes, where N is less than or equal to 3; the position point annotation is the position points of all target objects in the image. Refer to Figure 14 , Figure 14 , which is a schematic diagram of the annotation data provided by the embodiments of the present disclosure.

[0177] As Figure 14 shown, the target objects to be counted in the image are cows. The annotation data corresponding to the image may include three manually annotated rectangular bounding box annotations of cows, including A11, A12, and A13, and position point annotations of all cows, including B21, B22, B23, and B24. In the following embodiments, the objects in the manually annotated rectangular annotation boxes are called example objects.

[0178] In the inference scenario, after obtaining an image, it is necessary to obtain second annotation data for the target objects to be counted in the image. The second annotation data may include rectangular bounding box annotation. In some implementation scenarios, the rectangular bounding box annotation may include N rectangular bounding boxes, where N is less than or equal to 3.

[0179] Note that the example objects can come not only from the current image to be counted, but also from other images. For example, for the cows in the above Figure 14 example, the example objects can be not only the cows in Figure 14 , but also the cows in other images except Figure 14 .

[0180] The network module 132 is the core module of the above general counting system. Please refer to Figure 15 , which shows a schematic diagram of the processing flow of the network module 132. It can be seen that the network module 132 at least includes the following sub-modules: a feature extraction module 151, a feature fusion module 152, a correlation calculation module 153, a feature refinement module 154, a density regression module 155, and a post-processing module 156.

[0181] In some embodiments, the feature extraction module 151 can be a convolutional neural network pre-trained on the ImageNet dataset, such as VGG, ResNet, etc. Among them, the feature extraction module 151 takes an image as input and outputs a multi-layer feature map of the image. Then, using region of interest pooling (ROI Pooling), a multi-layer feature map of its example object is obtained.

[0182] In some embodiments, the feature fusion module 152 is used to fuse the features of different layers in the obtained multi-layer feature maps to obtain a fused feature with a consistent size. Among them, the feature fusion method can include but is not limited to a feature pyramid network, size transformation, and splicing, etc.

[0183] In this embodiment, the feature fusion module 152 can obtain the multi-layer feature map of the image output by the feature extraction module 151 and the multi-layer feature map of the example object. Since the feature maps are multi-layered and their sizes are inconsistent, they cannot be processed uniformly. Therefore, the feature fusion module 152 can perform feature fusion on the multi-layer feature map of the image and the multi-layer feature map of the example object respectively to obtain a fused feature map of the image and a fused feature map of the example object. For the convenience of description, in the following embodiments, the fused feature map of the image is simply referred to as the image feature map, and the fused feature map of the example object is simply referred to as the example feature map.

[0184] In some embodiments, the correlation calculation module 153 is used to calculate the correlation between the image feature map and the example feature map to obtain a corresponding correlation feature. The correlation feature represents the similarity between the image feature map and the example feature map at different positions, and the size of the correlation feature is the same as that of the image feature map.

[0185] Among them, the convolution method can be adopted, taking the image feature map as the object to be convolved and the example feature map as the convolution kernel to determine the correlation between the image feature map and the example feature map. The cosine similarity between the image feature map and the example feature map can also be calculated as the correlation between the image feature map and the example feature map.

[0186] In some embodiments, the correlation calculation module 153 needs to perform normalization processing on the obtained correlation features to obtain the correlation features after normalization processing.

[0187] In some embodiments, the feature refinement module 154 is used to obtain the target refined feature map corresponding to the example feature map based on the correlation features after normalization processing and the example feature map.

[0188] In this embodiment, the correlation feature map can be obtained first based on the correlation features after normalization processing and the example feature map. Among them, the obtained correlation features after normalization processing can be used as weights to fuse the example feature map to reconstruct the correlation feature map. The characteristic of the correlation feature map is that the closer the position is to the example feature map (example image), the greater the correlation corresponding to this position in the correlation features after normalization processing, and the better the features at this position will be preserved in the correlation feature map. Therefore, the correlation feature map is essentially a feature refined based on the correlation.

[0189] After that, by fusing the image feature map and the correlation feature map, a refined feature map can be obtained. Among them, the above feature map fusion methods can include but are not limited to at least one of the following: skip connection, convolution, layer regularization, etc. In this embodiment, the refined feature map can be directly used as the target refined feature map corresponding to the example feature map.

[0190] In some embodiments, after obtaining the refined feature map, the refined feature map can be re-input into the correlation calculation module 153 as the image feature map to obtain the new correlation features after normalization processing; the new correlation features after normalization processing are input into the feature refinement module 154 to obtain a new refined feature map. This embodiment is equivalent to performing further feature refinement. After at least two iterations, the obtained new refined feature map is used as the target refined feature map corresponding to the example feature map, where the obtained target refined feature map meets the refinement requirements for subsequent density regression.

[0191] In some embodiments, when the number of iterations is set to 2 to 4 times, the requirements of processing efficiency and refinement can be met simultaneously.

[0192] In some embodiments, the density regression module 155 is used to convert the target refined feature map obtained by the feature refinement module 154 into a density estimation map. The density regression module consists of multiple layers of convolutional layers, upsampling layers, and Relu activation layers. After the target refined feature map is input into the density regression module 155, an intermediate density map with 1 channel is generated, and then the intermediate density map is enlarged to the size of the original image through resampling to obtain a density estimation map corresponding to the original image.

[0193] In some embodiments, the post-processing module is used to generate a count estimation result and a position estimation result based on the density estimation map. The count estimation result is the number of objects in the entire original image, and the position estimation result is the position points of each object in the entire original image.

[0194] Among them, by summing the density estimation map, a count estimation result of the number of objects in the entire original picture can be obtained. Search for local peak points (i.e., points where the pixel value is greater than the surrounding 8 pixels) of the density estimation map to obtain a set of local peak points, and then perform non-maximum suppression (NMS) on the set of local peak points. The coordinates of the remaining points in the set of local peak points are the position estimation results of all objects.

[0195] Please refer to Figure 16 , which shows a schematic diagram of a network structure.

[0196] After obtaining the image feature map f I and the example feature map f e , it is necessary to use the example feature map f e as the convolution kernel, and use the image feature map f I as the object to be convolved. After convolution processing, the correlation feature (A) between the image feature map f I and the example feature map f e is obtained.

[0197] In the process of obtaining the correlation feature, the inventor found that due to the different scales of different image feature maps f I and example feature maps f e , it is not only unstable during implementation but also prone to divergence during training. Therefore, it is necessary to set a shared LN layer for the image feature map f I and the example feature map f e before convolution processing to solve the above-mentioned instability problem.

[0198] In some embodiments, the correlation feature A can be obtained by the following formula (1):

[0199]

[0200] Among them, A is the relevance feature, LN(.) is the shared LN layer, and f I is the image feature map, and f e is the exemplar feature map. K represents the number of exemplar feature maps, and H and W are the sizes of the relevant feature map. It should be noted that the size of this relevance feature is the same as the size of the original image.

[0201] In some embodiments, the relevance feature is respectively input into the Exemplar Norm (EN) and the Spatial Norm (SN), and the exemplar normalization feature A e corresponding to the relevance feature and the spatial normalization feature A s .

[0202] The Exemplar Norm (EN) is used to normalize the normalization feature from the dimensions of different exemplar objects. In some embodiments, the exemplar normalization feature A e can be obtained through formula (2):

[0203]

[0204] where softmax(., dim) is the softmax layer corresponding to the current dimension.

[0205] The inventors found that in the case of only EN, the sum of the relevance values of different samples at any position is 1, which means that the sum of the relevance values at high-relevance positions is equal to the sum of the relevance values at low-relevance positions. Therefore, in a further embodiment, the present disclosure embodiment solves this problem by introducing the Spatial Norm (SN). SN is used to normalize between different spatial positions. After introducing the SN, the sum of the relevance values of different samples in high-relevance positions is 1, and other positions are in [0, 1].

[0206] In some embodiments, the spatial normalization feature A s can be obtained through formula (3):

[0207]

[0208] where A s is the spatial normalization feature, and max(., dim) is used to find the maximum relevance value in the current dimension.

[0209] In some embodiments, after obtaining the above exemplar normalization feature A e and the spatial normalization feature A sAfter that, it is necessary to fuse the normalized features in these two dimensions to obtain the final normalized relevant feature map A. n Exemplarily, the above fusion process can be completed by element-wise multiplication, as shown in Formula (4):

[0210]

[0211] where A n is the final normalized correlation, which will be referred to as the normalized correlation feature in the following embodiments. denotes element-wise multiplication. the correlated feature map

[0212] In some embodiments, at the positions where the normalized correlation feature is more similar to the example image, the corresponding correlation is greater, and the correlation between the features at these positions and the samples is higher. Furthermore, the features at these positions need to be retained best. Therefore, the normalized correlation feature A n can be used as the weight for combining the example feature map to obtain the relevant feature map f c , as shown in Formulas (5) and (6).

[0213]

[0214]

[0215] where sum(.,dim) is used to determine the sum in the current dimension, and flip(f e ) is used to flip f e horizontally and vertically. It should be noted that the purpose of flipping the example feature map is to make the obtained relevant feature map f c have the same spatial structure as the example feature map.

[0216] In some embodiments, after obtaining the relevant feature map f c , it is necessary to fuse the relevant feature map f c and the above image feature map f I to obtain the fused refined feature map. Among them, the relevant feature map f c and the above image feature map f I can be added and passed through an LN layer to obtain the fused refined feature map. Through the skip connection method in the embodiments of the present disclosure, since the shallow image feature map f I and the deep relevant feature map f c, the problem of gradient disappearance is avoided.

[0217] Through the above process, based on the original image feature map f I and the example feature map f e the refined feature map f I ′ .

[0218] In a further embodiment, in order to further refine the refined feature map, the obtained refined feature map f I ′ can be used as a new image feature map, and the above process is iteratively executed using the new image feature map and the example feature map to obtain a new refined feature map f I ′ . Wherein, the number of iterations can be set to 2 to 4 times.

[0219] It should be noted that in the process of iteratively executing the above process using the new image feature map and the example feature map, the example feature map in the current iterative processing can be the original example feature map, or a new example feature map obtained based on the new image feature map obtained in the current iterative process. Exemplarily, taking the original image feature map and the example feature map as the first image feature map and the first example feature map respectively, after the first iterative processing, the first refined feature map corresponding to the current iterative processing can be obtained based on the first image feature map and the first example feature map. At this time, the second iterative processing is required. In some embodiments, the first refined feature map can be used as the second image feature map in the second iterative processing, and the second iterative processing process can be completed in combination with the original first example feature map to obtain the second refined feature map; in other embodiments, the first refined feature map can be used as the second image feature map in the second iterative processing, and a second example feature map in the second iterative processing can be generated based on the second image feature map, and the second refined feature map can be obtained by combining the second image feature map and the second example feature map.

[0220] In some embodiments, after at least two iterative processes, since the iterative process can refine and enhance the features related to the example and suppress the features unrelated to the example, a target refined feature map without interference from unrelated features can be obtained. Then, the target refined feature map can be directly converted into a density feature map through a regression head wherein, the regression head can include an iterative convolution layer, a ReLU activation layer, and a bi-linear upsample layer.

[0221] Figure 17 The following is a schematic structural diagram of an object statistics device provided by an embodiment of the present disclosure. As Figure 17 shown, the object statistics device 1700 includes:

[0222] A first acquisition module 1701, configured to acquire at least one example image corresponding to an image to be processed and an object to be counted;

[0223] An extraction module 1702, configured to extract an image feature map corresponding to the image to be processed and an example feature map corresponding to each of the at least one example image;

[0224] A generation module 1703, configured to generate a refined feature map based on the image feature map and each example feature map; the refined feature map retains features related to each example feature map on the basis of the image feature map;

[0225] A conversion module 1704, configured to convert the refined feature map into a density estimation map, and generate a statistical result for the object to be counted based on the density estimation map.

[0226] In some embodiments, generating the refined feature map based on the image feature map and each example feature map includes:

[0227] Generating a correlation feature corresponding to each example feature map based on the image feature map and each example feature map; the correlation feature is used to characterize the correlation degree between the image feature map and the example feature map;

[0228] Obtaining a correlation feature map based on the correlation feature corresponding to each example feature map and the corresponding example feature map;

[0229] Fusing the correlation feature map and the image feature map to obtain a refined feature map.

[0230] In some embodiments, generating the correlation feature corresponding to each example feature map based on the image feature map and each example feature map includes:

[0231] Performing convolution processing on the image feature map using each example feature map to obtain a correlation feature corresponding to each example feature map.

[0232] In some embodiments, performing convolution processing on the image feature map using each example feature map to obtain a correlation feature corresponding to each example feature map includes:

[0233] Processing the image feature map through a shared normalization layer to obtain an intermediate image feature map;

[0234] For each example feature map, process the example feature map through the shared normalization layer to obtain an intermediate example feature map; use the intermediate example feature map as a convolution kernel and the intermediate image feature map as the object to be convolved to obtain the relevance feature corresponding to the example feature map.

[0235] In some embodiments, before obtaining the correlation feature map based on the relevance feature corresponding to each example feature map and the corresponding example feature map, the method further includes: performing a normalization process on the relevance feature corresponding to each example feature map;

[0236] The performing a normalization process on the relevance feature corresponding to each example feature map includes: performing a normalization process on the relevance feature corresponding to each example feature map from at least one dimension to obtain the normalized relevance feature corresponding to each example feature map for each dimension in the at least one dimension; generating the normalized relevance feature corresponding to the example feature map based on the normalized relevance feature corresponding to each example feature map for each dimension.

[0237] In some embodiments, the at least one dimension includes an example dimension, and the performing a normalization process on the relevance feature corresponding to each example feature map from at least one dimension to obtain the normalized relevance feature corresponding to each example feature map for each dimension in the at least one dimension includes:

[0238] For each feature position in the example feature map, obtain the original relevance of the relevance feature corresponding to each example feature map at the feature position, perform a normalization process on the original relevance corresponding to each example feature map to obtain the updated relevance of the relevance feature corresponding to each example feature map at the feature position;

[0239] Generate the example normalization feature corresponding to each example feature map based on the updated relevance of the relevance feature corresponding to each example feature map at each feature position.

[0240] In some embodiments, the at least one dimension includes a spatial dimension, and the performing a normalization process on the relevance feature corresponding to each example feature map from at least one dimension to obtain the normalized relevance feature corresponding to each example feature map for each dimension in the at least one dimension includes:

[0241] For each of the example feature maps, obtain the original relevance at each feature position of the relevance feature corresponding to the example feature map, perform normalization processing on the original relevance at each feature position to obtain the updated relevance at each feature position; based on the updated relevance at each feature position, generate the spatially normalized feature corresponding to the example feature map.

[0242] In some embodiments, the generating the normalized relevance feature corresponding to the example feature map based on the normalized relevance features corresponding to each example feature map for each dimension includes:

[0243] For each of the example feature maps, fuse the normalized relevance features corresponding to each dimension based on element-wise multiplication to generate the normalized relevance feature corresponding to the example feature map.

[0244] In some embodiments, the obtaining the relevant feature map based on the relevance feature corresponding to each example feature map and the corresponding example feature map includes:

[0245] For each of the example feature maps, perform a flipping process on the example feature map to obtain the flipped example feature map; use the flipped example feature map to perform a convolution process on the relevance feature corresponding to the example feature map to obtain the sub-relevant feature map corresponding to the example feature map;

[0246] Based on the sub-relevant feature maps corresponding to each example feature map, generate the relevant feature map.

[0247] In some embodiments, the fusing the relevant feature map and the image feature map to obtain the refined feature map includes:

[0248] Perform a fusion process on the relevant feature map and the image feature map to obtain the refined feature map; the fusion process includes at least one of the following: skip connection, convolution, and normalization layer processing.

[0249] In some embodiments, before converting the refined feature map into a density estimation map, the method further includes:

[0250] Use the refined feature map as the image feature map and return to the step of generating the refined feature map based on the image feature map and the example feature map for iterative processing.

[0251] In some embodiments, the converting the refined feature map into a density estimation map includes:

[0252] Perform iterative convolution processing on the refined feature map;

[0253] Upsample the refined feature map after iterative convolution processing to obtain a density estimation map with the same size as the image to be processed.

[0254] In some embodiments, generating the statistical result for the object to be counted based on the density estimation map includes at least one of the following:

[0255] Sum the density estimation map to obtain the statistical quantity of the object to be counted in the image to be processed;

[0256] Obtain the local peak points in the density estimation map, and screen the local peak points through a non-maximum suppression algorithm to obtain the positions of the objects to be counted; the pixel values of the local peak points are greater than the pixel values of adjacent pixel points.

[0257] Figure 18 The following is a schematic structural diagram of a model training device provided by an embodiment of the present disclosure, as Figure 18 shown, the model training device 1800 includes:

[0258] A second acquisition module 1801, configured to acquire a sample image, the true density map corresponding to the sample image, and at least one example image corresponding to the object to be counted;

[0259] An input module 1802, configured to input the sample image and the at least one example image into a statistical model to be trained to obtain a predicted density map; the statistical model is used to extract the image feature map corresponding to the sample image and the example feature map corresponding to each of the at least one example image; generate a refined feature map based on the image feature map and each example feature map, and the refined feature map retains the features related to each example feature map on the basis of the image feature map; and convert the refined feature map into the predicted density map;

[0260] An adjustment module 1803, configured to adjust the parameters of the statistical model to be trained based on the predicted density map and the true density map to obtain a trained statistical model.

[0261] In some embodiments, the statistical model includes a feature extraction network, a correlation calculation network, a feature refinement network, and a density regression network. The inputting the sample image and the at least one example image into the statistical model to be trained to obtain a predicted density map includes:

[0262] Input the sample image and the at least one example image into the feature extraction network to obtain the image feature map corresponding to the sample image and the example feature map corresponding to each of the at least one example image;

[0263] Input the image feature map and each of the example feature maps into the relevance calculation network to obtain a relevance feature corresponding to each of the example feature maps; the relevance feature is used to characterize the degree of relevance between the image feature map and the example feature map;

[0264] Input the image feature map, each of the example feature maps, and the relevance feature corresponding to each of the example feature maps into the feature refinement network to obtain the refined feature map; the refined feature map retains the features related to each of the example feature maps on the basis of the image feature map;

[0265] Input the refined feature map into the density regression network to obtain the predicted density map.

[0266] In some embodiments, the relevance calculation network includes a first normalization layer and a first convolutional layer; the step of inputting the image feature map and each of the example feature maps into the relevance calculation network to obtain a relevance feature corresponding to each of the example feature maps includes:

[0267] Input the image feature map and each of the example feature maps into the first normalization layer to obtain an intermediate image feature map corresponding to the image feature map and intermediate example feature maps corresponding to each of the example feature maps;

[0268] Use each of the intermediate example feature maps as a convolution kernel and use the intermediate image feature map as the object to be convolved to obtain a relevance feature corresponding to each of the example feature maps.

[0269] In some embodiments, the relevance calculation network further includes a second normalization layer, and the method further includes: inputting the relevance feature corresponding to each of the example feature maps into the second normalization layer to obtain a normalized relevance feature corresponding to each of the example feature maps;

[0270] Wherein, the second normalization layer includes sub-normalization layers corresponding to at least one dimension and a first sub-fusion layer; the step of inputting the relevance feature corresponding to each of the example feature maps into the second normalization layer to normalize the relevance feature corresponding to each of the example feature maps includes: inputting the relevance feature corresponding to each of the example feature maps into the sub-normalization layer corresponding to each dimension to obtain a normalized relevance feature corresponding to each of the example feature maps corresponding to each dimension; inputting the normalized relevance feature corresponding to each of the example feature maps corresponding to each dimension into the first sub-fusion layer to obtain a normalized relevance feature corresponding to each of the example feature maps.

[0271] In some embodiments, the sub-normalization layer corresponding to at least one dimension includes an example sub-normalization layer corresponding to an example dimension. The step of inputting the relevance features corresponding to each example feature map into the sub-normalization layer corresponding to each dimension to obtain the normalized relevance features corresponding to each example feature map in each dimension includes:

[0272] Inputting the relevance features corresponding to each example feature map into the example sub-normalization layer to obtain the example-normalized features corresponding to each example feature map;

[0273] Wherein, the example sub-normalization layer is used to obtain the original relevance of the relevance features corresponding to each example feature map at each feature position in the example feature map, perform normalization processing on the original relevance corresponding to each example feature map to obtain the updated relevance of the relevance features corresponding to each example feature map at each feature position; and generate the example-normalized features corresponding to each example feature map based on the updated relevance of the relevance features corresponding to each example feature map at each feature position.

[0274] In some embodiments, the sub-normalization layer corresponding to at least one dimension includes a spatial sub-normalization layer corresponding to a spatial dimension. The step of inputting the relevance features corresponding to each example feature map into the sub-normalization layer corresponding to each dimension to obtain the normalized relevance features corresponding to each example feature map in each dimension includes:

[0275] Inputting the relevance features corresponding to each example feature map into the spatial sub-normalization layer to obtain the spatial-normalized features corresponding to each example feature map;

[0276] Wherein, the spatial sub-normalization layer is used to obtain the original relevance of the relevance features corresponding to each example feature map at each feature position for each example feature map, perform normalization processing on the original relevance at each feature position to obtain the updated relevance at each feature position; and generate the spatial-normalized features corresponding to the example feature map based on the updated relevance at each feature position.

[0277] In some embodiments, the step of inputting the normalized relevance features corresponding to each example feature map in each dimension into the first sub-fusion layer to obtain the normalized relevance features corresponding to each example feature map includes:

[0278] For each of the example feature maps, input the correlation features after normalization corresponding to each dimension into the first sub-fusion layer to obtain the correlation features after normalization corresponding to each example feature map; the first sub-fusion layer is used to fuse the correlation features after normalization corresponding to each dimension based on element-wise multiplication to generate the correlation features after normalization corresponding to the example feature map.

[0279] In some embodiments, the feature refinement network includes a correlation feature map generation layer and an iterative fusion layer; the step of inputting the image feature map, each of the example feature maps, and the correlation features corresponding to each example feature map into the feature refinement network to obtain the refined feature map includes:

[0280] Input each of the example feature maps and the correlation features corresponding to each example feature map into the correlation feature map generation layer to obtain a correlation feature map;

[0281] Input the correlation feature map and the image feature map into the iterative fusion layer to obtain the refined feature map.

[0282] In some embodiments, the correlation feature map generation layer includes a feature flipping layer, a second convolutional layer, and a second sub-fusion layer; the step of inputting each of the example feature maps and the correlation features corresponding to each example feature map into the correlation feature map generation layer to obtain a correlation feature map includes:

[0283] For each of the example feature maps, input the example feature map into the feature flipping layer to obtain a flipped example feature map;

[0284] Input the flipped example feature map and the correlation features corresponding to the example feature map into the second convolutional layer to obtain a sub-correlation feature map corresponding to the example feature map;

[0285] Input the sub-correlation feature maps corresponding to each of the example feature maps into the second sub-fusion layer to generate the correlation feature map.

[0286] In some embodiments, the iterative fusion layer is used to perform a fusion process on the correlation feature map and the image feature map to obtain the refined feature map; the fusion process includes at least one of the following: skip connection, convolution, and normalization layer processing.

[0287] In some embodiments, before inputting the refined feature map into the density regression network to obtain the predicted density map, the method further includes:

[0288] Use the refined feature map as the image feature map, and return the step of inputting the image feature map and each of the example feature maps into the relevance calculation network to obtain the relevance feature corresponding to each of the example feature maps for iterative processing.

[0289] In some embodiments, the density regression network includes an iterative convolutional layer and a bilinear upsampling layer. The step of inputting the refined feature map into the density regression network to obtain the predicted density map includes:

[0290] Input the refined feature map into the iterative convolutional layer to obtain a refined feature map after iterative convolutional processing;

[0291] Input the refined feature map after iterative convolutional processing into the bilinear upsampling layer to obtain a predicted density map with the same size as the image to be processed.

[0292] In some embodiments, the step of adjusting the parameters of the statistical model to be trained based on the predicted density map and the true density map to obtain a trained statistical model includes:

[0293] Input the predicted density map and the true density map into a preset loss function to determine a loss value;

[0294] Use the loss value to adjust the parameters of the statistical model to be trained until a preset training stop condition is reached, and output the trained statistical model. The training stop condition includes at least one of the following: the loss value converges, the preset number of training times is reached.

[0295] The description of the above device embodiments is similar to the description of the above method embodiments, and has beneficial effects similar to those of the method embodiments. For technical details not disclosed in the device embodiments of the present disclosure, please refer to the description of the method embodiments of the present disclosure for understanding.

[0296] It should be noted that in the embodiments of the present disclosure, if the above object statistics method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present disclosure essentially or the part that contributes to the related technology can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a device to execute all or part of the methods of the various embodiments of the present disclosure. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), magnetic disks, or optical discs that can store program codes. In this way, the embodiments of the present disclosure are not limited to any combination of hardware and software for any purpose.

[0297] Figure 19 The following is a schematic diagram of the hardware entity of an object statistics device provided by an embodiment of the present disclosure. As Figure 19 shown, the hardware entity of the object statistics device 1900 includes: a processor 1901 and a memory 1902. Among them, the memory 1902 stores a computer program that can run on the processor 1901, and when the processor 1901 executes the program, it implements the steps in the method of any of the above embodiments.

[0298] The memory 1902 stores a computer program that can run on the processor. The memory 1902 is configured to store instructions and applications executable by the processor 1901, and can also cache data to be processed or already processed by the processor 1901 and each module in the object statistics device 1900 (for example, image data, audio data, voice communication data, and video communication data), and can be implemented by flash memory (FLASH) or random access memory (Random Access Memory, RAM).

[0299] When the processor 1901 executes the program, it implements the steps of the object statistics method of any of the above. The processor 1901 generally controls the overall operation of the object statistics device 1900.

[0300] An embodiment of the present disclosure provides a computer storage medium. The computer storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the object statistics method of any of the above embodiments.

[0301] It should be pointed out here that: the descriptions of the above storage medium and device embodiments are similar to the descriptions of the above method embodiments, and have beneficial effects similar to those of the method embodiments. For the technical details not disclosed in the storage medium and device embodiments of the present disclosure, please refer to the descriptions of the method embodiments of the present disclosure for understanding.

[0302] The above-mentioned processor may be at least one of an Application Specific Integrated Circuit (ASIC), a Digital Signal Processor (DSP), a Digital Signal Processing Device (DSPD), a Programmable Logic Device (PLD), a Field Programmable Gate Array (FPGA), a Central Processing Unit (CPU), a controller, a microcontroller, and a microprocessor. It can be understood that the electronic device implementing the functions of the above-mentioned processor may also be other devices, and the embodiments of the present disclosure do not make specific limitations.

[0303] The above-mentioned computer storage medium / memory may be a Read Only Memory (ROM), a Programmable Read-Only Memory (PROM), an Erasable Programmable Read-Only Memory (EPROM), an Electrically Erasable Programmable Read-Only Memory (EEPROM), a Ferromagnetic Random Access Memory (FRAM), a Flash Memory, a magnetic surface memory, an optical disc, or a Compact Disc Read-Only Memory (CD-ROM), etc.; it may also be various terminals including one or any combination of the above-mentioned memories, such as a mobile phone, a computer, a tablet device, a personal digital assistant, etc.

[0304] It should be understood that the "one embodiment" or "an embodiment" or "an embodiment of the present disclosure" or "the foregoing embodiment" or "some embodiments" mentioned throughout the specification means that the target features, structures or characteristics related to the embodiment are included in at least one embodiment of the present disclosure. Therefore, the "in one embodiment" or "in an embodiment" or "an embodiment of the present disclosure" or "the foregoing embodiment" or "some embodiments" that appear throughout the specification do not necessarily refer to the same embodiment. In addition, these target features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present disclosure, the magnitudes of the serial numbers of the above processes do not mean the sequence of execution, and the execution sequence of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present disclosure. The serial numbers of the embodiments of the present disclosure above are only for description and do not represent the advantages or disadvantages of the embodiments.

[0305] Without special instructions, when the object statistical device executes any step in the embodiments of the present disclosure, it can be the processor of the object statistical device that executes this step. Unless otherwise specified, the embodiments of the present disclosure do not limit the sequence of execution of the following steps by the object statistical device. In addition, the methods used to process data in different embodiments can be the same method or different methods. It should also be noted that any step in the embodiments of the present disclosure can be independently executed by the object statistical device, that is, when the object statistical device executes any step in the above embodiments, it can be independent of the execution of other steps.

[0306] In several embodiments provided by the present disclosure, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the couplings between the components shown or discussed, or direct couplings, or communication connections can be through some interfaces, and the indirect couplings or communication connections of devices or units can be electrical, mechanical or other forms.

[0307] The units described as separate components above may or may not be physically separated, and the components shown as units may or may not be physical units; they can be located in one place or distributed to multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0308] In addition, in each embodiment of the present disclosure, each functional unit can be entirely integrated into one processing unit, or each unit can be separately regarded as one unit, or two or more units can be integrated into one unit; the above integrated unit can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.

[0309] The methods disclosed in several method embodiments provided by the present disclosure can be arbitrarily combined without conflict to obtain new method embodiments.

[0310] The features disclosed in several product embodiments provided by the present disclosure can be arbitrarily combined without conflict to obtain new product embodiments.

[0311] The features disclosed in several method or device embodiments provided by the present disclosure can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0312] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: removable storage devices, read-only memory (ROM), magnetic disks, or optical disks and other various media that can store program codes.

[0313] Alternatively, if the above integrated unit of the present disclosure is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of the present disclosure, in essence, or the part that contributes to the related technology can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, an object statistical device, or a network device, etc.) to execute all or part of the methods described in various embodiments of the present disclosure. And the foregoing storage medium includes: removable storage devices, ROM, magnetic disks, or optical disks and other various media that can store program codes.

[0314] In the embodiments of the present disclosure, the descriptions of the same steps and the same content in different embodiments can be referred to each other. In the embodiments of the present disclosure, the term "and" does not affect the sequence of steps.

[0315] As described above, it is only the implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present disclosure can easily think of changes or substitutions, which should all be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure shall be subject to the protection scope of the claims.

Claims

1. An object counting method, characterized in that, The method includes: Obtaining at least one example image corresponding to the image to be processed and the object to be counted; Extracting the image feature map corresponding to the image to be processed and the example feature map corresponding to each of the at least one example image; Generating a refined feature map based on the image feature map and each of the example feature maps; the refined feature map retains the features related to each of the example feature maps on the basis of the image feature map; Taking the refined feature map as the image feature map and returning the step of generating the refined feature map based on the image feature map and the example feature maps for iterative processing; Converting the refined feature map after iterative processing into a density estimation map, and generating a statistical result for the object to be counted based on the density estimation map.

2. The method according to claim 1, wherein The generating a refined feature map based on the image feature map and each of the example feature maps includes: Generating a correlation feature corresponding to each of the example feature maps based on the image feature map and each of the example feature maps; the correlation feature is used to characterize the correlation degree between the image feature map and the example feature map; Obtaining a correlation feature map based on the correlation feature corresponding to each of the example feature maps and the corresponding example feature map; Fusing the correlation feature map and the image feature map to obtain a refined feature map.

3. The method according to claim 2, characterized in that, The generating a correlation feature corresponding to each of the example feature maps based on the image feature map and each of the example feature maps includes: Performing convolution processing on the image feature map using each of the example feature maps to obtain the correlation feature corresponding to each of the example feature maps.

4. The method according to claim 3, characterized in that The performing convolution processing on the image feature map using each of the example feature maps to obtain the correlation feature corresponding to each of the example feature maps includes: Processing the image feature map through a shared normalization layer to obtain an intermediate image feature map; For each example feature map, processing the example feature map through the shared normalization layer to obtain an intermediate example feature map; Taking the intermediate example feature map as a convolution kernel and taking the intermediate image feature map as the object to be convolved to obtain the correlation feature corresponding to the example feature map.

5. The method according to any one of claims 2 to 4, characterized in that Before obtaining a correlation feature map based on the correlation feature corresponding to each of the example feature maps and the corresponding example feature map, the method further includes: performing normalization processing on the correlation feature corresponding to each of the example feature maps; The performing normalization processing on the correlation feature corresponding to each of the example feature maps includes: performing normalization processing on the correlation feature corresponding to each of the example feature maps from at least one dimension to obtain a correlation feature to be fused for each of the example feature maps in each dimension; generating a normalized correlation feature corresponding to each of the example feature maps based on the correlation feature to be fused for each of the example feature maps in each dimension.

6. The method according to claim 5, wherein The at least one dimension includes an example dimension, and the performing normalization processing on the correlation feature corresponding to each of the example feature maps from at least one dimension to obtain a correlation feature to be fused for each of the example feature maps in each dimension includes: For each feature position in the relevance feature corresponding to the example feature map, obtain the original relevance of the relevance feature corresponding to each said example feature map at the feature position; perform normalization processing on the original relevance corresponding to each said example feature map to obtain the updated relevance of the relevance feature corresponding to each said example feature map at the feature position; Based on the updated relevance of the relevance feature corresponding to each said example feature map at each said feature position, generate the example normalization feature corresponding to each said example feature map.

7. The method according to claim 5, wherein The at least one dimension includes a spatial dimension, and the normalization processing of the relevance feature corresponding to each said example feature map from the at least one dimension to obtain the relevance feature to be fused corresponding to each said example feature map in each said dimension includes: For each said example feature map, obtain the original relevance of the relevance feature corresponding to the example feature map at each feature position, perform normalization processing on the original relevance at each said feature position to obtain the updated relevance at each said feature position; based on the updated relevance at each said feature position, generate the spatial normalization feature corresponding to the example feature map.

8. The method according to claim 5, wherein The generating the relevance feature after normalization processing corresponding to each said example feature map based on the relevance feature to be fused corresponding to each said example feature map in each said dimension includes: For each said example feature map, fuse the relevance features to be fused corresponding to the example feature map in each said dimension based on element-wise multiplication to generate the relevance feature after normalization processing corresponding to the example feature map.

9. The method according to any one of claims 2 to 4, 6 to 8, characterized in that The obtaining the relevant feature map based on the relevance feature corresponding to each said example feature map and the corresponding example feature map includes: For each said example feature map, perform a flipping process on the example feature map to obtain the flipped example feature map; use the flipped example feature map to perform a convolution process on the relevance feature corresponding to the example feature map to obtain the sub-relevant feature map corresponding to the example feature map; Based on the sub-relevant feature map corresponding to each said example feature map, generate the relevant feature map.

10. The method according to any one of claims 2 to 4 and 6 to 8, characterized in that The fusing the relevant feature map and the image feature map to obtain the refined feature map includes: Perform a fusion process on the relevant feature map and the image feature map to obtain the refined feature map; the fusion process includes at least one of the following: skip connection, convolution, and normalization layer processing.

11. The method according to any one of claims 1 to 4, 6 to 8, characterized in that, The converting the refined feature map after iterative processing into a density estimation map includes: Perform an iterative convolution process on the refined feature map; Perform upsampling on the refined feature map after iterative convolution processing to obtain a density estimation map with the same size as the image to be processed.

12. The method according to any one of claims 1 to 4, 6 to 8, characterized in that, The generating the statistical result for the object to be counted based on the density estimation map includes at least one of the following: Sum the density estimation map to obtain the statistical quantity of the object to be counted in the image to be processed; Obtain local peak points in the density estimation map, and screen the local peak points through a non-maximum suppression algorithm to obtain the position of the object to be counted; the pixel value of the local peak point is greater than the pixel values of adjacent pixel points.

13. A model training method, characterized in that, The method includes: Obtain a sample image, the true density map corresponding to the sample image, and at least one example image corresponding to the object to be counted; Input the sample image and the at least one example image into a statistical model to be trained to obtain a predicted density map; the statistical model is used to extract the image feature map corresponding to the sample image and the example feature map corresponding to each example image in the at least one example image; generate a refined feature map based on the image feature map and each example feature map, and the refined feature map retains features related to each example feature map on the basis of the image feature map; take the refined feature map as the image feature map, and return the step of inputting the image feature map and each example feature map into a relevance calculation network to obtain the relevance feature corresponding to each example feature map for iterative processing, and convert the refined feature map after iterative processing into the predicted density map; Based on the predicted density map and the true density map, adjust the parameters of the statistical model to be trained to obtain a trained statistical model.

14. The method according to claim 13, wherein The statistical model includes a feature extraction network, a relevance calculation network, a feature refinement network, and a density regression network. The step of inputting the sample image and the at least one example image into the statistical model to be trained to obtain a predicted density map includes: Input the sample image and the at least one example image into the feature extraction network to obtain the image feature map corresponding to the sample image and the example feature map corresponding to each example image in the at least one example image; Input the image feature map and each example feature map into the relevance calculation network to obtain the relevance feature corresponding to each example feature map; the relevance feature is used to characterize the relevance degree between the image feature map and the example feature map; Input the image feature map, each example feature map, and the relevance feature corresponding to each example feature map into the feature refinement network to obtain the refined feature map; the refined feature map retains features related to each example feature map on the basis of the image feature map; Input the refined feature map into the density regression network to obtain the predicted density map.

15. The method according to claim 14, wherein The relevance calculation network includes a first normalization layer and a first convolutional layer. The step of inputting the image feature map and each example feature map into the relevance calculation network to obtain the relevance feature corresponding to each example feature map includes: Input the image feature map and each example feature map into the first normalization layer to obtain the intermediate image feature map corresponding to the image feature map and the intermediate example feature map corresponding to each example feature map; Using each of the intermediate example feature maps as a convolution kernel and the intermediate image feature map as the object to be convolved, a relevance feature corresponding to each of the example feature maps is obtained.

16. The method according to claim 14 or 15, characterized in that, The relevance calculation network further includes a second normalization layer, and the method further includes: inputting the relevance feature corresponding to each of the example feature maps into the second normalization layer to obtain a normalized relevance feature corresponding to each of the example feature maps; wherein, the second normalization layer includes sub-normalization layers corresponding to at least one dimension and a first sub-fusion layer; the step of inputting the relevance feature corresponding to each of the example feature maps into the second normalization layer to perform normalization processing on the relevance feature corresponding to each of the example feature maps includes: inputting the relevance feature corresponding to each of the example feature maps into the sub-normalization layer corresponding to each dimension to obtain a to-be-fused relevance feature of each of the example feature maps in each dimension; inputting the obtained to-be-fused relevance feature of each of the example feature maps in each dimension into the first sub-fusion layer to obtain a normalized relevance feature corresponding to each of the example feature maps.

17. The method according to claim 16, wherein The sub-normalization layer corresponding to the at least one dimension includes an example sub-normalization layer corresponding to the example dimension, and the step of inputting the relevance feature corresponding to each of the example feature maps into the sub-normalization layer corresponding to each dimension to obtain a to-be-fused relevance feature of each of the example feature maps in each dimension includes: inputting the relevance feature corresponding to each of the example feature maps into the example sub-normalization layer to obtain an example-normalized feature corresponding to each of the example feature maps; wherein, the example sub-normalization layer is used to obtain the original relevance of the relevance feature corresponding to the example feature map at each feature position for each example feature map; performing normalization processing on the original relevance corresponding to each example feature map to obtain the updated relevance of the relevance feature corresponding to each example feature map at the feature position; generating the example-normalized feature corresponding to each of the example feature maps based on the updated relevance of the relevance feature corresponding to each example feature map at each feature position.

18. The method according to claim 16, characterized in that, The sub-normalization layer corresponding to the at least one dimension includes a spatial sub-normalization layer corresponding to the spatial dimension, and the step of inputting the relevance feature corresponding to each of the example feature maps into the sub-normalization layer corresponding to each dimension to obtain a to-be-fused relevance feature of each of the example feature maps in each dimension includes: inputting the relevance feature corresponding to each of the example feature maps into the spatial sub-normalization layer to obtain a spatial-normalized feature corresponding to each of the example feature maps; Among them, the spatial sub-normalization layer is used to obtain the original relevance of the relevance features corresponding to each example feature map at each feature position for each example feature map, perform normalization processing on the original relevance at each feature position to obtain the updated relevance at each feature position; based on the updated relevance at each feature position, generate the spatial normalization feature corresponding to the example feature map.

19. The method according to claim 16, characterized in that, Inputting the relevance features to be fused of each example feature map in each dimension into the first sub-fusion layer to obtain the normalized relevance features corresponding to each example feature map includes: For each example feature map, inputting the relevance features to be fused of each example feature map in each dimension into the first sub-fusion layer to obtain the normalized relevance features corresponding to each example feature map; the first sub-fusion layer is used to fuse the relevance features to be fused of each example feature map in each dimension based on element-wise multiplication to generate the normalized relevance features corresponding to the example feature map.

20. The method according to any one of claims 14 to 15, 17 to 19, characterized in that The feature refinement network includes a correlation feature map generation layer and an iterative fusion layer; inputting the image feature map, each example feature map, and the relevance features corresponding to each example feature map into the feature refinement network to obtain the refined feature map includes: Inputting each example feature map and the relevance features corresponding to each example feature map into the correlation feature map generation layer to obtain a correlation feature map; Inputting the correlation feature map and the image feature map into the iterative fusion layer to obtain the refined feature map.

21. The method according to claim 20, wherein The correlation feature map generation layer includes a feature flipping layer, a second convolutional layer, and a second sub-fusion layer; inputting each example feature map and the relevance features corresponding to each example feature map into the correlation feature map generation layer to obtain a correlation feature map includes: For each example feature map, inputting the example feature map into the feature flipping layer to obtain a flipped example feature map; Inputting the flipped example feature map and the relevance features corresponding to the example feature map into the second convolutional layer to obtain a sub-correlation feature map corresponding to the example feature map; Inputting the sub-correlation feature maps corresponding to each example feature map into the second sub-fusion layer to generate the correlation feature map.

22. The method according to claim 20, characterized in that, The iterative fusion layer is used to perform fusion processing on the correlation feature map and the image feature map to obtain the refined feature map; the fusion processing includes at least one of the following: skip connection, convolution, and normalization layer processing.

23. The method according to any one of claims 14 to 15, 17 to 19, characterized in that, The density regression network includes an iterative convolutional layer and a bilinear upsampling layer. Inputting the iteratively processed refined feature map into the density regression network to obtain the predicted density map includes: Inputting the refined feature map into the iterative convolutional layer to obtain an iteratively convolutionally processed refined feature map; Inputting the iteratively convolutionally processed refined feature map into the bilinear upsampling layer to obtain a predicted density map with the same size as the image to be processed.

24. The method according to any one of claims 14 to 15, 17 to 19, characterized in that, Adjusting the parameters of the statistical model to be trained based on the predicted density map and the ground truth density map to obtain the trained statistical model, including: Inputting the predicted density map and the ground truth density map into a preset loss function to determine a loss value; Adjusting the parameters of the statistical model to be trained using the loss value until a preset training stop condition is reached, and outputting the trained statistical model; the training stop condition includes at least one of the following: the loss value converges, a preset number of training times is reached.

25. An object counting device, characterized in that, Including: A first acquisition module, configured to acquire at least one example image corresponding to the image to be processed and the object to be counted; An extraction module, configured to extract an image feature map corresponding to the image to be processed and an example feature map corresponding to each of the at least one example image; A generation module, configured to generate a refined feature map based on the image feature map and each example feature map; the refined feature map retains features related to each example feature map on the basis of the image feature map; Taking the refined feature map as the image feature map and returning to the step of generating the refined feature map based on the image feature map and the example feature maps for iterative processing; A conversion module, configured to convert the refined feature map after iterative processing into a density estimation map, and generate a statistical result for the object to be counted based on the density estimation map.

26. A model training device, characterized in that, Including: A second acquisition module, configured to acquire a sample image, the ground truth density map corresponding to the sample image, and at least one example image corresponding to the object to be counted; An input module, configured to input the sample image and the at least one example image into the statistical model to be trained to obtain a predicted density map; the statistical model is configured to extract an image feature map corresponding to the sample image and an example feature map corresponding to each of the at least one example image; generate a refined feature map based on the image feature map and each example feature map, and the refined feature map retains features related to each example feature map on the basis of the image feature map; Taking the refined feature map as the image feature map, and returning to the step of inputting the image feature map and each example feature map into a relevance calculation network to obtain a relevance feature corresponding to each example feature for iterative processing, and converting the refined feature map after iterative processing into the predicted density map; An adjustment module, configured to adjust the parameters of the statistical model to be trained based on the predicted density map and the ground truth density map to obtain the trained statistical model.

27. An object statistical device, characterized in that, Including: A memory and a processor, The memory stores a computer program that can run on the processor, When the processor executes the computer program, the steps in the method according to any one of claims 1 to 12, 13 to 24 are implemented.

28. A computer storage medium, characterized in that, The computer storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the method according to any one of claims 1 to 12, 13 to 24.

Citation Information

Patent Citations

  • Image processing method and device, model training method and device, equipment and storage medium

    CN114612414A