Method for detecting the number of people based on a crowd image template
By obtaining the image of the population to be detected and determining its similarity to the candidate population template image, and using the number of people in the target population template image to determine the number of people in the image of the population to be detected, the problem of low number detection efficiency in the prior art is solved, and fast and accurate number detection is achieved.
Patent Information
- Application Number
- CN202210699257.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-20
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-06-20
AI Technical Summary
In the prior art, the efficiency of performing number detection of people in the image is low.
By obtaining the image of the population to be detected, the candidate population template image with the highest similarity is determined, and it is used as the target population template image, and the number of people in the image of the population to be detected is determined based on the number of people corresponding to the target population template image.
Improves the efficiency of number detection and enables the number of people in the image to be quickly and accurately determined.
Smart Images

Figure CN115205779B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of human counting, and specifically relates to a human counting method based on a crowd image template. Background Art
[0002] Currently, using computer vision related technologies to count the number of people in an image is a relatively popular research direction. The research goal is to generate a corresponding crowd density map through an algorithm and detect the number of people in the crowd density map given an image or a video in a crowd scene. In related technologies, the efficiency of counting the number of people in an image is low. Summary of the Invention
[0003] An embodiment of this application provides a human counting method based on a crowd image template, which can improve the efficiency of counting the number of people in an image.
[0004] In a first aspect, an embodiment of this application provides a human counting method based on a crowd image template, including:
[0005] Obtain a to-be-detected crowd image that needs to be counted;
[0006] Determine a candidate crowd template image with the highest similarity to the to-be-detected crowd image from multiple candidate crowd template images;
[0007] Use the candidate crowd template image with the highest similarity as the target crowd template image;
[0008] Determine the number of people in the to-be-detected crowd image according to the number of people corresponding to the target crowd template image.
[0009] In a second aspect, an embodiment of this application further provides a human counting device based on a crowd image template, including:
[0010] An obtaining module, configured to obtain a to-be-detected crowd image that needs to be counted;
[0011] A first determination module, configured to determine a candidate crowd template image with the highest similarity to the to-be-detected crowd image from multiple candidate crowd template images;
[0012] A second determination module, configured to use the candidate crowd template image with the highest similarity as the target crowd template image;
[0013] A third determination module, configured to determine the number of people in the to-be-detected crowd image according to the number of people corresponding to the target crowd template image.
[0014] Thirdly, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program runs on a processor, the computer is enabled to execute the method for detecting the number of people based on a crowd image template provided in any embodiment of the present application.
[0015] Fourthly, an embodiment of the present application further provides an electronic device, including a processor and a memory. The memory has a computer program, and the processor is used to execute the method for detecting the number of people based on a crowd image template provided in any embodiment of the present application by calling the computer program.
[0016] The technical solution provided by the embodiment of the present application determines, from multiple candidate crowd template images, a candidate crowd template image with the highest similarity to the to-be-detected crowd image by obtaining the to-be-detected crowd image that needs to be counted for the number of people, uses the candidate crowd template image with the highest similarity as the target crowd template image, and determines the number of people in the to-be-detected crowd image according to the number of people corresponding to the target crowd template image. The present application directly determines the number of people in the to-be-detected crowd image by obtaining the target crowd template image with the highest similarity to the to-be-detected crowd image and according to the number of people corresponding to the target crowd template image, which can improve the efficiency of number detection. Description of the Drawings
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0018] Figure 1 It is a schematic diagram of the scenario of the number detection system provided by the embodiment of the present application.
[0019] Figure 2 It is the first flowchart of the method for detecting the number of people based on a crowd image template provided by the embodiment of the present application.
[0020] Figure 3 It is the second flowchart of the method for detecting the number of people based on a crowd image template provided by the embodiment of the present application.
[0021] Figure 4 It is the structural diagram of the number detection model provided by the embodiment of the present application.
[0022] Figure 5 It is the structural diagram of the attention scale network of the method for detecting the number of people based on a crowd image template provided by the embodiment of the present application.
[0023] Figure 6Schematic structural diagram of the number detection device based on the population image template provided by the embodiment of the present application.
[0024] Figure 7 The first schematic structural diagram of the electronic device provided by the embodiment of the present application.
[0025] Figure 8 The second schematic structural diagram of the electronic device provided by the embodiment of the present application. Detailed implementation manners
[0026] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present application.
[0027] As used herein, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present application. The phrase does not necessarily refer to the same embodiment at various places in the specification, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0028] Currently, detecting the number of people in an image through computer vision-related technologies is a relatively popular research direction. Its research goal is to generate a corresponding crowd density map and detect the number of people in the crowd density map given an image or a video in a crowd scene. In the related art, the efficiency of detecting the number of people in an image is relatively low.
[0029] In order to improve the efficiency of detecting the number of people in an image, the embodiment of the present application provides a number detection method based on a population image template. Among them, the execution subject of the number detection method based on the population image template can be the number detection device based on the population image template provided by the embodiment of the present application, or an electronic device integrated with the number detection device based on the population image template, where the number detection device based on the population image template can be implemented in a hardware or software manner.
[0030] Please refer to Figure 1 , the present application also provides a number detection system, as Figure 1As shown in the figure, the number detection system includes an electronic device 10, and the number detection device based on the crowd image template provided by this application is integrated in the electronic device 10. For example, when the electronic device 10 is also equipped with a camera, the camera can be directly used to capture the crowd to be counted, so as to obtain the to-be-detected crowd image for crowd counting. Then, the target crowd template image matching the to-be-detected crowd image is determined, and the number of people in the to-be-detected crowd image is determined according to the number of people corresponding to the target crowd template image. By directly obtaining the target crowd template image matching the to-be-detected crowd image and determining the number of people in the to-be-detected crowd image according to the number of people corresponding to the target crowd template image, the efficiency of detecting the number of people in the image can be improved in this application.
[0031] The electronic device 10 can be any device with a processor and thus processing capabilities, such as mobile electronic devices with a processor like smartphones, tablets, handheld computers, laptop computers, etc., or fixed electronic devices with a processor like desktop computers, TVs, servers, etc.
[0032] In addition, as Figure 1 shown, the number detection system may further include a memory 20 for storing data. For example, the electronic device 10 can store data such as the to-be-detected crowd image, the target crowd template image, and the number of people in the to-be-detected crowd image into the memory 20.
[0033] It should be noted that Figure 1 the scene schematic diagram of the number detection system shown is only an example. The number detection system and the scene described in the embodiments of this application are for more clearly explaining the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those of ordinary skill in the art know that with the evolution of the number detection system and the emergence of new business scenarios, the technical solutions provided by the embodiments of this application are equally applicable to similar technical problems.
[0034] The following will be described in detail respectively. It should be noted that the serial numbers of the following embodiments do not limit the preferred order of the embodiments.
[0035] Please refer to Figure 2 , Figure 2 which is the first flow schematic diagram of the number detection method based on the crowd image template provided by the embodiments of this application. The specific process of the number detection method based on the crowd image template provided by the embodiments of this application can be as follows:
[0036] 101. Obtain the to-be-detected crowd image for crowd counting.
[0037] Among them, the image of the population to be detected refers to the image for which the number of people needs to be detected. This image of the population to be detected can be obtained by shooting with an image acquisition device such as a camera, or by an image pre-shot and stored locally in an electronic device, or by obtaining the image resources stored on the server side.
[0038] For example, when obtaining the image of the population to be detected for which the number of people needs to be counted, the place where the number of people needs to be calculated can be photographed to obtain the image of the population to be detected. Among them, the place where the number of people needs to be calculated can be places such as shopping malls, banks, stations, scenic spots during holidays, etc.
[0039] For example, in a bank, in order to avoid potential safety hazards caused by too many people entering the bank, it is necessary to count the number of people in the bank, so as to facilitate the bank staff to manage the personnel flow in the bank. Then, camera devices can be set up in the bank to photograph each area of the bank, and the photographed photos can be used as the images of the population to be detected.
[0040] 102. Determine the candidate population template image with the highest similarity to the image of the population to be detected from multiple candidate population template images.
[0041] Among them, the candidate population template image refers to the population template image that can be selected. This population template image includes population images with different population distribution situations, and this population template image can be obtained through various channels. For example, it can be obtained by shooting with an image acquisition device such as a camera, or by obtaining the image resources stored on the server side.
[0042] In this embodiment, after obtaining the image of the population to be detected, determine the candidate population template image with the highest similarity to the image of the population to be detected from multiple candidate population template images.
[0043] It should be noted that the image content in the target population template image has a relatively high similarity to the image content in the image of the population to be detected.
[0044] Exemplarily, the candidate population template image with the highest similarity to the image of the population to be detected can be determined from multiple candidate population template images through an image similarity algorithm as the target population template image.
[0045] 103. Use the candidate population template image with the highest similarity as the target population template image.
[0046] In this embodiment, after determining the candidate population template image with the highest similarity to the image of the population to be detected from multiple candidate population template images, use the candidate population template image with the highest similarity as the target population template image.
[0047] 104. Determine the number of people in the image of the population to be detected according to the number of people corresponding to the template image of the target population.
[0048] Exemplarily, in this embodiment, the number of people corresponding to the template image of the target population can be directly used as the number of people in the image of the population to be detected.
[0049] In this embodiment, the number of people in the template image of the population can be detected in advance, and the corresponding relationship between the template image of the population and the number of people corresponding to it can be established. When obtaining the number of people corresponding to the template image of the target population, the above-mentioned corresponding relationship established in advance can be used to quickly obtain the number of people corresponding to the template image of the target population.
[0050] Among them, when detecting the number of people in the template image of the population, the density map corresponding to the template image of the population can be obtained, and the number of people in the template image of the population can be determined according to the density map. In addition, the faces in the template image of the population can also be recognized by face recognition, and the number of people in the template image of the population can be determined according to the number of recognized faces, etc. Among them, there are various ways to obtain the number of people in the template image of the population, which are not specifically limited here.
[0051] In specific implementation, this application is not limited by the execution order of the described steps. Without conflict, some steps can also be performed in other orders or simultaneously.
[0052] As can be seen from the above, the method for detecting the number of people based on the template image of the population provided by the embodiment of this application obtains the image of the population to be detected that needs to be counted, determines the candidate template image of the population with the highest similarity to the image of the population to be detected from multiple candidate template images of the population, uses the candidate template image of the population with the highest similarity as the template image of the target population, and determines the number of people in the image of the population to be detected according to the number of people corresponding to the template image of the target population. This application directly determines the number of people in the image of the population to be detected by obtaining the template image of the target population with the highest similarity to the image of the population to be detected and according to the number of people corresponding to the template image of the target population, which can improve the efficiency of detecting the number of people.
[0053] According to the method described in the previous embodiment, the following will give a further detailed example.
[0054] Please refer to Figure 3 , Figure 3 which is the second process schematic diagram of the method for detecting the number of people based on the template image of the population provided by the embodiment of this application. The method includes:
[0055] 201. Obtain the image of the population to be detected that needs to be counted.
[0056] Among them, the image of the population to be detected refers to the image for which the number of people needs to be detected. This image of the population to be detected can be obtained by shooting with an image acquisition device such as a camera, or by an image pre-shot and stored locally in an electronic device, or by obtaining the image resources stored on the server side.
[0057] For example, in a bank, in order to avoid potential safety hazards caused by an excessive number of people entering the bank, it is necessary to count the number of people in the bank, so as to facilitate the bank staff to manage the personnel flow in the bank. Then, imaging devices can be set up in the bank to shoot each area of the bank, and the photos obtained by shooting can be used as the images of the population to be detected.
[0058] 202. Construct virtual population images with different population distribution situations.
[0059] For example, use the objects in the GTA5 virtual scene to construct a population scene with different population distribution situations, and then capture stable images from the constructed scene through a data collector, so as to obtain virtual population images with different population distribution situations, and then use the virtual population images as population template images.
[0060] In this embodiment, the rendering data corresponding to each constructed virtual population image can be obtained, the number of people in the population template image can be determined according to the rendering data, the corresponding relationship between the population template image and the number of people can be established according to the population template image and the number of people in the population template image, and the number of people corresponding to the target population template image can be determined according to the corresponding relationship.
[0061] Since the number of people in the image of the population to be detected in this application is determined according to the number of people in the target population template image, determining the number of people in the population template image according to the rendering data can ensure the accuracy of the number of people in the obtained population template image, thereby improving the accuracy of the method for detecting the number of people based on the population image template provided in this application.
[0062] 203. Use the virtual population image as the population template image.
[0063] For example, after constructing virtual population images with different population distribution situations, these constructed virtual population images can be used as population template images.
[0064] 204. Determine the candidate population template image with the highest similarity to the image of the population to be detected from multiple candidate population template images.
[0065] Among them, the candidate population template image refers to the selectable population template image. The population template image includes population images with different population distribution situations, and the population template image can be obtained through various channels. For example, it can be obtained by shooting with an image acquisition device such as a camera, or by obtaining the image resources stored on the server side.
[0066] In one implementation, when determining the candidate population template image with the highest similarity to the to-be-detected population image from multiple candidate population template images, the to-be-detected population image can be input into the population image similarity model to obtain the similarity between the to-be-detected population image and each candidate population template image, and the candidate population template image with the highest similarity to the to-be-detected population image is determined according to the similarity.
[0067] Among them, the population image similarity model is configured to obtain the similarity between the input to-be-detected population image and multiple candidate population template images. The population image similarity model can extract the image features in the population image and perform comparison of the image features, so as to obtain the similarity between the images.
[0068] Specifically, the population image similarity model can adopt an image similarity algorithm, and no specific limitation is imposed on the image similarity algorithm here.
[0069] 205. Use the population template image with the highest similarity as the target population template image.
[0070] It can be understood that the higher the similarity between a certain candidate population template image and the to-be-detected population image, the higher the similarity of the image content between the candidate population template image and the to-be-detected population image. In this embodiment, using the population template image with the highest similarity as the target population template image can ensure the accuracy of the number of people in the to-be-detected population image obtained subsequently.
[0071] 206. Determine the number of people in the to-be-detected population image according to the number of people corresponding to the target population template image.
[0072] Exemplarily, in this embodiment, the number of people corresponding to the target population template image can be directly used as the number of people in the to-be-detected population image.
[0073] In addition, in one implementation, when determining the number of people in the to-be-detected population image according to the number of people corresponding to the target population template image, the to-be-detected population image can be input into the number detection model to obtain the reference number of people corresponding to the to-be-detected population image, and the number of people in the to-be-detected population image is determined according to the number of people corresponding to the target population template image and the reference number of people.
[0074] Among them, the number detection model is configured to perform number detection processing on the to-be-detected population image and obtain the reference number of people in the to-be-detected population image.
[0075] Specifically, the average value of the number of people corresponding to the target population template image and the above reference number can be obtained, and the average value is used as the number of people in the image of the population to be detected. That is to say, by using the number of people corresponding to the target population template image as a reference value, and also using the reference number obtained by inputting the image of the population to be detected into the number detection model as another reference value, the number of people in the image of the population to be detected is determined based on these two reference values. Since the reference values of the number of people in the image of the population to be detected obtained by these two methods are from different dimensions, in this embodiment, the average value of these two reference values is calculated to determine, which can reduce the error between the measured value of the number of people in the image of the population to be detected and the true value of the number of people, and improve the accuracy of number detection.
[0076] In one implementation, when determining the number of people in the image of the population to be detected according to the number of people corresponding to the target population template image, the number of people corresponding to the target population template image and the reference number can be weighted and summed according to a preset strategy to obtain a weighted sum value, and the weighted sum value is used as the number of people in the image of the population to be detected.
[0077] Among them, the preset strategy can be: calculating the confidence level of the method used to obtain the number of people corresponding to the target population template image through a neural network model, and calculating the confidence level of the number detection model through a neural network model. Using the confidence levels of the above two number detection methods as their corresponding weight values, the number of people corresponding to the target population template image and the reference number are weighted and summed to obtain a weighted sum value.
[0078] For example, when obtaining the number of people corresponding to the target population template image, if the method of determining the number of people corresponding to the target population template image by face and the number of faces is used, the confidence level of this face and the number of faces recognition method can be calculated.
[0079] It should be noted that the confidence level refers to the credibility of the measured value.
[0080] Among them, the preset strategy in this application is not limited and can be set by those skilled in the art according to the actual situation.
[0081] Among them, this embodiment of the application also provides a number detection system, which includes a number detection model. Please refer to Figure 4 , Figure 4Schematic diagram of a structure of an optional number detection model provided by an embodiment of the present application. Among them, the number detection model includes a density segmentation network, an attention scale network, a density map correction module, a density map fusion module, and a number detection module. When the number detection model provided by the present application performs number detection processing on an image of a crowd to be detected, the image of the crowd to be detected is input into the density segmentation network to obtain attention masks of different density levels corresponding to the image of the crowd to be detected, and the image of the crowd to be detected is input into the attention scale network to obtain density maps of different density levels and scale factors corresponding to the image of the crowd to be detected. Then, the attention masks of different density levels, the density maps of different density levels, and the scale factors of different density levels are input into the density map correction module to obtain corrected density maps corresponding to each density level. Then, the corrected density maps of each density level are input into the density map fusion module to obtain the target density map of the image of the crowd to be detected. Finally, the target density map corresponding to the image of the crowd to be detected is input into the number detection module to obtain the reference number corresponding to the image of the crowd to be detected.
[0082] Among them, the density segmentation network can use VGG-19 as the backbone network to perform semantic segmentation of density levels on the image of the crowd to be detected, classify each pixel into a specific density level, and pixels of the same density level form a region of an attention mask.
[0083] In one implementation, before inputting the image of the crowd to be detected into the density segmentation network to obtain the attention masks corresponding to different density levels, it may further include: obtaining crowd sample images with different density distributions, and training the density segmentation network according to the crowd sample images using a loss function.
[0084] Among them, the loss function may include an adaptive pyramid loss function.
[0085] For example, the loss function used in this embodiment may be as shown in the following formula. Since the single mean square loss function (MSE) ignores the influence of different levels of density on the network training process. And the low-density and high-density distribution regions are usually quite unbalanced, and the corresponding estimation errors will cause bias in the trained counting network, which will weaken the generalization ability of the counting network. In this embodiment, an adaptive pyramid loss (APLoss) is added to the mean square loss function, which can alleviate the training bias and strengthen the generalization ability of the counting network at the same time.
[0086] L = L MSE + λL AP
[0087]
[0088] During the processing of this loss function, the density map is divided into 2×2 grids each time, and the number of people is detected for each grid. If the local count of each grid is greater than the threshold T, this grid is further divided into 2X2, and the above operation is repeated; if the given threshold T is not exceeded, no further division is performed. The local loss of the sub-region corresponding to the divided grid is calculated as follows, X k is the k-th input image, D k is its corresponding ground truth density map, is the predicted density map, is the sub-region after the n-th division.
[0089]
[0090] The total L AP is shown in the following formula, where M is the size of the training set
[0091]
[0092] In one implementation, when obtaining population sample images with different density distributions, virtual population images with different density distributions can be constructed and used as population sample images.
[0093] For example, a data collector and labeler can be designed based on the dataset of a game GTA5. It can synthesize crowd scenes and automatically label them. Thanks to the excellent game engine, its scene rendering, texture details, weather effects, etc. are very close to the real-world situation. Therefore, in this application, complex and crowded crowd scenes can be constructed by using the objects in the GTA5 virtual scene, and then the data collector captures stable images from the constructed scenes to obtain virtual population images with different density distributions. Finally, by analyzing the data from the game rendering template, automatic annotation of the head positions of people can be achieved. Through the designed collector and labeler, a large-scale and diverse synthetic crowd scene dataset can be constructed.
[0094] In addition, in this embodiment, for the generation of the density level labels of the ground truth, the specific process can be as follows:
[0095] (1) By using a 64×64 sliding window to scan the gt crowd map in the training set pixel by pixel, all local counts are obtained.
[0096] (2) Calculate the number of people in all non-zero regions to obtain the average value AvgCnt 11 and find the minimum count MinCnt and the maximum count MaxCnt.
[0097] (3) In this way, a set of density level thresholds {MinCnt, AvgCnt11 , MaxCnt}, the density can thus be divided into two levels: low density and high density. Then, the average value AvgCnt of all low-density counts can be iteratively calculated. 21 and the average value AvgCnt of all high-density counts 22 . Then, a new set of thresholds {MinCnt, AvgCnt 21 , AvgCnt 11 , AvgCnt 22 , MaxCnt} is obtained, and in this way, the density can be divided into four levels, and so on.
[0098] (4) In this way, labels can be obtained for training. Given N density levels, there are N + 1 density labels, including an additional background label, and then each pixel point in the gt crowd image is labeled with its density level according to the count.
[0099] Among them, please refer to Figure 5 , Figure 5 , which is a schematic structural diagram of the attention scale network of the method for detecting the number of people based on a crowd image template provided by the embodiment of the present application. The attention scale network includes a feature extraction backbone, a scale factor (AS, Attention Scaling) branch, and a density estimation (DE, Density Estimation) branch. The feature extraction backbone is used to extract the image features of the crowd image to be detected, and VGG-19 can be used as the feature extraction backbone. The scale factor branch is used to learn the scale factor, and then the scale factor is used to automatically adjust the estimated density of each corresponding sub-region, so as to reduce the local estimation error. The density estimation branch is used to output the density map that needs to be corrected.
[0100] Among them, the density estimation branch includes a spatial pyramid pooling layer and a convolutional layer. When the image features are input into the density estimation branch to obtain the density maps corresponding to different density levels, the image features can be input into the spatial pyramid pooling layer to obtain the spatial pyramid pooling features, and then the pyramid pooling features are input into the convolutional layer to obtain the density maps corresponding to different density levels.
[0101] In the present application, a spatial pyramid pooling layer is added to the density estimation branch. By calculating the features of different scales through the spatial pyramid pooling layer, on the basis of the features extracted by the feature extraction backbone, multi-scale context information is extracted, which is equivalent to combining the context information to predict the crowd density, making the number detection more accurate.
[0102] It should be noted that, due to the limitation of the features extracted by the VGG network used in the feature extraction backbone, which is that it encodes the same receptive field across the entire image, to solve this problem, a spatial pyramid pooling layer is added in the density estimation branch in this embodiment. By performing spatial pyramid pooling (SPP), features of different scales are calculated, and on the basis of the features extracted by the VGG network, multi-scale context information is extracted. The calculation formula is as follows:
[0103] f j = U bi (F j (P ave (f v , j), θ j )))
[0104] Where, for each scale j, P ave (f v , j) divides the feature f v extracted by the VGG network into k(j)×k(j) blocks. F j is a convolutional network with a convolutional kernel size of 1, which is used to combine context features across channels without changing its size. U bi represents bilinear interpolation, which is used to sample the output feature map f j to the same size as f v . The system adopts 4 different scales, corresponding to k(j)∈{1, 2, 3, 4}. Experiments prove that such a setting is the most effective for improving the network performance. Finally, the context features of different scales are fused with the original VGG features for outputting density maps of different density levels.
[0105] Where, when the density map correction module corrects the density maps of each density level according to the attention mask and scale factor corresponding to each density level, it multiplies the attention mask, scale factor, and density map corresponding to each density level respectively to obtain the corrected density map corresponding to each density level.
[0106] Where, when the density map fusion module fuses the corrected density maps of each density level, it can add the corrected density maps of each density level to obtain the target density map of the image of the crowd to be detected.
[0107] Where, the number detection module can perform number detection processing according to the target density map corresponding to the image of the crowd to be detected to obtain the reference number of the image of the crowd to be detected.
[0108] The above-mentioned number detection model provided by the embodiments of the present application can obtain attention masks with different density levels corresponding to the to-be-detected crowd image for which crowd counting is required, and obtain density maps with different density levels and scale factors corresponding to the to-be-detected crowd image for which crowd counting is required. Then, according to the attention masks and scale factors corresponding to each density level, correct the density maps of each density level to obtain corrected density maps of each density level, fuse the corrected density maps of each density level to obtain the target density map of the to-be-detected crowd image, and finally, determine the reference number of people in the to-be-detected crowd image according to the target density map. This number detection model can avoid the influence of the crowd distribution with different densities in different image regions of the to-be-detected crowd image on the number detection, so that the accuracy of the number measurement value obtained by this number detection model is relatively high.
[0109] As can be seen from the above, the number detection method based on the crowd image template proposed by the embodiments of the present application obtains the to-be-detected crowd image for which crowd counting is required, constructs virtual crowd images with different crowd distribution situations, uses the virtual crowd images as crowd template images, determines the candidate crowd template image with the highest similarity degree to the to-be-detected crowd image from multiple candidate crowd template images, uses the crowd template image with the highest similarity degree as the target crowd template image, and determines the number of people in the to-be-detected crowd image according to the number corresponding to the target crowd template image. The present application directly determines the number of people in the to-be-detected crowd image by obtaining the target crowd template image with the highest similarity degree to the to-be-detected crowd image and according to the number corresponding to the target crowd template image, which can improve the efficiency of number detection. In addition, when determining the number of people in the to-be-detected crowd image, the present application can also use the number corresponding to the target crowd template image as a reference value, and also use the reference number of people obtained by inputting the to-be-detected crowd image into the number detection model as a reference value, and determine the number of people in the to-be-detected crowd image according to these two reference values, improving the accuracy of number detection.
[0110] In an embodiment, a number detection device based on a crowd image template is also provided. Please refer to Figure 6 , Figure 6 is a schematic structural diagram of the number detection device 300 based on the crowd image template provided by the embodiments of the present application. The number detection device 300 based on the crowd image template is applied to an electronic device. The number detection device 300 based on the crowd image template includes an acquisition module 301, a first determination module 302, a second determination module 303, and a third determination module 304, as follows:
[0111] The acquisition module 301 is configured to acquire a to-be-detected crowd image for which crowd counting is required;
[0112] The first determination module 302 is configured to determine, from multiple candidate population template images, the candidate population template image with the highest similarity to the to-be-detected population image;
[0113] The second determination module 303 is configured to use the candidate population template image with the highest similarity as the target population template image;
[0114] The third determination module 304 is configured to determine the number of people in the to-be-detected population image according to the number of people corresponding to the target population template image.
[0115] In one implementation, the first determination module 302 may be configured to: input the to-be-detected population image into a population image similarity model to obtain the similarity between the to-be-detected population image and each candidate population template image; determine the candidate population template image with the highest similarity to the to-be-detected population image according to the similarity.
[0116] In one implementation, the acquisition module 301 may further be configured to: construct virtual population images with different population distribution situations; use the virtual population images as population template images.
[0117] In one implementation, the acquisition module 301 may further be configured to: acquire the rendering data corresponding to each constructed virtual population image; determine the number of people in the population template image according to the rendering data; establish a correspondence between the population template image and the number of people according to the population template image and the number of people in the population template image; determine the number of people corresponding to the target population template image according to the correspondence.
[0118] In one implementation, the third determination module 304 may be configured to: acquire the average value of the number of people corresponding to the target population template image and the reference number of people, and use the average value as the number of people in the to-be-detected population image.
[0119] In one implementation, the third determination module 304 may be configured to: perform a weighted summation process on the number of people corresponding to the target population template image and the reference number of people according to a preset policy to obtain a weighted sum value, and use the weighted sum value as the number of people in the to-be-detected population image.
[0120] It should be noted that the number detection device based on the population image template provided in the embodiments of the present application and the number detection method based on the population image template in the above embodiments belong to the same concept. Through the number detection device based on the population image template, any method provided in the embodiments of the number detection method based on the population image template can be implemented. The specific implementation process is detailed in the embodiments of the number detection method based on the population image template, and will not be elaborated here.
[0121] As described above, the number detection device based on the population image template proposed in the embodiments of the present application obtains the to-be-detected population image that needs to be counted by the acquisition module 301, determines the candidate population template image with the highest similarity to the to-be-detected population image from multiple candidate population template images by the first determination module 302, uses the candidate population template image with the highest similarity as the target population template image by the second determination module 303, and determines the number of people in the to-be-detected population image according to the number of people corresponding to the target population template image by the third determination module 304. The present application directly obtains the target population template image that matches the to-be-detected population image, and determines the number of people in the to-be-detected population image according to the number of people corresponding to the target population template image, which can improve the efficiency of number detection.
[0122] The embodiments of the present application also provide an electronic device. The electronic device may be a device such as a smart phone or a tablet computer. Please refer to Figure 7 , Figure 7 which is the first structural schematic diagram of the electronic device provided by the embodiments of the present application. The electronic device 400 includes a processor 401 and a memory 402. Among them, the processor 401 is electrically connected to the memory 402.
[0123] The processor 401 is the control center of the electronic device 400, connects various parts of the entire electronic device through various interfaces and lines, executes various functions of the electronic device and processes data by running or calling the computer program stored in the memory 402 and calling the data stored in the memory 402, so as to monitor the electronic device as a whole.
[0124] The memory 402 can be used to store computer programs and data. The computer program stored in the memory 402 contains instructions that can be executed in the processor. The computer program can form various functional modules. The processor 401 executes various functional applications and data processing by calling the computer program stored in the memory 402.
[0125] In this embodiment, the processor 401 in the electronic device 400 will load the instructions corresponding to the processes of one or more computer programs into the memory 402 according to the following steps, and the processor 401 will run the computer program stored in the memory 402 to implement various functions:
[0126] Obtain the to-be-detected population image that needs to be counted;
[0127] Determine the candidate population template image with the highest similarity to the to-be-detected population image from multiple candidate population template images;
[0128] Use the candidate population template image with the highest similarity as the target population template image;
[0129] Determine the number of people in the to-be-detected population image according to the number of people corresponding to the target population template image.
[0130] In one implementation, refer to Figure 8 , Figure 8 which is the second structural schematic diagram of the electronic device provided by the embodiment of the present application. The electronic device 400 further includes: a radio frequency circuit 403, a display screen 404, a control circuit 405, an input unit 406, an audio circuit 407, a sensor 408, and a power supply 409. Among them, the processor 401 is electrically connected to the radio frequency circuit 403, the display screen 404, the control circuit 405, the input unit 406, the audio circuit 407, the sensor 408, and the power supply 409 respectively.
[0131] The radio frequency circuit 403 is used to receive and transmit radio frequency signals to communicate with a network device or other electronic devices through wireless communication.
[0132] The display screen 404 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of the electronic device. These graphical user interfaces can be composed of images, texts, icons, videos, and any combination thereof.
[0133] The control circuit 405 is electrically connected to the display screen 404 and is used to control the display screen 404 to display information.
[0134] The input unit 406 can be used to receive input digital, character information, or user feature information (such as fingerprints), and generate keyboard, mouse, joystick, optical, or trackball signal inputs related to user settings and function controls. Among them, the input unit 406 can include a fingerprint recognition module.
[0135] The audio circuit 407 can provide an audio interface between the user and the electronic device through a speaker and a microphone. Among them, the audio circuit 407 includes a microphone. The microphone is electrically connected to the processor 401. The microphone is used to receive voice information input by the user.
[0136] The sensor 408 is used to collect external environmental information. The sensor 408 can include one or more of sensors such as an ambient light sensor, an acceleration sensor, and a gyroscope.
[0137] The power supply 409 is used to supply power to each component of the electronic device 400. In one implementation, the power supply 409 can be logically connected to the processor 401 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system.
[0138] Although not shown in the figure, the electronic device 400 may further include a camera, a Bluetooth module, etc., which will not be elaborated here.
[0139] In this embodiment, the processor 401 in the electronic device 400 will load the instructions corresponding to the processes of one or more computer programs into the memory 402 according to the following steps, and the processor 401 will run the computer programs stored in the memory 402 to implement various functions:
[0140] Obtain a to-be-detected crowd image for which crowd counting is to be performed;
[0141] Determine the candidate crowd template image with the highest similarity to the to-be-detected crowd image from multiple candidate crowd template images;
[0142] Use the candidate crowd template image with the highest similarity as the target crowd template image;
[0143] Determine the number of people in the to-be-detected crowd image according to the number of people corresponding to the target crowd template image.
[0144] In one implementation, when the processor 401 executes to determine the candidate crowd template image with the highest similarity to the to-be-detected crowd image from multiple candidate crowd template images, it may execute: input the to-be-detected crowd image into a crowd image similarity model to obtain the similarity between the to-be-detected crowd image and each candidate crowd template image; determine the candidate crowd template image with the highest similarity to the to-be-detected crowd image according to the similarity.
[0145] In one implementation, before the processor 401 executes to determine the candidate crowd template image with the highest similarity to the to-be-detected crowd image from multiple candidate crowd template images, it may also execute: obtain the rendering data corresponding to each virtual crowd image; determine the number of people in the crowd template image according to the rendering data; establish a correspondence between the crowd template image and the number of people according to the crowd template image and the number of people in the crowd template image; determine the number of people corresponding to the target crowd template image according to the correspondence.
[0146] In one implementation, when the processor 401 executes to determine the number of people in the to-be-detected crowd image according to the target number of people, it may execute: input the to-be-detected crowd image into a number detection model to obtain the reference number of people corresponding to the to-be-detected crowd image; determine the number of people in the to-be-detected crowd image according to the number of people corresponding to the target crowd template image and the reference number of people.
[0147] In one embodiment, when the processor 401 determines the number of people in the to-be-detected crowd image according to the number of people corresponding to the target crowd template image and the reference number of people, it may perform: obtaining the average value of the number of people corresponding to the target crowd template image and the reference number of people, and using the average value as the number of people in the to-be-detected crowd image.
[0148] In one embodiment, when the processor 401 determines the number of people in the to-be-detected crowd image according to the number of people corresponding to the target crowd template image and the reference number of people, it may perform: performing a weighted summation process on the number of people corresponding to the target crowd template image and the reference number of people according to a preset policy to obtain a weighted sum value, and using the weighted sum value as the number of people in the to-be-detected crowd image.
[0149] The embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. When the computer program runs on a processor, the computer executes the method for detecting the number of people based on a crowd image template described in any one of the above embodiments.
[0150] It should be noted that those of ordinary skill in the art can understand that all or part of the steps in the above various methods can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the storage medium may include, but is not limited to: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.
[0151] In addition, the terms "first", "second", "third", etc. in the present application are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or modules is not limited to the listed steps or modules, but some embodiments further include steps or modules that are not listed, or some embodiments further include other steps or modules inherent to these processes, methods, products or devices.
[0152] The method for detecting the number of people based on a crowd image template provided by the embodiments of the present application has been introduced in detail above. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A method for detecting the number of people based on a crowd image template, characterized in that Including: Obtaining a to-be-detected crowd image that needs to be counted; Constructing virtual crowd images with different crowd distribution situations; Using the virtual crowd image as a crowd template image; Determining a candidate crowd template image with the highest similarity degree to the to-be-detected crowd image from multiple candidate crowd template images; Using the candidate crowd template image with the highest similarity degree as the target crowd template image; Inputting the to-be-detected crowd image into a number detection model to obtain a reference number corresponding to the to-be-detected crowd image; Determining the number of people in the to-be-detected crowd image according to the number corresponding to the target crowd template image and the reference number; Wherein, the number detection model includes a density segmentation network, an attention scale network, a density map correction module, a density map fusion module, and a number detection module. The attention scale network includes a feature extraction backbone, a scale factor branch, and a density estimation branch. The density estimation branch includes a spatial pyramid pooling layer and a convolutional layer. The feature extraction backbone is used to extract image features of the to-be-detected crowd image. The scale factor branch is used to learn a scale factor and use the scale factor to adjust the estimated density of each corresponding sub-region. The density estimation branch is used to output a density map that needs to be corrected. The density segmentation network is trained through a loss function, and the loss function includes an adaptive pyramid loss function; The step of inputting the to-be-detected crowd image into a number detection model to obtain a reference number corresponding to the to-be-detected crowd image includes: Inputting the to-be-detected crowd image into the density segmentation network to obtain attention masks of different density levels corresponding to the to-be-detected crowd image; Inputting the to-be-detected crowd image into the attention scale network to obtain density maps and scale factors of different density levels corresponding to the to-be-detected crowd image; Inputting the attention masks, the density maps, and the scale factors of different density levels into the density map correction module to obtain corrected density maps corresponding to each density level; Inputting the corrected density maps of each density level into the density map fusion module to obtain a target density map of the to-be-detected crowd image; Inputting the target density map corresponding to the to-be-detected crowd image into the number detection module to obtain the reference number corresponding to the to-be-detected crowd image.
2. The method for detecting the number of people based on the crowd image template according to claim 1, wherein The step of determining a candidate crowd template image with the highest similarity degree to the to-be-detected crowd image from multiple candidate crowd template images includes: Inputting the to-be-detected crowd image into a crowd image similarity model to obtain the similarity degree between the to-be-detected crowd image and each candidate crowd template image; Determining a candidate crowd template image with the highest similarity degree to the to-be-detected crowd image according to the similarity degree.
3. The method for detecting the number of people based on the population image template according to claim 1, wherein Before determining the number of people in the to-be-detected crowd image according to the number corresponding to the target crowd template image, it further includes: Obtaining rendering data corresponding to each virtual crowd image constructed; Determining the number of people in the crowd template image according to the rendering data; Establishing a corresponding relationship between the crowd template image and the number of people according to the crowd template image and the number of people in the crowd template image; Determine the number of people corresponding to the target population template image according to the corresponding relationship.
4. The method for detecting the number of people based on the crowd image template according to claim 1, wherein Determining the number of people in the to-be-detected population image according to the number of people corresponding to the target population template image and the reference number of people includes: Obtain the average value of the number of people corresponding to the target population template image and the reference number of people, and use the average value as the number of people in the to-be-detected population image.
5. The method for detecting the number of people based on the population image template according to claim 1, wherein, Determining the number of people in the to-be-detected population image according to the number of people corresponding to the target population template image and the reference number of people includes: Perform weighted summation processing on the number of people corresponding to the target population template image and the reference number of people according to a preset strategy to obtain a weighted sum value, and use the weighted sum value as the number of people in the to-be-detected population image.
6. A device for detecting the number of people based on a crowd image template, characterized in that, Includes: An acquisition module, configured to acquire a to-be-detected population image that needs to be counted for the number of people; A first determination module, configured to construct virtual population images with different population distribution situations; Use the virtual population image as a population template image; determine the candidate population template image with the highest similarity to the to-be-detected population image from multiple candidate population template images; A second determination module, configured to use the candidate population template image with the highest similarity as the target population template image; A third determination module, input the to-be-detected population image into a number detection model to obtain the reference number of people corresponding to the to-be-detected population image; determine the number of people in the to-be-detected population image according to the number of people corresponding to the target population template image and the reference number of people; Wherein, the number detection model includes a density segmentation network, an attention scale network, a density map correction module, a density map fusion module, and a number detection module, and the attention scale network includes a feature extraction backbone, a scale factor branch, and a density estimation branch; The density estimation branch includes a spatial pyramid pooling layer and a convolutional layer. The feature extraction backbone is used to extract the image features of the to-be-detected population image. The scale factor branch is used to learn a scale factor and use the scale factor to adjust the estimated density of each corresponding sub-region. The density estimation branch is used to output a density map that needs to be corrected; The density segmentation network is trained through a loss function, and the loss function includes an adaptive pyramid loss function; Inputting the to-be-detected population image into the number detection model to obtain the reference number of people corresponding to the to-be-detected population image includes: Input the to-be-detected population image into the density segmentation network to obtain attention masks of different density levels corresponding to the to-be-detected population image; Input the to-be-detected population image into the attention scale network to obtain density maps and scale factors of different density levels corresponding to the to-be-detected population image; Input the attention masks, the density maps, and the scale factors of different density levels into the density map correction module to obtain corrected density maps corresponding to each density level; Input the corrected density maps of each density level into the density map fusion module to obtain the target density map of the to-be-detected population image; Input the target density map corresponding to the image of the population to be detected into the number detection module to obtain the reference number corresponding to the image of the population to be detected.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program runs on a processor, it causes the computer to execute the method for detecting the number of people based on a population image template according to any one of claims 1 to 5.
8. An electronic device, comprising a processor and a memory, the memory storing a computer program, characterized in that, The processor is configured to execute the method for detecting the number of people based on a population image template according to any one of claims 1 to 5 by calling the computer program.
Citation Information
Patent Citations
Counting model processing method and device based on target detection and computer equipment
CN114241411A
Method for determining stationary crowds
WO2016019973A1