Crowd Counting Prediction Method, System, Computer Device and Storage Medium

Optimizing the convolutional neural network through point-annotated probability regression and the loss function of optimal transmission theory, the problem of target scale changes and high computational cost in dense crowd counts is solved, and more accurate crowd count prediction is achieved.

CN114580731BActive Publication Date: 2025-07-22XI AN JIAOTONG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210192825.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-28
Publication Date
2025-07-22
Estimated Expiration
2042-02-28

AI Technical Summary

Technical Problem

The prior art has problems with lack of target scale information and large-scale scale changes in dense population counts, resulting in inaccurate population counts and high computational cost.

Method used

The point annotation probability regression method is used to extract the point annotation probability map of the image through a convolutional neural network, and a loss function based on the optimal transmission theory is constructed, the transmission cost of the significance region is defined, and the point annotation of the training image is directly used as the supervision signal to optimize the convolutional neural network parameters to generate an accurate point annotation probability distribution.

Benefits of technology

It improves the accuracy of population counting, reduces dependence on target scale information, reduces calculation costs, and improves the prediction effect of population counting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114580731B_ABST
    Figure CN114580731B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of computer vision, and discloses a crowd counting prediction method, system, computer device and storage medium, including: obtaining an image to be predicted; inputting the image to be predicted into a preset point annotation feature extraction network to obtain a point annotation probability map of the image to be predicted; summing up the point annotation probability map of the image to be predicted to obtain a crowd counting result of the image to be predicted. By formulating crowd counting as a point annotation probability regression problem, using the point annotation probability map as the prediction target, representing the statistical probability distribution of the manual annotation process, and realizing crowd counting based on the predicted point annotation probability map, due to the characteristic that the point annotation probability map does not need to consider the scale information of the target, it is easier to generate than the density map, and can better realize crowd counting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision and relates to a method, system, computer device and storage medium for crowd counting prediction. Background Art

[0002] Crowd counting is the estimation of the number of people in a static image. Due to its wide applications in crowd analysis, urban planning and video surveillance, crowd counting methods have received increasing attention. Although crowd counting is useful, its practical applications are still limited because there are some challenges in crowd counting in dense contexts. In dense crowds, providing point annotations requires less labor than bounding box annotations. Therefore, supervised methods in dense scenes generally use point annotations as labels. A point annotation is a sparse matrix with a large number of 0s and a small number of 1s, and it does not contain information about the target scale. Considering the sparsity of points, it is difficult to define a pixel-level loss function to convert image observations (dense real-value matrices) into point annotations (sparse binary matrices). Therefore, one of the greatest challenges in crowd counting lies in how to design effective learning objectives to address the discreteness between image observations and point annotations.

[0003] Some algorithms consider crowd counting as a density map regression task, and they place a Gaussian kernel on each point annotation to alleviate this discreteness. From the perspective of density map estimation, researchers need to design a ground-truth density map based on image observations: design a large Gaussian kernel for large targets and a small Gaussian kernel for small targets.

[0004] However, these density map-based methods are subject to two limitations: inaccurate ground-truth density maps and large-scale variation problems. First, due to the lack of target scale information, it is difficult to generate accurate ground-truth density maps. To solve this problem, some algorithms scale the Gaussian kernel according to the scene perspective or design an adaptive Gaussian by using the local congestion level (calculating the distance to the nearest neighbor). There are also some methods that propose an adaptive density map generator for generating a learnable density map representation. Although these methods have achieved excellent counting performance in crowd counting, they still cannot set a suitable Gaussian kernel for each target. Second, because the scale of the target varies greatly within an image and between different images, designing a density map based on the scale of the observed target will naturally introduce the large-scale variation problem. Researchers have proposed various algorithms and network architectures to handle the large-scale variation problem, but due to the continuous variation of the scale throughout the image, it is difficult for these methods to completely solve this large-scale variation problem. In addition, extracting multi-scale feature information will bring more computational costs. Summary of the Invention

[0005] The object of the present invention is to overcome the above-mentioned disadvantages of the prior art and provide a method, system, computer device and storage medium for crowd counting prediction.

[0006] To achieve the above object, the present invention is implemented by the following technical solutions:

[0007] In a first aspect of the present invention, a method for crowd counting prediction includes:

[0008] Obtain an image to be predicted;

[0009] Input the image to be predicted into a preset point annotation feature extraction network to obtain a point annotation probability map of the image to be predicted;

[0010] Sum up the point annotation probability map of the image to be predicted to obtain the crowd counting result of the image to be predicted.

[0011] Optionally, the point annotation feature extraction network is obtained by the following method:

[0012] Obtain a convolutional neural network, training images, and point annotations of the training images;

[0013] Use the convolutional neural network to extract features from the training images to obtain a point annotation probability map of the training images;

[0014] Normalize the point annotation probability map of the training images to obtain a point annotation probability distribution of the training images;

[0015] Construct a loss function based on the optimal transport theory, and according to the loss function and the point annotations of the training images, obtain the transport cost of converting the point annotations into the point annotation probability distribution. Iteratively optimize the parameters of the convolutional neural network according to the transport cost until the preset number of iterations;

[0016] Obtain the convolutional neural network with the minimum transport cost to obtain the point annotation feature extraction network.

[0017] Optionally, the convolutional neural network is a VGG convolutional neural network.

[0018] Optionally, the normalization of the point annotation probability map of the training images includes:

[0019] Normalize the point annotation probability map of the training images by the following formula:

[0020]

[0021] where P is the point annotation probability map of the training images, is the point annotation probability distribution of the training images, and ||·||1 is the L1 norm.

[0022] Optionally, when constructing the loss function based on the optimal transport theory, construct the transport cost function of the loss function according to the relative entropy.

[0023] Optionally, the transmission cost function is:

[0024]

[0025] where is the transmission cost function, representing the cost of transmitting pixel point a i to pixel point ; κ δ is a Gaussian kernel with bandwidth δ, is the point annotation of the training image in the d-dimensional space, a i is the pixel point among them, and n is the number of pixel points of the point annotation of the training image; is the probability distribution of the point annotation of the training image in the d-dimensional space, is the pixel point among them, and m is the number of pixel points of the probability distribution of the point annotation of the training image; ||·|| 2 is the L2 norm.

[0026] Optionally, the loss function based on the optimal transport theory is:

[0027]

[0028]

[0029] where is the loss function, is the optimal transport loss in the loss function, S is the transport matrix for mapping each pixel point value in A to P, P is the probability map of the point annotation of the training image, U is the set of all transport matrices, and ||·||1 is the L1 norm.

[0030] In the second aspect of the present invention, a crowd counting prediction system includes:

[0031] An acquisition module for acquiring an image to be predicted;

[0032] A prediction module for inputting the image to be predicted into a preset point annotation feature extraction network to obtain a probability map of the point annotation of the image to be predicted;

[0033] A summation module for summing the probability map of the point annotation of the image to be predicted to obtain the crowd counting result of the image to be predicted.

[0034] In the third aspect of the present invention, a computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned crowd counting prediction method are implemented.

[0035] In a fourth aspect of the present invention, a computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned crowd counting prediction method are implemented.

[0036] Compared with the prior art, the present invention has the following beneficial effects:

[0037] In the crowd counting prediction method of the present invention, the obtained image to be predicted is input into a preset point annotation feature extraction network to obtain the point annotation probability map of the image to be predicted, and then the point annotation probability maps of the image to be predicted are summed to obtain the crowd counting result of the image to be predicted. The crowd counting is formulated as a point annotation probability regression problem, with the point annotation probability map as the prediction target, representing the statistical probability distribution of the manual annotation process, and the crowd counting is realized based on the predicted point annotation probability map. Based on the characteristic that the scale information of the target does not need to be considered for the point annotation probability map, it is easier to generate than the density map, and the crowd counting can be better realized, ensuring the accuracy of the crowd counting result.

[0038] Furthermore, a transmission cost function of the loss function is constructed according to the relative entropy to establish the significant region when setting the annotation target. Considering that the significant regions of targets of different sizes are roughly fixed, this method of establishing the significant region of the annotation target will not introduce large-scale change problems. And different transmission costs are defined for different annotation positions (inside / outside the significant region), thereby effectively improving the training effect of the convolutional neural network and finally improving the crowd counting prediction accuracy. At the same time, the point annotation of the training image is directly used as the supervision signal of the convolutional neural network, so that the supervision signal is more accurate, making the point annotation probability distribution of the training image obtained by the convolutional neural network approach the distribution of the point annotation of the training image, and further making the point annotation probability distribution predicted by the convolutional neural network more accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 is a flowchart of the crowd counting prediction method according to an embodiment of the present invention;

[0040] Figure 2 is a flowchart of the construction of the point annotation feature extraction network according to an embodiment of the present invention;

[0041] Figure 3 is a schematic diagram of the principle of the construction of the point annotation feature extraction network according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0043] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0044] The present invention will be further described in detail below with reference to the accompanying drawings:

[0045] See Figure 1 , in an embodiment of the present invention, a crowd counting prediction method is provided to achieve accurate estimation of crowd counting. Specifically, the crowd counting prediction method includes the following steps:

[0046] S1: Obtain the image to be predicted.

[0047] Specifically, the image to be predicted is generally an RGB image.

[0048] S2: Input the image to be predicted into a preset point annotation feature extraction network to obtain the point annotation probability map of the image to be predicted.

[0049] Specifically, when predicting the image to be predicted, it is necessary to pre-obtain the point annotation feature extraction network, and then extract the point annotation probability map of the image to be predicted based on the point annotation feature extraction network.

[0050] See Figure 2 and 3 , in this embodiment, the following construction method is provided for the point annotation feature extraction network:

[0051] S201: Obtain a convolutional neural network, training images, and point annotations of the training images.

[0052] In a possible implementation, the convolutional neural network adopts the VGG (Visual Geometry Group, the computer vision group of the University of Oxford) convolutional neural network. VGG explores the relationship between the depth of the convolutional neural network and its performance. By repeatedly stacking small convolutional kernels of 3×3 and max pooling layers of 2×2, the VGG convolutional neural network has successfully constructed a convolutional neural network with a depth of 16 to 19 layers, effectively improving the accuracy of the convolutional neural network, and its generalization on various image data is also very good.

[0053] S202: Use a convolutional neural network to extract features from the training image to obtain the point annotation probability map of the training image.

[0054] Specifically, use the convolutional neural network f to extract features from the input image I to obtain the point annotation probability map P. The specific calculation formula is as follows:

[0055] P = f(I, θ)

[0056] where I is the input training image, f(·, θ) is the convolutional neural network with parameter θ, and P is the point annotation probability map of the training image output by the convolutional neural network.

[0057] S203: Normalize the point annotation probability map of the training image to obtain the point annotation probability distribution of the training image.

[0058] Specifically, the point annotation probability map of the training image is normalized by the following formula:

[0059]

[0060] where P is the point annotation probability map of the training image, is the point annotation probability distribution of the training image, and ||·||1 is the L1 norm.

[0061] S204: Construct a loss function based on the optimal transport theory, and according to the loss function and the point annotation of the training image, obtain the transport cost of converting the point annotation into the point annotation probability distribution. Iteratively optimize the parameters of the convolutional neural network according to the transport cost until the preset number of iterations.

[0062] Specifically, in this embodiment, when constructing the loss function based on the optimal transport theory, a transport cost function of the loss function is constructed according to the relative entropy. The transport cost function is:

[0063]

[0064] where, is the transport cost function, indicating the cost of transporting the pixel point a i to the pixel point ; κδ is a Gaussian kernel with a bandwidth of δ, where δ controls the size of the significant region of the image; is a coefficient related to the bandwidth of the Gaussian kernel, is the point annotation of the training image in the d-dimensional space, a i is the pixel point among them, and n is the number of pixel points of the point annotation of the training image; is the probability distribution of the point annotation of the training image in the d-dimensional space, is the pixel point among them, and m is the number of pixel points of the probability distribution of the point annotation of the training image; ||·|| 2 is the L2 norm.

[0065] To reflect the marking habits of annotators, a transportation cost function for constructing the loss function is established based on the relative entropy, so as to establish the significant region when defining the annotation target. Considering that the significant regions of targets of different sizes are roughly fixed, this method of establishing the significant region of the annotation target will not introduce large-scale change problems. And different transportation costs are defined for different annotation positions (inside / outside the significant region), thereby effectively improving the training effect of the convolutional neural network and ultimately improving the accuracy of crowd counting prediction.

[0066] In this embodiment, the loss function based on the optimal transportation theory is:

[0067]

[0068]

[0069] where, is the loss function, is the optimal transportation loss in the loss function, S is the transportation matrix used to map each pixel point value in A to P, and U is the set of all transportation matrices.

[0070] Specifically, the above loss function based on the optimal transportation theory directly uses the point annotation of the training image as the supervision signal of the convolutional neural network, so that the supervision signal is more accurate, making the probability distribution of the point annotation of the training image obtained by the convolutional neural network approach the distribution of the point annotation of the training image, and thus making the probability distribution of the point annotation predicted by the convolutional neural network more accurate.

[0071] Finally, the parameters of the convolutional neural network are modified according to the transportation cost, and the above steps are repeated using the modified convolutional neural network to achieve the iterative optimization of the convolutional neural network. In this embodiment, the iteration number threshold of the iterative optimization is set, that is, the preset iteration number. When the current iteration number is the preset iteration number, it indicates that the iterative optimization process is completed.

[0072] S205: Obtain the convolutional neural network with the minimum transmission cost to get the point annotation feature extraction network.

[0073] Specifically, after completing the iterative optimization process, select the convolutional neural network with the minimum transmission cost from among convolutional neural networks with different parameters. This indicates that the convolutional neural network has the best prediction effect on images. Therefore, this convolutional neural network is used as the final point annotation feature extraction network.

[0074] S3: Sum the point annotation probability maps of the image to be predicted to obtain the crowd counting result of the image to be predicted.

[0075] In summary, for the crowd counting prediction method of the present invention, the obtained image to be predicted is input into a preset point annotation feature extraction network to obtain the point annotation probability map of the image to be predicted, and then the point annotation probability maps of the image to be predicted are summed to obtain the crowd counting result of the image to be predicted. The crowd counting is formulated as a point annotation probability regression problem, with the point annotation probability map as the prediction target, representing the statistical probability distribution of the manual annotation process, and the crowd counting is realized based on the predicted point annotation probability map. Due to the characteristic that the scale information of the target does not need to be considered for the point annotation probability map, it is easier to generate than the density map and can better achieve crowd counting, ensuring the accuracy of the crowd counting result.

[0076] In a possible implementation manner, the pre-trained VGG convolutional neural network on the image-net dataset is used as the backbone network for feature extraction. The point annotation is directly used as the supervision signal of the VGG convolutional neural network, and a loss function based on the optimal transport theory is designed to measure the similarity between the potential point annotation probability distribution and the point annotation distribution. On the UCF-QNRF dataset, experiments are carried out with σ = 16 as the bandwidth of the relative entropy. The method of the present invention and the existing methods are used to train and model the dense crowd data of UCF-QNRF, and the UCF-QNRF test set is tested. The experimental results are shown in Table 1.

[0077] Table 1

[0078]

[0079]

[0080] It can be seen from the experimental results that, compared with the existing methods, the method of the present invention achieves higher test accuracy on the same dataset.

[0081] The following is an apparatus embodiment of the present invention, which can be used to execute the method embodiment of the present invention. For details not disclosed in the apparatus embodiment, please refer to the method embodiment of the present invention.

[0082] In another embodiment of the present invention, a crowd counting prediction system is provided, which can be used to implement the above-mentioned crowd counting prediction method. The crowd counting prediction system includes an acquisition module, a prediction module, and a summation module.

[0083] Among them, the acquisition module is used to acquire the image to be predicted; the prediction module is used to input the image to be predicted into a preset point annotation feature extraction network to obtain the point annotation probability map of the image to be predicted; the summation module is used to sum the point annotation probability map of the image to be predicted to obtain the crowd counting result of the image to be predicted.

[0084] In a possible implementation manner, the point annotation feature extraction network is obtained by the following method: acquiring a convolutional neural network, training images, and point annotations of the training images; using the convolutional neural network to extract features from the training images to obtain the point annotation probability map of the training images; normalizing the point annotation probability map of the training images to obtain the point annotation probability distribution of the training images; constructing a loss function based on the optimal transport theory, and according to the loss function and the point annotations of the training images, obtaining the transport cost of converting the point annotations into the point annotation probability distribution, and iteratively optimizing the parameters of the convolutional neural network according to the transport cost until a preset number of iterations; acquiring the convolutional neural network with the minimum transport cost to obtain the point annotation feature extraction network.

[0085] In a possible implementation manner, the convolutional neural network is a VGG convolutional neural network.

[0086] In a possible implementation manner, the normalization of the point annotation probability map of the training images includes: normalizing the point annotation probability map of the training images through the following formula:

[0087]

[0088] where P is the point annotation probability map of the training images, is the point annotation probability distribution of the training images, and ||·||1 is the L1 norm.

[0089] In a possible implementation manner, when constructing the loss function based on the optimal transport theory, a transport cost function of the loss function is constructed according to the relative entropy.

[0090] In a possible implementation manner, the transport cost function is:

[0091]

[0092] where is the transport cost function, indicating the cost of transporting the pixel point a i to the pixel point ; κ δ is a Gaussian kernel with a bandwidth of δ, is the point annotation of the training image in the d-dimensional space, a i is the pixel point therein, and n is the number of pixel points of the point annotation of the training image; is the probability distribution of the point annotation of the training image in the d-dimensional space, is the pixel point therein, and m is the number of pixel points of the probability distribution of the point annotation of the training image; ||·|| 2 is the L2 norm.

[0093] In a possible implementation manner, the loss function based on the optimal transport theory is:

[0094]

[0095]

[0096] wherein, is the loss function, is the optimal transport loss in the loss function, S is the transport matrix for mapping each pixel point value in A to P, P is the probability map of the point annotation of the training image, U is the set of all transport matrices, and ||·||1 is the L1 norm.

[0097] All relevant contents of each step involved in the embodiments of the foregoing crowd counting prediction method can be cited in the function descriptions of the corresponding functional modules of the crowd counting prediction system in the embodiments of the present invention, and will not be elaborated herein. The division of modules in the embodiments of the present invention is illustrative, only a logical function division, and there may be other division methods in actual implementation. In addition, in each embodiment of the present invention, the functional modules can be integrated in a processor, or can exist separately physically, or two or more modules can be integrated in one module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules.

[0098] In another embodiment of the present invention, a computer device is provided. The computer device includes a processor and a memory. The memory is used to store a computer program, and the computer program includes program instructions. The processor is used to execute the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions in the computer storage medium to implement the corresponding method flow or corresponding function. The processor described in the embodiment of the present invention can be used for the operation of the crowd counting prediction method.

[0099] In another embodiment of the present invention, a storage medium is also provided, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in the computer device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and, of course, the extended storage medium supported by the computer device. The computer-readable storage medium provides a storage space, and the operating system of the terminal is stored in this storage space. And, one or more instructions suitable for being loaded and executed by the processor are also stored in this storage space. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. One or more instructions stored in the computer-readable storage medium can be loaded and executed by the processor to implement the corresponding steps of the crowd counting prediction method in the above embodiments.

[0100] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0101] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0102] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0103] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0104] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: still can modify the specific implementation manners of the present invention or make equivalent replacements, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the protection scope of the claims of the present invention.

Claims

1. A crowd counting prediction method, characterized in that, It includes: Obtain the image to be predicted; Input the image to be predicted into a preset point annotation feature extraction network to obtain the point annotation probability map of the image to be predicted; Sum up the point annotation probability maps of the image to be predicted to obtain the crowd counting result of the image to be predicted; The point annotation feature extraction network is obtained through the following method: Obtain a convolutional neural network, training images, and the point annotations of the training images; Use the convolutional neural network to extract features from the training images to obtain the point annotation probability maps of the training images; Normalize the point annotation probability maps of the training images to obtain the point annotation probability distributions of the training images; Construct a loss function based on the optimal transport theory, and according to the loss function and the point annotations of the training images, obtain the transport cost of converting the point annotations into point annotation probability distributions. Iteratively optimize the parameters of the convolutional neural network according to the transport cost until the preset number of iterations; Obtain the convolutional neural network with the minimum transport cost to obtain the point annotation feature extraction network; When constructing the loss function based on the optimal transport theory, construct the transport cost function of the loss function according to the relative entropy; The transport cost function is: Among them, is the transmission cost function, representing the cost of transmitting pixel point a i to pixel point ; κ δ is a Gaussian kernel with bandwidth δ, is the point annotation of the training image in the d-dimensional space, a i is the pixel point among them, and n is the number of pixel points of the point annotation of the training image; is the probability distribution of the point annotation of the training image in the d-dimensional space, is the pixel point among them, and m is the number of pixel points of the probability distribution of the point annotation of the training image; ||·|| 2 is the L2 norm; The loss function based on the optimal transport theory is: Among them, is the loss function, is the optimal transport loss in the loss function, S is the transport matrix for mapping each pixel value in A to P, P is the point annotation probability map of the training image, U is the set of all transport matrices, and ||·||1 is the L1 norm.

2. The population count prediction method according to claim 1, characterized in that The convolutional neural network is a VGG convolutional neural network.

3. The population counting and prediction method according to claim 1, wherein The normalization of the point annotation probability maps of the training images includes: Normalize the point annotation probability maps of the training images through the following formula: where P is the point annotation probability map of the training image, is the point annotation probability distribution of the training image, and ||·||1 is the L1 norm.

4. A crowd counting prediction system based on the crowd counting prediction method according to claim 1, characterized in that, It includes: An acquisition module for obtaining the image to be predicted; A prediction module for inputting the image to be predicted into a preset point annotation feature extraction network to obtain the point annotation probability map of the image to be predicted; A summation module for summing up the point annotation probability maps of the image to be predicted to obtain the crowd counting result of the image to be predicted.

5. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the crowd counting prediction method according to any one of claims 1 to 3.

6. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the crowd counting prediction method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Crowd density and quantity estimation method based on convolutional neural network

    CN111209892A