A Video Image Crowd Counting Method Based on Multi-Scale Attention Mechanism
By constructing a multi-scale attention mechanism module and VGG-16 neural network, combined with expansion convolution and loss function training, the accuracy of population counting in complex scenarios is solved, and high-precision population counting in crowd scale changes and complex contexts are achieved.
Patent Information
- Application Number
- CN202211088471.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-07
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-09-07
AI Technical Summary
In complex scenarios, the prior art is difficult to ensure the accuracy of population counting, especially the counting accuracy problems caused by large population size changes and complex background.
The video image population counting method based on the multi-scale attention mechanism is adopted. By constructing a multi-scale attention mechanism module, combining VGG-16 neural network and expanded convolution, a population counting model is designed, and the weighted functions of multi-column variance loss and Euclidean loss are trained to reduce the impact of background noise and enhance the scale diversity of the model.
It improves the counting accuracy in large population size changes and complex background scenarios, enhances the robustness of the model, and can predict the population number more accurately.
Smart Images

Figure CN115631454B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video image processing, and more specifically, to a method for counting the number of people in video images based on a multi-scale attention mechanism. Background Art
[0002] Crowd gatherings are likely to occur in public places. With the development and popularization of video surveillance technology, it is possible to use surveillance images to predict the number of people in the surveillance images, and issue warning messages or take evacuation measures when the number of people exceeds the warning threshold.
[0003] The application of crowd counting based on computer vision is in the fields of video surveillance, intelligent transportation, etc. Since the scale of people in the crowd images captured by outdoor surveillance varies greatly, and the background is complex, it greatly interferes with the accuracy of crowd counting. For the problem of crowd counting in outdoor scene images, the crowd density, the change of crowd scale, and the noise of the crowd background all affect the accuracy of crowd counting.
[0004] In some crowd scenes with complex backgrounds, leaves, vehicles, human faces in pictures, or other small targets far from the camera are often misjudged as human heads by the regression quantity. This phenomenon will significantly reduce the accuracy of the prediction results. Therefore, researchers need to remove these interference areas with complex backgrounds to avoid unnecessary counting. Image segmentation is a research hotspot in the field of image processing and also the basis of image analysis. With the improvement of segmentation technology, this technology can be used to clarify the foreground and background information when counting the number of people, so as to improve the accuracy.
[0005] However, the problem of large-scale variation of the crowd is a key issue that people focus on in crowd counting. Due to different shooting angles and the distance difference between the crowd and the camera, there will be a situation where the crowd sizes at different positions in the same image are different. If this scale variation is ignored during the experiment and filters with the same receptive field size are still used to process the pictures, the multi-scale features of the crowd cannot be accurately captured, and the counting accuracy and the quality of the density map will also decrease accordingly, which will further affect crowd analysis and other deeper applications.
[0006] Therefore, the accuracy of crowd counting in outdoor captured images still faces great challenges, attracting researchers to study and explore. Summary of the Invention
[0007] The purpose of the present invention is to solve the defect that it is difficult to ensure the accuracy of crowd counting under the interference of complex scenes in the prior art, and provide a method for counting the number of people in video images based on a multi-scale attention mechanism to solve the above problems.
[0008] To achieve the above purpose, the technical solution of the present invention is as follows:
[0009] A method for counting the number of people in video images based on a multi-scale attention mechanism, comprising the following steps:
[0010] Obtaining and preprocessing crowd images: Obtaining crowd images and preprocessing them to generate a training dataset;
[0011] Generating a true crowd density map: Generating a true crowd density map based on the preprocessed crowd images;
[0012] Constructing a multi-scale attention mechanism module: Constructing a multi-scale attention mechanism module for input feature maps and output scale diversity weight channel feature maps;
[0013] Constructing a crowd counting model: Constructing a crowd counting model based on the multi-scale attention mechanism module;
[0014] Training the crowd counting model: Corresponding the crowd images and the generated true crowd density maps to the input and output of the crowd counting model, and through the training of the neural network, fitting the parameters in the crowd counting model to make the similarity between the estimated crowd density map output by the crowd counting model and the true crowd density map reach the set requirements;
[0015] Obtaining the video images to be detected: Obtaining the video images to be detected and performing preprocessing;
[0016] Counting the number of people in the video images: Inputting the preprocessed video images to be detected into the trained crowd counting model to obtain a crowd prediction density map; Obtaining the number of people counted by integrating the crowd prediction density map, and the integration formula is shown as follows m and n respectively represent the length and width of the generated crowd density map, and P pixel (x i ,y i ) represents the pixel value at the pixel position (x i ,y i ) in the crowd density map, the interval size of the pixel value is [0, 1], and C count represents the obtained predicted number of people;
[0017] Finally, the crowd counting result of the video images is obtained.
[0018] The generation of the true crowd density map includes the following steps:
[0019] Obtaining a training dataset and performing head marking, and recording the coordinates x of the center of each head in the image i ;
[0020] Randomly mirroring the images in the training dataset with marked head center coordinates at a ratio of 0.5 and performing gamma contrast transformation at a ratio of 0.3;
[0021] Generate a real population density map using the geometric adaptive Gaussian convolution method, and its expression is as follows:
[0022]
[0023] Among them, x i represents the coordinates of the center position of the marked head, N is the total number of people in the whole image, δ(x - x i ) represents the impulse function, represents the Gaussian kernel function at pixel coordinate x, and the standard deviation σ i is calculated by multiplying the average distance of K nearest neighbors by a constant.
[0024] The construction of the multi-scale attention mechanism module includes the following steps:
[0025] Set the input of the multi-scale attention mechanism module as the feature map;
[0026] Perform convolutions with different convolution kernel sizes on the input feature map respectively to form four scale branches: the first scale branch is a 3×3 convolution, the second scale branch first performs a 5×5 convolution and fuses with the first branch, the third scale branch is a 7×7 convolution and fuses with the second scale branch, and the fourth scale branch is a 9×9 convolution and then fuses with the third scale branch;
[0027] Adjust the four scale branches to have equal number of channels through 1×1 convolution;
[0028] Generate descriptors of different channels of different scale branches for the feature maps output by the four different scale branches through global average pooling in the channel dimension,
[0029]
[0030] is the descriptor of the Cth channel of the Xth branch, where X and C represent the scale branch and the channel respectively, X ∈ {1, 2, 3, 4}, C ∈ {1, 2, 3,.., m}, H and W represent the height and width of the feature map respectively, represents the element value of the i-th row and j-th column of the Cth channel feature map of the Xth scale branch;
[0031] For all descriptors under all scale branches and channels First perform a fully connected layer, then activate it with the Relu function, activate the activation value with another fully connected layer, and activate it with the Sigmoid function to obtain the attention descriptor of the Cth channel of the Xth scale branch
[0032] The weights of the fully connected layer and the parameters of the two activation functions Relu and Sigmoid are iterated during training, and the iteration method adopts the Adam gradient descent method;
[0033] Normalize the attention descriptor After normalization, the attention descriptor of the C-th channel in the X-th scale branch is denoted as As shown in the formula:
[0034]
[0035] exp is the exponential function with the natural constant e as the base, which is the attention descriptor of the C-th channel in the X-th scale branch; is the attention descriptor of the C-th channel in the X-th scale branch, and m represents the number of channels;
[0036] Multiply the normalized attention descriptor as the weight by the feature map corresponding to the scale and channel, and fuse the weighted feature maps of each scale and channel as the output of the multi-scale attention mechanism module.
[0037] The construction of the crowd counting model includes the following steps:
[0038] Set the input of the crowd counting model as the training dataset;
[0039] Set the first part of the crowd counting model as:
[0040] In the crowd counting model, the training dataset is first convolved and pooled by the pre-trained VGG16 neural network to output the feature map of the image, with a size of 1 / 8 of the original input image, as the feature map output by the first part;
[0041] Set the second part of the crowd counting model as:
[0042] The feature map output by the first part is processed by a serial multi-scale attention module to output the second part of the feature map;
[0043] Set the third part of the crowd counting model as:
[0044] The second part of the feature map is further processed by a serial multi-scale attention module to output the third part of the feature map;
[0045] Set the fourth part of the crowd counting model as:
[0046] The third part of the feature map is further processed by a serial multi-scale attention module to output the fourth part of the feature map;
[0047] Set the fifth part of the crowd counting model to use dilated convolution and the corresponding activation function to regress and generate the crowd density map:
[0048] The fourth part inputs the dilated convolutional layer with feature maps. There are 4 layers of dilated convolutions, and the sizes of the convolutional kernels of each layer are 3×512, 3×256, 3×128, and 3×64 respectively. The dilation rate of the dilated convolution of each layer is 2, and the fifth part of the feature map is output;
[0049] Use a 1×1 convolutional layer to merge the channels of the fifth part of the feature map and regress to output a high-resolution crowd density map.
[0050] The training of the crowd counting model includes the following steps:
[0051] Input the crowd image into the crowd counting model. Among them, the parameters of the VGG16 network part in the crowd counting model adopt the parameters that have been trained on Imagenet and are not updated iteratively;
[0052] Set the loss function L of the crowd counting model to measure the error between the trained and fitted crowd density map and the true labeled crowd density map. The loss function L is defined as follows and is the weighted sum of the Euclidean loss and the multi-column variance loss.
[0053] L = L E + λL M
[0054] L E = ||G i - D i || 2
[0055]
[0056] where L E is the Euclidean loss, which is used to measure the pixel-level error between the estimated density map and the true density map. G i represents the estimated crowd density map output by the network, and D i represents the corresponding true density map. L M is the multi-column variance loss, whose purpose is to reduce the similarity of the features extracted by the multi-scale branch structure, force each branch to extract as different features as possible, and alleviate the problem of information redundancy in the features extracted by the multi-scale branch structure. λ is an empirical value and takes 1, 0.01, or 0.001;
[0057] y att_X_S ∈ R H×W represents the feature vector obtained by performing average pooling operation on the feature map y X output by the X-scale branch on the channel axis and then flattening. S represents the number of multi-scale attention modules, and y att_sum_S represents the sum of the y att_X_S vectors corresponding to each branch in the multi-scale attention module. ε is a fixed value used to avoid division by 0 and is set to 1×10 -6 ;
[0058] Iteratively update the convolution parameters, attention weights, and fully connected weights of the multi-scale attention mechanism module according to the Adam gradient descent method;
[0059] After each update, calculate the loss function of the crowd counting model with the updated parameters;
[0060] When the loss function is lower than the threshold or the number of iterations is greater than 1000, the iteration stops, and the parameters after iteration are the values of the undetermined parameters of the crowd number network.
[0061] Beneficial effects
[0062] A video image crowd counting method based on the multi-scale attention mechanism according to the present invention designs a multi-scale attention module compared with the prior art, and embeds the attention mechanism in different scale branches to reduce the influence of irrelevant background noise of the model at different scales, and at the same time increases the scale diversity of the model.
[0063] Based on the multi-scale attention module, the present invention proposes a new crowd counting network, which consists of the VGG-16 as the backbone network, three multi-scale attention modules in series, plus a regression network composed of dilated convolution and ordinary convolution. The loss function of the crowd counting network adopts the weighted sum of the multi-column variance loss and the Euclidean loss to limit the similarity of multi-scale features in the process of fitting the crowd density map.
[0064] The present invention can predict the number of people in an image. By inputting the image into the crowd counting network of the present invention, a predicted crowd density map is output, so as to obtain the number of people. It has good accuracy and robustness in processing crowd images with large scale changes and complex backgrounds, and improves the crowd counting accuracy in scenarios with large crowd scale changes and complex backgrounds. Description of the drawings
[0065] Figure 1 It is the sequence diagram of the method of the present invention;
[0066] Figure 2 It is the structure diagram of the multi-scale attention module involved in the present invention;
[0067] Figure 3 It is the structure diagram of the crowd counting network based on the multi-scale attention module involved in the present invention;
[0068] Figure 4a It is the original image on the UCF_QNRF dataset
[0069] Figure 4b It is the corresponding ground truth crowd density map on the UCF_QNRF dataset
[0070] Figure 4cThe crowd density map predicted by the CSRNet method
[0071] Figure 4d The crowd density map predicted by the present invention. Detailed implementation manners
[0072] To further understand and recognize the structural features and achieved effects of the present invention, the following is a detailed description in conjunction with preferred embodiments and accompanying drawings:
[0073] As Figure 1 shown, a video image crowd counting method based on a multi-scale attention mechanism according to the present invention includes the following steps:
[0074] First step, acquisition and preprocessing of crowd images: Acquire crowd images and preprocess them to generate a training data set.
[0075] Second step, generation of real crowd density map: Generate a real crowd density map according to the preprocessed crowd images.
[0076] (1) Obtain a training data set and perform head marking, and record the coordinate x of the head center for each head in the image i .
[0077] (2) For the images in the training data set with marked head center coordinates, perform random mirroring at a ratio of 0.5 and gamma contrast transformation at a ratio of 0.3.
[0078] (3) Use the method of geometric adaptive Gaussian convolution to generate a real crowd density map, and its expression is as follows:
[0079]
[0080] where x i represents the coordinate of the marked head center position, N is the total number of people in the whole image, δ(x - x i ) represents the impulse function, represents the Gaussian kernel function at pixel coordinate x, and the standard deviation σ i is calculated by multiplying the average distance of K nearest neighbors by a constant.
[0081] Third step, construct a multi-scale attention mechanism module: Construct a multi-scale attention mechanism module for input feature maps and output scale diversity weight channel feature maps.
[0082] As Figure 2 and Figure 3As shown in the figure, the present invention uses a multi-scale branch structure to extract information at different scales. Adjacent branches merge different multi-scale information in a scale fusion manner, and propose diverse features at different scales. The structure of the multi-scale branch similar to the residual network is beneficial to enhancing the network's ability to extract diverse scale features. An attention mechanism is added to different scale branches to emphasize the features closely related to the crowd information at different scales and reduce the influence of background features unrelated to the crowd. By increasing the weights of the attention mechanism for different scales and different channels, the crowd counting network model suppresses the influence of the background unrelated to the crowd from the perspective of scale and better fits the real crowd density map.
[0083] The specific steps are as follows:
[0084] (1) Set the input of the multi-scale attention mechanism module as the feature map.
[0085] (2) Perform convolutions with different kernel sizes on the input feature map respectively to form four scale branches: the first scale branch is a 3×3 convolution, the second scale branch first performs a 5×5 convolution and then fuses with the first branch, the third scale branch is a 7×7 convolution and fuses with the second scale branch, and the fourth scale branch is a 9×9 convolution and then fuses with the third scale branch.
[0086] (3) Adjust the four scale branches to have equal number of channels through 1×1 convolutions.
[0087] (4) Generate descriptors for different channels of different scale branches by performing global average pooling on the feature maps output by the four different scale branches in the channel dimension.
[0088]
[0089] is the descriptor for the C-th channel of the X-th branch, where X and C represent the scale branch and the channel respectively, X ∈ {1, 2, 3, 4}, C ∈ {1, 2, 3,.., m}, and H and W represent the height and width of the feature map respectively. represents the element value at the i-th row and j-th column of the feature map of the C-th channel of the X-th scale branch.
[0090] (4) For all descriptors under all scale branches and channels First perform a fully connected layer, then activate it with the Relu function, activate the activation value with another fully connected layer, and activate it with the Sigmoid function to obtain the attention descriptor for the C-th channel of the X-th scale branch.
[0091] The weights of the fully connected layer and the parameters of the two activation functions Relu and Sigmoid are iterated during training, and the iteration method uses the Adam gradient descent method.
[0092] (5) Normalize the attention descriptor After normalization, the attention descriptor of the C-th channel of the X-th scale branch is denoted as As shown in the formula:
[0093]
[0094] exp is the exponential function with the natural constant e as the base, is the attention descriptor of the C-th channel of the X-th scale branch; is the attention descriptor of the C-th channel of the X-th scale branch, and m represents the number of channels.
[0095] (6) Multiply the normalized attention descriptor as the weight by the feature map corresponding to the scale and channel, and fuse the weighted feature maps of each scale and channel as the output of the multi-scale attention mechanism module.
[0096] Fourth step, construction of the crowd counting model: Construct a crowd counting model based on the multi-scale attention mechanism module. This model extracts general features of the image through the general VGG16 network, and then concatenates three multi-scale attention modules designed by the present invention. A high-quality and robust crowd density map is generated through the regression network. This crowd counting model can well handle the interference of complex backgrounds and inconsistent crowd scales, and has good crowd counting accuracy and robustness.
[0097] (1) Set the input of the crowd counting model as the training data set.
[0098] (2) Set the first part of the crowd counting model as:
[0099] In the crowd counting model, the training data set is first convolved and pooled by the pre-trained VGG16 neural network, and the feature map of the output image, with a size of 1 / 8 of the original input image, is used as the feature map output in the first part.
[0100] (3) Set the second part of the crowd counting model as:
[0101] The feature map output in the first part is processed by a serial multi-scale attention module, and the feature map of the second part is output.
[0102] (4) Set the third part of the crowd counting model as:
[0103] The feature map of the second part is further processed by a serial multi-scale attention module, and the feature map of the third part is output.
[0104] (5) Set the fourth part of the crowd counting model as:
[0105] The third part of the feature map is further processed by a serial multi-scale attention module to output the fourth part of the feature map.
[0106] (6) Set the fifth part of the crowd counting model to generate a crowd density map by regression using dilated convolution and the corresponding activation function:
[0107] A1) The fourth part of the feature map is input into the dilated convolution layer. The dilated convolution has 4 layers, and the sizes of the convolution kernels of each layer are 3×512, 3×256, 3×128, and 3×64 respectively. The dilation rate of the dilated convolution of each layer is 2, and the fifth part of the feature map is output;
[0108] A2) Use a 1×1 convolution layer to merge each channel of the fifth part of the feature map and regress to output a high-resolution crowd density map.
[0109] Step 5, training of the crowd counting model: Corresponding the crowd image and the generated true crowd density map to the input and output of the crowd counting model, and through the training of the neural network, fitting the parameters in the crowd counting model to make the similarity between the estimated crowd density map output by the crowd counting model and the true crowd density map reach the set requirements.
[0110] (1) Input the crowd image into the crowd counting model. Among them, the parameters of the VGG16 network part in the crowd counting model adopt the parameters that have been trained on Imagenet and are not updated iteratively.
[0111] (2) Set the loss function L of the crowd counting model to measure the error between the trained and fitted crowd density map and the truly labeled crowd density map. The loss function L is defined as follows and is the weighted sum of the Euclidean loss and the multi-column variance loss.
[0112] L = L E + λL M
[0113] L E =||G i - D i || 2
[0114]
[0115] where L E is the Euclidean loss, used to measure the pixel-level error between the estimated density map and the true density map. G i represents the estimated crowd density map output by the network, D i represents the corresponding true density map, and L MIt is the multi-column variance loss, aiming to reduce the feature similarity extracted by the multi-scale branch structure, forcing each branch to extract as different features as possible, and alleviating the problem of information redundancy in the multi-scale branch structure. λ is an empirical value, taking 1, 0.01 or 0.001;
[0116] y att_X_S ∈R H×W is expressed as the feature vector obtained by performing average pooling operation on the output feature map y of the X-scale branch and then flattening it. S represents the number of multi-scale attention modules, and y X represents the sum of the y vectors corresponding to each branch in the multi-scale attention module. ε is a fixed value used to avoid division by zero and is set to 1×10 att_sum_S represents the sum of the y vectors corresponding to each branch in the multi-scale attention module. ε is a fixed value used to avoid division by zero and is set to 1×10 att_X_S . -6 .
[0117] (3) Iterate the convolution parameters, attention weights, and fully connected weights of the multi-scale attention mechanism module according to the Adam gradient descent method.
[0118] (4) Update the parameters of the crowd counting model each time and calculate its loss function.
[0119] (5) When the loss function is lower than the threshold or the number of iterations is greater than 1000 times, the iteration stops, and the iterated parameters are the values of the undetermined parameters of the crowd number network.
[0120] The sixth step, obtaining the video image to be detected: Obtain the video image to be detected and perform preprocessing.
[0121] The seventh step, counting the people in the video image: Input the preprocessed video image to be detected into the trained crowd counting model to obtain the crowd prediction density map; Integrate the crowd prediction density map to obtain the crowd count, and the integration formula is shown as follows m and n respectively represent the length and width of the generated crowd density map, and P pixel (x i ,y i ) represents the pixel value at the pixel position (x i ,y i ) in the crowd density map. The interval size of the pixel value is [0, 1], and C count represents the obtained predicted number of people;
[0122] Finally, the crowd counting result of the video image is obtained.
[0123] Such as Figure 4a , Figure 4b , Figure 4c , Figure 4dAs shown, the comparison between the partial population density maps predicted by the present invention in the UCF_QNRF dataset and the population density maps predicted by the CSRNet method. The original images are three randomly selected images from the UCF_QNRF dataset, where the scenes in the images are complex, the people are crowded and overlapping with each other, and the scale of the people varies greatly. From the comparison, the quality of the population density maps predicted by the method of the present invention is higher than that of the population density maps predicted by the CSRNet method, and the counting accuracy is also better.
[0124] The present invention achieves good accuracy in the public dataset for population counting, superior to the classical population counting methods. The following are the experimental results in the ShanghaiTech A and ShanghaiTech B population counting datasets.
[0125] Table 1 Experimental result table of ShanghaiTech A dataset
[0126]
[0127] Table 2 Experimental result table of ShanghaiTech B dataset
[0128]
[0129] From Table 1 and Table 2, the mean absolute error and mean square error of the population counting results are the most common evaluation criteria for evaluating population counting methods. It can be seen that in ShanghaiTech A and ShanghaiTech B, both the mean absolute error and mean square error of the population counting results of the present invention are smaller than those of the existing advanced population counting methods, indicating that the population counting accuracy of the present method is higher than the above methods.
[0130] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification is only the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection required by the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for counting the number of people in video images based on a multi-scale attention mechanism, characterized in that It includes the following steps: 11) Acquisition and preprocessing of crowd images: Acquire crowd images and preprocess them to generate a training dataset; 12) Generation of real crowd density maps: Generate real crowd density maps based on the preprocessed crowd images; 13) Construction of a multi-scale attention mechanism module: Construct a multi-scale attention mechanism module that inputs a feature map and outputs a feature map with diverse scale weights in the channels; The construction of the multi-scale attention mechanism module includes the following steps: 131) Set the input of the multi-scale attention mechanism module as a feature map; 132) Perform convolutions on the input feature map with different convolution kernel sizes to form four scale branches: The first scale branch is a 3×3 convolution, the second scale branch is a 5×5 convolution fused with the first branch, the third scale branch is a 7×7 convolution fused with the second scale branch, and the fourth scale branch is a 9×9 convolution fused with the third scale branch; 133) Adjust the four scale branches to have equal numbers of channels through 1×1 convolutions; 134) Generate descriptors for different channels of different scale branches by performing global average pooling on the feature maps output from the four different scale branches in the channel dimension, Descriptor for the C-th channel of the X-th branch, where X and C represent the scale branch and the channel respectively, X ∈ {1, 2, 3, 4}, C ∈ {1, 2, 3,.., m}, and H and W represent the height and width of the feature map respectively. Represents the element value at the i-th row and j-th column of the feature map of the C-th channel of the X-th scale branch; 134) For the descriptors under all scale branches and channels First, perform a fully connected operation, then activate it with the Relu function. The activation value is then subjected to another fully connected operation and activated with the Sigmoid function to obtain the attention descriptor for the C-th channel of the X-th scale branch The weights of the fully connected layer and the parameters of the two activation functions Relu and Sigmoid are iterated during training, and the iteration method uses the Adam gradient descent method; 135) Normalize the attention descriptor After normalization, the attention descriptor of the C-th channel of the X-th scale branch is expressed as As shown in the formula: exp is the exponential function with the natural constant e as the base and is the attention descriptor for the C-th channel of the X-th scale branch; It is the attention descriptor for the C-th channel of the X-th scale branch, and m represents the number of channels; 136) Multiply the normalized attention descriptor by the feature map corresponding to the scale and channel as the weight point, and fuse the feature maps of each scale and channel after weighting as the output of the multi-scale attention mechanism module; 14) Construction of a crowd counting model: Construct a crowd counting model based on the multi-scale attention mechanism module; The construction of the crowd counting model includes the following steps: 141) Set the input of the crowd counting model as the training dataset; 142) Set the first part of the crowd counting model as: In the crowd counting model, the training dataset first undergoes convolution and pooling using the pre-trained VGG16 neural network to output a feature map of the image, with a size of 1 / 8 of the original input image, as the feature map output by the first part; 143) Set the second part of the crowd counting model as: The feature map output by the first part is processed through a serial multi-scale attention module to output the second part of the feature map; 144) Set the third part of the crowd counting model as: The second part of the feature map is further processed through a serial multi-scale attention module to output the third part of the feature map; 145) Set the fourth part of the crowd counting model as: The third part of the feature map is further processed through a serial multi-scale attention module to output the fourth part of the feature map; 146) Set the fifth part of the crowd counting model to use dilated convolution and the corresponding activation function to regress and generate a crowd density map: 1461) The fourth part of the feature map is input into the dilated convolution layer. This dilated convolution has 4 layers, and the sizes of the convolution kernels for each layer are 3×512, 3×256, 3×128, and 3×64 respectively. The dilation rate of the dilated convolution for each layer is 2, and the fifth part of the feature map is output; 1462) Use a 1×1 convolution layer to merge the channels of the fifth part of the feature map and regress to output a high-resolution crowd density map; 15) Training of the crowd counting model: Corresponding the crowd image and the generated real crowd density map to the input and output of the crowd counting model. Through the training of the neural network, the parameters in the crowd counting model are fitted so that the similarity between the estimated crowd density map output by the crowd counting model and the real crowd density map reaches the set requirements; 16) Acquisition of the video image to be detected: Acquire the video image to be detected and perform preprocessing; 17) Counting the people in the video image: Input the preprocessed video image to be detected into the trained people counting model to obtain the people prediction density map; Integrate the people prediction density map to get the people count, and the integration formula is shown as follows m and n respectively represent the length and width of the generated people density map, P pixel (x i , y i ) represents the pixel value at the pixel position (x i , y i ) in the people density map, the interval size of the pixel value is [0, 1], and C count represents the obtained predicted number of people; Finally, the crowd counting result of the video image is obtained.
2. The method for counting the number of people in a video image based on a multi-scale attention mechanism according to claim 1, wherein The generation of the real crowd density map includes the following steps: 21) Obtain a training data set and perform head marking, and record the coordinate x of the head center for each head on the image i ; 22) For the images in the training dataset with the center coordinates of the labeled human heads, perform random mirroring at a ratio of 0.5 and gamma contrast transformation at a ratio of 0.3; 23) Use the method of geometric adaptive Gaussian convolution to generate the real crowd density map, and its expression is as follows: Among them, x i represents the coordinate of the center position of the marked head. N is the total number of people in the whole image. δ(x - x i ) represents the impulse function, represents the Gaussian kernel function at pixel coordinate x, and the standard deviation σ i is calculated by multiplying the average distance of K nearest neighbors by a constant.
3. The method for counting the number of people in video images based on a multi-scale attention mechanism according to claim 1, wherein, The training of the crowd counting model includes the following steps: 31) Input the crowd image into the crowd counting model. Among them, the parameters of the VGG16 network part in the crowd counting model adopt the parameters that have been trained on Imagenet and are not updated iteratively; 32) Set the loss function L of the crowd counting model to measure the error between the trained and fitted crowd density map and the real labeled crowd density map. The loss function L is defined as follows and is the weighted sum of the Euclidean loss and the multi-column variance loss. Among them, L E is the Euclidean loss, which is used to measure the pixel-level error between the estimated density map and the ground-truth density map. G i represents the estimated crowd density map output by the network, and D i represents the corresponding ground-truth density map. L M is the multi-column variance loss, which aims to reduce the feature similarity extracted by the multi-scale branch structure, forcing each branch to extract as different features as possible and alleviating the problem of information redundancy in the multi-scale branch structure. λ is an empirical value, taking 1, 0.01 or 0.001; y att_X_S ∈R H×W is expressed as the feature map y output from the X-scale branch X which is the feature vector obtained by performing average pooling operation on the channel axis and then flattening. S represents the number of multi-scale attention modules, and y att_sum_S represents the y corresponding to each branch in the multi-scale attention module att_X_S which is the sum of vectors. ε is a fixed value used to avoid division by zero and is set to 1×10 -6 ; 33) Iterate the parameters of the convolution parameters, attention weights, and fully connected weights of the multi-scale attention mechanism module according to the Adam gradient descent method; 34) After each update of the parameters of the crowd counting model, calculate its loss function; 35) When the loss function is lower than the threshold or the number of iterations is greater than 1000 times, the iteration stops, and the iterated parameters are the values of the undetermined parameters of the crowd counting network.
Citation Information
Patent Citations
Image processing method for crowd counting, apparatus, device, and storage medium
JP2022052714A
Person re-identification method combining reverse attention and multi-scale deep supervision
US20210232813A1