A method and system for estimating crowd density based on visual converter

By adopting a neural network model based on vision converter in population density estimation, combining convolutional neural network and vision converter, the problem of low accuracy of existing methods is solved, higher accuracy density estimation is achieved, and the model construction and training process is simplified.

CN114519844BActive Publication Date: 2025-05-13FUDAN UNIVERSITY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210133298.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-09
Publication Date
2025-05-13
Estimated Expiration
2042-02-09

AI Technical Summary

Technical Problem

The detection accuracy of existing population density estimation methods is low and cannot be applied to practical general estimation tasks. Increasing the amount of training data will lead to a longer training time of the model and may even be unable to be actually completed.

Method used

The population density estimation method based on vision converter is adopted, and the population density estimation is estimated by building a neural network model including a front-end convolutional neural network and a back-end vision converter, combining the multi-level convolution mechanism and the vision converter to extract local and global information.

Benefits of technology

The accuracy of population density estimation is improved, suitable for high-density population scenarios, the model structure is simple, the construction is fast, and the training and calculation is small.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114519844B_ABST
    Figure CN114519844B_ABST
Patent Text Reader

Abstract

The present invention discloses a crowd density estimation method based on a visual converter, which has the following characteristics, including: step 1, pre-processing the image to be tested to obtain a pre-processed image, and then building a coding and decoding layer; step 2, building a neural network model based on the visual converter; step 3, inputting the training data into the neural network model based on the visual converter for model training, and obtaining a trained neural network model based on the visual converter mechanism; step 4, inputting the pre-processed image into the trained neural network model based on the visual converter mechanism, and respectively obtaining the crowd density results in each pre-processed image and outputting them, wherein the neural network model based on the visual converter includes a local information extraction module based on a convolutional neural network at the front end and a global information extraction module based on the visual converter at the back end. The present invention also discloses a crowd density estimation system based on a visual converter, including a pre-processing unit and a density prediction unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer vision and artificial intelligence, and in particular to a crowd density estimation method and system based on a visual converter. Background Art

[0002] With the rapid improvement of machine learning technology and computer hardware performance, breakthroughs have been made in application fields such as computer vision, natural language processing, and speech detection in recent years. As a basic task in the field of computer vision, the accuracy of crowd density estimation has also been greatly improved.

[0003] The crowd density estimation task can be specifically described as follows:

[0004] For the captured pictures or recorded videos and crowd scenes captured by the camera, a density map is generated to represent the density of the crowd per unit area. Based on the density map, the density of the crowd per unit area in the density map is summed up to obtain the final crowd density of the overall scene or the change in crowd density of the entire video.

[0005] Crowd density estimation is of great significance to the field of computer vision and practical applications. In the past few decades, it has inspired a large number of researchers to pay close attention to and invest in research. With the development of powerful machine learning theory and feature analysis technology, research activities related to crowd density estimation have increased in the past decade, and the latest research results and practical applications are published and announced every year. In addition, crowd density estimation has also been applied to many practical tasks, such as intelligent video surveillance, crowd situation analysis, etc. However, the detection accuracy of various crowd density estimation methods in the existing technology is still low and cannot be applied to practical general estimation tasks. Therefore, crowd density estimation is far from being perfectly solved and is still an important and challenging research topic.

[0006] In order to improve the accuracy of density estimation, the commonly used method is to increase the training data when training the prediction model. However, on the one hand, collecting a large amount of training data is an extremely difficult task. On the other hand, the increase in the amount of training data also leads to a longer model training time, and it is even possible that the training cannot be actually completed. Summary of the invention

[0007] The present invention is made to solve the above-mentioned problem, and aims to provide a crowd density estimation method and system based on visual converter.

[0008] The present invention provides a crowd density estimation method based on a visual converter, which has the following characteristics and comprises the following steps: step 1, preprocessing an image to be tested to obtain a preprocessed image, and then building a coding and decoding layer; step 2, building a neural network model based on a visual converter; step 3, inputting training data into the neural network model based on the visual converter for model training, and obtaining a trained neural network model based on the visual converter mechanism; step 4, inputting the preprocessed image into the trained neural network model based on the visual converter mechanism, and respectively obtaining and outputting crowd density results in each preprocessed image, wherein the neural network model based on the visual converter comprises a local information extraction module based on a convolutional neural network at the front end and a global information extraction module based on the visual converter at the back end.

[0009] The crowd density estimation method based on visual transformer provided by the present invention may also have the following features: wherein the neural network model based on visual transformer also introduces a multi-level convolution mechanism, and each level of convolution operation is followed by a visual transformer to extract global connections.

[0010] The crowd density estimation method based on the visual converter provided by the present invention may also have the following feature: wherein the image to be measured is a high-density crowd image.

[0011] In the crowd density estimation method based on visual converter provided by the present invention, it can also have the following characteristics: wherein, in step 1, the preprocessing step is: step 1-1, image segmentation is performed from the image to be tested; step 1-2, regularization is performed on the segmented image, and the image segmentation method is to scale the image to be tested according to a certain scale, randomly select a point in the image to be tested as the center point, and cut the image to be tested at a certain ratio.

[0012] The crowd density estimation method based on visual transformer provided by the present invention may also have the following features: wherein, in step 1, the encoding and decoding layer is based on the VGG-16 network and the visual transformer network.

[0013] The crowd density estimation method based on the visual converter provided by the present invention may also have the following features: wherein, in step 2, when building the neural network model of the visual converter, a visual transformer structure is added after the front-end multi-level convolution layer.

[0014] In the crowd density estimation method based on visual converter provided by the present invention, it can also have the following characteristics: wherein step 3 includes the following steps: step 3-1, constructing a neural network model based on visual converter, wherein the model optimizer included is stochastic gradient descent, and the learning rate is 10 -7; Step 3-2, input each training image in the training set into the visual converter-based neural network model in turn and perform one iteration; Step 3-3, after the iteration, use the model parameters of the last layer to calculate the loss error respectively, and then back-propagate the calculated loss error to update the model parameters; Step 3-4, repeat steps 3-2 to 3-3 until the training completion conditions are met, and obtain the trained visual converter-based neural network model.

[0015] The present invention provides a crowd density estimation system based on a visual converter, which has the following characteristics: using a neural network model based on a visual converter to detect crowd density from an image to be tested, including: a preprocessing unit, preprocessing the image to be tested to obtain a preprocessed image; a density prediction unit, building a neural network model of the visual converter, inputting training data into the built neural network model based on the visual converter, thereby performing model training, and inputting the preprocessed image into the trained neural network model based on the visual converter mechanism, thereby obtaining and outputting crowd density results in each preprocessed image, wherein the neural network model based on the visual converter includes a local information extraction module based on a convolutional neural network at the front end and a global information extraction module based on the visual converter at the back end.

[0016] Functions and Effects of the Invention

[0017] According to the crowd density estimation method based on visual converter involved in the present invention, the estimation steps are as follows: step 1, preprocessing the image to be tested to obtain a preprocessed image, and then building a coding and decoding layer; step 2, building a neural network model based on the visual converter; step 3, inputting the training data into the neural network model based on the visual converter for model training, and obtaining a trained neural network model based on the visual converter mechanism; step 4, inputting the preprocessed image into the trained neural network model based on the visual converter mechanism, and respectively obtaining the crowd density results in each preprocessed image and outputting them, wherein the neural network model based on the visual converter includes a local information extraction module based on a convolutional neural network at the front end and a global information extraction module based on the visual converter at the back end.

[0018] Therefore, according to the crowd density estimation method based on visual converter provided by the present invention, since a neural network model combining a multi-level convolutional fusion network and a visual converter is introduced as a prediction model, the hybrid neural network structure of the prediction model can enable the model to better locate the crowd and identify the density of the crowd. Therefore, this model can learn more features and express features better, which is more suitable for the crowd density estimation task of high-density crowds and can ultimately improve the accuracy of crowd density estimation.

[0019] In addition, the model structure is simple and does not require the use of model mixing, multi-task training, and metric learning methods. Therefore, compared with existing high-precision models, the model of the present invention is fast and convenient to build, and the amount of computation consumed in the training process is also small. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 It is a flow chart of a crowd density estimation method based on a visual converter in an embodiment of the present invention.

[0021] Figure 2 It is a structural schematic diagram of a crowd density estimation method based on a visual converter in an embodiment of the present invention.

[0022] Figure 3 It is a structural diagram of the visual converter module in an embodiment of the present invention.

[0023] Figure 4 It is a structural diagram of multi-level fused convolution in an embodiment of the present invention. DETAILED DESCRIPTION

[0024] In order to make the technical means, creative features, objectives and effects achieved by the present invention easy to understand, the following embodiments and accompanying drawings specifically illustrate a crowd density estimation method and system based on a visual converter of the present invention.

[0025] In this embodiment, a crowd density estimation method based on a visual converter is provided.

[0026] The dataset that can be used in this embodiment is UCF-QNRF. UCF-QNRF is a challenging high-density scene dataset, which contains 1535 labeled images and annotation information of a total of 1251642 heads. The number of people in each image in the dataset ranges from 49 to 12865, and the average resolution of the images in the dataset is 2013 times 2902, which is a high-resolution dataset.

[0027] The dataset that can also be used in this embodiment is ShangehaiTech. ShangehaiTech is a challenging high-density scene dataset, which contains 1,198 labeled images and annotation information of a total of 330,165 heads. This dataset consists of two parts: Part A and Part B. Part A includes 482 pictures randomly downloaded from the Internet. The train group has 300 images and the test group has 182 images. Part B has a total of 716 images of different scenes taken by drones in Shanghai.

[0028] The dataset that can also be used in this embodiment is UCF CC 50. UCF CC 50 is a challenging high-density scene dataset, which contains 50 annotated images of extremely dense crowds. The number of people annotated in each image ranges from 94 to 4543, with an average of 1280. In addition, the hardware platform implemented in this embodiment requires an NVIDIA 1080Ti graphics card (GPU acceleration).

[0029] This embodiment first preprocesses the data set images, then trains a convolutional neural network model based on the attention mechanism, and finally obtains the crowd density of the image through the neural network model. Specifically, it includes four processes: preprocessing, model building, model training, and density prediction.

[0030] Figure 1 4 is a flow chart of a crowd density estimation method based on a visual converter in an embodiment of the present invention.

[0031] like Figure 1 As shown, a crowd density estimation method based on a visual converter according to this embodiment includes the following steps:

[0032] Step S1, preprocessing the image to be tested to obtain a preprocessed image, and then building a coding and decoding layer.

[0033] In this embodiment, the image to be tested is an image obtained from the UCF-QNRF dataset. Due to the high score ratio of the dataset, the image cannot be directly input into the model. Downsampling is performed, and a random center point is selected to divide the image into multiple sub-blocks, which are then input into the model.

[0034] In the above process of this embodiment, downsampling is a measure for the high resolution of the image; selecting a random center point to divide multiple sub-block training data is to increase the number of images obtained, realize data expansion so that the amount of data obtained from the image to be tested is richer, and then increase the epoch of iteration. In other embodiments, the image to be tested can also be a single image (such as a photo, etc.), in which case no segmentation operation is required. In addition, in other embodiments, the image may not be copied, or other data expansion methods in the prior art may be used (such as vertical flipping, horizontal flipping combined with vertical flipping, etc.).

[0035] Construct the encoding and decoding layers. The encoding layer is based on the first ten layers of VGG-16. Each of the last four layers consists of a maximum pooling and two dilated convolutions. The kernel size of each layer is 3x3, and the dilation factor is 2. The output channels of the third layer are 512, the output channels of the fourth layer are 512 and 256 respectively, and the output channels of the fifth layer are 128 and 64. In order to adjust the output features of each layer, a convolution layer is connected to the 3rd, 4th, 5th, and 6th layers respectively, and the output channels of these four convolution layers are all 128.

[0036] Step S2, building a neural network model based on the visual converter.

[0037] First, we use the existing deep learning framework PyTorch to build a neural network model based on a visual converter. The neural network model based on a visual converter introduces a neural network that combines multi-level fused convolution and a visual converter.

[0038] Specifically, the model of this embodiment consists of a convolutional local information extractor based on multi-level fusion and a global information extractor of transformer, wherein in the multi-level fusion convolutional local information module, each convolution operation of the convolutional layer is followed by a visual converter.

[0039] Figure 2 It is a structural diagram of a neural network model of a visual converter according to an embodiment of the present invention.

[0040] like Figure 2 As shown, the visual converter neural network model of the present invention includes an input layer I, a multi-level fusion convolution layer based on the VGG-16 front end, and a visual transformer module which are arranged in sequence.

[0041] like Figure 2 As shown, the neural network model based on the visual extractor specifically includes the following structure:

[0042] (1) Input layer I, used to input each preprocessed and encoded image.

[0043] (2) The first two layers of the CNN part of the network use the first 10 layers of the pre-trained VGG16 as the backbone. The last four layers each consist of a maximum pooling and two dilated convolutions. The kernel size of each layer is 3x3 and the dilation factor is 2. The output channels of the third layer are 512, the output channels of the fourth layer are 512 and 256 respectively, and the output channels of the fifth layer are 128 and 64. In order to adjust the output features of each layer, a convolution layer is connected to each of the 3rd, 4th, 5th, and 6th layers. The output channels of these four convolution layers are all 128.

[0044] (3) Visual transformer module: The output of multi-level convolution is fed into this module. Finally, all the results are merged into the final result graph.

[0045] Figure 3 It is a structural diagram of the visual converter module of an embodiment of the present invention.

[0046] like Figure 3 As shown in the figure, in the visual converter neural network model, channel adjustment operation is performed after each multi-level fusion convolution layer.

[0047] Figure 4 It is a structural diagram of a multi-stage fused convolution according to an embodiment of the present invention.

[0048] like Figure 4 As shown in the figure, each layer of multi-level convolution is followed by an adjustment layer, which adjusts the output feature map to the size and scale required by the back-end visual converter.

[0049] like Figure 4 As shown in the figure, firstly, a sub-image of the center point of the input image is randomly selected, and then the sub-image and the original image are compared. Figure 1 Then, they are fed into the model for training.

[0050] Step S3, input the training data into the neural network model based on the visual converter to perform model training, and obtain a trained neural network model based on the visual converter mechanism.

[0051] This embodiment uses the crowd dataset UCF-QNRF as training data. Using the same method as step S1, 1535 images containing 1251642 human heads are obtained from the dataset; these images are segmented to achieve data enhancement, and then regularized, and the multiple images obtained are the training set of this embodiment.

[0052] The images in the above training set enter the network model in batches for training. The batch size of the training images entering the network model each time is 1, and the training is iterated 2000 times in total.

[0053] Wherein, step S3 comprises the following steps:

[0054] Step S3-1, construct a neural network model based on the visual converter, which includes a model optimizer of stochastic gradient descent with a learning rate of 10 -7 .

[0055] Step S3-2, each training image in the training set is sequentially input into the visual converter-based neural network model and an iteration is performed.

[0056] Step S3-3, after iteration, the model parameters of the last layer are used to calculate the loss error respectively, and then the calculated loss error is back-propagated to update the model parameters.

[0057] Step S3-4, repeat steps S3-2 to S3-3 until the training completion condition is met, and obtain the trained visual converter-based neural network model.

[0058] Step S4, input the pre-processed image into the trained neural network model based on the visual converter mechanism, obtain the crowd density results in each pre-processed image and output them.

[0059] Among them, the neural network model based on the visual converter includes a local information extraction module based on the convolutional neural network at the front end and a global information extraction module based on the visual converter at the back end.

[0060] In this embodiment, the UCF-QNRF test set is used as the test image to test the model of this embodiment, wherein the scene is a high-density crowd scene.

[0061] The specific process is as follows: using the UCF-QNRF dataset, multiple images in its dataset are preprocessed as described in step S1 to obtain 334 images (i.e., preprocessed images after preprocessing) as the test set, which are sequentially input into the trained convolutional neural network model based on the attention mechanism to generate the corresponding density map and calculate the crowd density result.

[0062] In this embodiment, the trained visual converter-based neural network model generates a mean absolute error (MAE) of 99.7 and a mean square error (MSE) of 172.3 for the crowd density estimation of the test set.

[0063] In this embodiment, other crowd density estimation models in the prior art are also used to conduct comparative tests on the same test set, and the results are shown in Table 1 below.

[0064] Table 1 shows the comparative test results of the method of this embodiment and other methods in the prior art for estimating crowd density on the UCF-QNRF dataset.

[0065] Table 1

[0066]

[0067] In Table 1, MCNN, CP-CNN, TDF-CNN, ic-CNN, D-ConvNet, and CSRNet are several models with high accuracy in crowd density estimation in the existing technology. In addition, MAE stands for mean absolute error and MSE stands for mean square error.

[0068] The above test process shows that the crowd density estimation method based on the visual converter neural network model of this embodiment can achieve a very high accuracy on the UCF-QNRF dataset.

[0069] This embodiment also provides a crowd density estimation system based on a visual converter, which uses a neural network model based on a visual converter to detect crowd density from a test image, including:

[0070] The preprocessing section performs preprocessing using the method in step S1 in this embodiment.

[0071] The density prediction unit performs density prediction using the method in steps S1 to S4 in this embodiment to obtain a prediction result.

[0072] Functions and Effects of the Embodiments

[0073] According to the crowd density estimation method and system based on visual converter involved in this embodiment, because, step 1, preprocess the image to be tested to obtain the preprocessed image, and then build the encoding and decoding layer; step 2, build the neural network model based on the visual converter; step 3, input the training data into the neural network model based on the visual converter for model training, and obtain the trained neural network model based on the visual converter mechanism; step 4, input the preprocessed image into the trained neural network model based on the visual converter mechanism, and respectively obtain the crowd density results in each preprocessed image and output them, wherein the neural network model based on the visual converter includes a local information extraction module based on the convolutional neural network at the front end and a global information extraction module based on the visual converter at the back end.

[0074] Therefore, according to the crowd density estimation method and system based on visual converter provided by the above-mentioned embodiments, since a neural network model combining a multi-level convolutional fusion network and a visual converter is introduced as a prediction model, the hybrid neural network structure of the prediction model can enable the model to better locate the crowd and identify the density of the crowd. Therefore, this model can learn more features and express features better, which is more suitable for the crowd density estimation task of high-density crowds and can ultimately improve the accuracy of crowd density estimation.

[0075] In addition, the model structure is simple and does not require the use of methods such as model mixing, multi-task training, and metric learning. Therefore, compared with existing high-precision models, the model construction of the above embodiment is fast and convenient, and the amount of computation consumed in the training process is also small.

[0076] The above-mentioned embodiments are preferred examples of the present invention and are not intended to limit the protection scope of the present invention.

Claims

1. A crowd density estimation method based on visual converter, characterized in that: The crowd density is detected from the image to be tested by using a visual converter and a multi-level convolution combined with a neural network model, including the following steps: Step 1: preprocess the image to be tested to obtain a preprocessed image, and then build the encoding and decoding layer; Step 2: Build a neural network model based on the visual converter; Step 3, inputting the training data into the neural network model based on the visual converter to perform model training, thereby obtaining a trained neural network model based on the visual converter mechanism; Step 4, inputting the pre-processed image into the trained neural network model based on the visual converter mechanism, respectively obtaining the crowd density results in each of the pre-processed images and outputting them, The neural network model based on the visual converter includes a local information extraction module based on a convolutional neural network at the front end and a global information extraction module based on a visual converter at the back end. In step 1, the encoding and decoding layers are based on the VGG-16 network and the visual transformer network. The first two layers of the CNN part of the network use the first 10 layers of the pre-trained VGG16 as the backbone. The last four layers each consist of a maximum pooling and two dilated convolutions. The kernel size of each layer is 3x3, and the dilation factor is 2. The output channels of the third layer are 512, the output channels of the fourth layer are 512 and 256 respectively, and the output channels of the fifth layer are 128 and 64. In order to adjust the output features of each layer, a convolution layer is connected to the 3rd, 4th, 5th, and 6th layers respectively. The output channels of these four convolution layers are all 128. The output of the multi-level convolution is fed into the visual transformer module, and all the results are combined into the final result graph.

2. The crowd density estimation method based on visual converter according to claim 1 is characterized in that: in, The visual transformer-based neural network model also introduces a multi-level convolution mechanism, where each level of convolution operation is followed by a visual transformer to extract global connections.

3. The crowd density estimation method based on visual converter according to claim 1, characterized in that: in, The image to be tested is a high-density crowd image.

4. The crowd density estimation method based on visual converter according to claim 1, characterized in that: in, In step 1, the pre-processing steps are: Step 1-1, performing image segmentation from the image to be tested; Step 1-2, regularize the segmented image. The image segmentation method is to scale the image to be tested according to a certain scale, randomly select a point in the image to be tested as a center point, and cut the image to be tested at a certain ratio.

5. The crowd density estimation method based on visual converter according to claim 1, characterized in that: in, In step 2, when building the neural network model of the visual transformer, a visual transformer structure is added after the front-end multi-level convolution layer.

6. The crowd density estimation method based on visual converter according to claim 1, characterized in that: in, Step 3 includes the following steps: Step 3-1, construct the neural network model based on the visual converter, the model optimizer included is stochastic gradient descent, and the learning rate is 10 -7 ; Step 3-2, inputting each training image in the training set into the visual converter-based neural network model in sequence and performing one iteration; Step 3-3, after the iteration, using the model parameters of the last layer to respectively calculate the loss error, and then back-propagating the calculated loss error to update the model parameters; Step 3-4, repeat steps 3-2 to 3-3 until the training completion condition is met, and obtain the trained visual converter-based neural network model.

7. A crowd density estimation system based on visual converter, characterized in that: A neural network model based on visual converter is used to detect crowd density from the image to be tested, including: The preprocessing unit preprocesses the image to be tested to obtain a preprocessed image, and then builds a coding and decoding layer. The coding and decoding layer is based on the VGG-16 network and the visual transformer network. The output of the multi-level convolution is sent to the visual transformer module, and all the results are merged into the final result image; The density prediction unit builds a neural network model of the visual converter, inputs the training data into the built neural network model based on the visual converter, thereby training the model, and inputs the pre-processed image into the trained neural network model based on the visual converter mechanism, thereby obtaining the crowd density results in each of the pre-processed images and outputting them. Among them, the neural network model based on the visual converter includes a local information extraction module based on a convolutional neural network at the front end and a global information extraction module based on the visual converter at the back end.

Citation Information

Patent Citations

  • Crowd density estimation method and device based on attention type deep neural network

    CN110889343A