A Method for Estimating Population Density by Fusing Features of Three Arrays Based on Residual Network

Through the three-array feature fusion population density estimation algorithm based on residual network, the improved ResNet34 network and hollow convolution process population density maps are solved, and the accuracy and robustness of density estimation are significantly improved.

CN114972936BActive Publication Date: 2025-05-27CHANGCHUN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110597361.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-05-31
Publication Date
2025-05-27
Estimated Expiration
2041-05-31

AI Technical Summary

Technical Problem

The multi-scale problems caused by camera perspective in the crowd density estimation task are difficult to effectively deal with in the prior art, resulting in a decrease in estimation accuracy.

Method used

The three-array feature fusion population density estimation algorithm based on residual network is used to extract picture features through the improved ResNet34 network, and the output is processed using three columns of hollow convolutions with different hollow rates. Finally, the three columns of outputs are cascaded to generate high-quality density prediction maps.

Benefits of technology

It effectively solves the multi-scale problem, improves the accuracy and robustness of population density estimation, and the generated density map is more accurate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure QLYQS_1
    Figure QLYQS_1
  • Figure QLYQS_2
    Figure QLYQS_2
  • Figure 210527175030
    Figure 210527175030
Patent Text Reader

Abstract

The technical problem solved by the present invention is the multi-scale problem caused by the camera perspective in the crowd density estimation task. A crowd density estimation algorithm based on three-array feature fusion of residual network is proposed. The improved ResNet34 network is used to extract the image feature information, and the output is processed by three dilated convolutions with different dilation rates. The principle of dilated convolution is to perform zero-padding operations at intervals on the convolutional kernel based on ordinary convolution, so that the receptive field can be expanded without increasing the network parameters, thereby improving the network's ability to capture multi-scale information of images. Finally, the outputs of the three arrays are cascaded to obtain a high-quality density prediction map. The present invention can effectively solve the influence of crowd density, occlusion, and multi-scale on the detection accuracy in the crowd density estimation work.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Based on deep learning, the present invention proposes a three-array feature fusion crowd density estimation algorithm based on a residual network. The improved residual network is used at the front end of the network to extract the shallow information of the picture well. At the back end of the network, dilated convolutions with different dilation rates in three arrays are used to process the output of the front end. The advantage of dilated convolution that can increase the receptive field without increasing network parameters can better capture the multi-scale information of the image. The three-column outputs are cascaded to obtain a high-quality and high-precision predicted density map, so as to make the crowd density estimation more accurate. Background Art

[0002] Crowd density, as a basic attribute of the crowd, reflects the degree of crowding in the space. The more crowded the crowd is, the greater the possibility of personnel casualties, resulting in greater casualties and property losses. In order to prevent similar incidents, real-time monitoring of crowd density can timely detect abnormal situations and notify relevant departments to take preventive measures. Therefore, accurate detection of crowd density is of great significance and research value for preventing the occurrence of malignant events. The density, occlusion, and multi-scale nature in the crowd have a great impact on the detection accuracy, especially the scale diversity in the crowd brings great challenges to the research in this field.

[0003] In recent years, deep learning has developed rapidly. Due to the excellent performance of convolutional neural networks in the field of image processing, researchers in the field of crowd density estimation have also applied convolutional neural networks to this field. In particular, MCNN designs three networks with different convolutional kernel sizes to deal with the problem of different scales in crowd data and shows good results. However, when it comes to dealing with a large number of people (many scales), many scene occlusions, and perspective problems, the effect drops significantly. Later, networks represented by CSRNet successfully applied dilated convolution to this field, reduced the impact of crowd density discretization, maintained the continuity of local areas, and further improved the accuracy of crowd density estimation. Summary of the Invention

[0004] The technical problem solved by the present invention is the multi-scale problem caused by the camera perspective in the crowd density estimation task. A three-array feature fusion crowd density estimation algorithm based on a residual network is proposed. The improved ResNet34 network is used to extract the feature information of the picture, and dilated convolutions with different dilation rates in three columns are used to process the output. The principle of dilated convolution is to perform an operation of padding zeros at intervals on the convolutional kernel based on ordinary convolution, so that the receptive field can be expanded without increasing network parameters, thereby improving the ability of the network to capture multi-scale information of the image. Finally, the outputs of the three arrays are cascaded to obtain a high-quality density prediction map.

[0005] To achieve the above object, the technical solution of the present invention is as follows:

[0006] A three-array feature fusion crowd density estimation algorithm based on a residual network, comprising the following steps:

[0007] Step 1: Data preprocessing, resizing the dataset to a size that can be trained by the network according to requirements;

[0008] Step 2: Construct an improved ResNet34 network model, and use this model to extract features from the test images, and use the method of transfer learning without updating the network parameters;

[0009] Step 3: Construct a three-array dilated convolution network to process the output of the improved ResNet34 network, extract more multi-scale information of the image, and continuously update the parameter weights of the dilated convolution network during training;

[0010] Step 4: Concatenate the output results of the three columns of dilated convolutions to obtain a density prediction map;

[0011] Step 5: After the network is trained, input the test set for testing to verify the network effect.

[0012] The specific situation in Step 1 is as follows:

[0013] The Shanghaitech dataset is divided into two parts, Part_A and Part_B. Different from Part_B, the crowd pictures in Part_B are all 1064×768 in size, and the crowd pictures in Part_A are of different sizes, which can better detect the performance of the network under different conditions. The pictures are reduced to 1 / 8 of the original size to meet the requirements of the network.

[0014] The specific situation in Step 2 is as follows:

[0015] (1) Construct an improved ResNet34 network to extract shallow feature information of the picture:

[0016] ① Replace the 7×7 convolution in the first layer of ResNet34 with a 3×3 convolution and remove the second pooling layer and the last fully connected layer;

[0017] ② A residual block can be expressed as:

[0018]

[0019] The residual block is divided into two parts, the direct mapping part and the residual part, where is the direct mapping part, is the residual part;

[0020] The training uses a hybrid loss function, the Euclidean distance, to measure the difference between the true density map and the density map generated by the network. The Euclidean calculation formula is:

[0021]

[0022] Among them, represents the loss of the crowd density estimation task, S represents the number of images in a training batch, that is, the batch size, represents the estimated density value of the i-th training sample x i with parameters θ, represents the true density value;

[0023] The constructed front-end network model uses 3 times of downsampling to reduce the size of the picture to 1 / 8 of the original. When residual blocks with different numbers of channels are connected to each other, a 1×1 convolution layer is added to change the output dimension, so that the subsequent convolution layers can retain part of the original information, preventing information loss and network degradation due to the excessive depth of the network;

[0024] The number of channels of the last convolution layer is 512 and it is connected to the back-end network;

[0025] (2) For transfer learning, directly call the ResNet34 network parameters and do not update the parameters during training.

[0026] The specific situation in the third step is as follows:

[0027] (1) Build a three-column dilated convolutional neural network with the number of channels [512, 512, 512, 256, 128, 64], and input the outputs of the previous network into these three columns of convolutions for processing;

[0028] (2) Adjust the dilation rates of the three-column dilated convolutional neural network to 1, 2, and 4 respectively.

[0029] The specific situation in the fourth step is as follows:

[0030] Concatenate the outputs of the three-column dilated convolutions, input them into a 1×1 convolution to reduce the dimension to 1, and then use it as the final result output of the entire network

[0031] output.

[0032] The specific situation in the fifth step is as follows:

[0033] The test set images are input into the trained network to generate crowd density maps and output the number of people in the maps. After comparing the estimated number of people with the actual number of people in the original images, the Mean Absolute Error (MAE) and Mean Squared Error (MSE) are used as evaluation metrics. MAE can reflect the accuracy of the algorithm, while MSE reflects the robustness of the algorithm. The formulas are as follows:

[0034]

[0035]

[0036] Among them, M is the total number of test images, Q(i) represents the predicted number of people in the test set, and P(i) represents the actual number of people marked in the test set.

[0037] Compared with the prior art, the beneficial effects of the technical solution of the present invention are as follows:

[0038] (1) Based on the ResNet network structure, the present invention retains some original information in the output data, effectively preventing the problem of gradient dispersion during backpropagation;

[0039] (2) The present invention introduces dilated convolution to extract image features. The zero-padding operation of dilated convolution can increase the receptive field without increasing the network parameters. Compared with pooling operations that require reducing the image size to achieve the same effect, dilated convolution can effectively reduce information loss;

[0040] (3) Dilated convolutions with different dilation rates can capture more scale detail information, and the effect on dealing with multi-scale tasks such as dense crowd estimation is very significant. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 is the schematic diagram of the residual block used in the present invention;

[0042] Figure 2 is the effect diagram of the dilated convolution used in the present invention;

[0043] Figure 3 is the network model based on ResNet34 used in the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0044] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted. The technical solution of the present invention will be further described below with reference to the drawings and embodiments.

[0045] The present invention provides a three-array feature fusion crowd density estimation algorithm based on a residual network. This method effectively solves the multi-scale problem of crowded images, and the generated high-quality density map significantly improves the accuracy of crowd density work.

[0046] Figure 1 It is the schematic diagram of the residual block of the residual network applied in the present invention. First is image preprocessing, which transforms the ShanghaiTech dataset into the input required by the network. Then an improved ResNet34 network model is constructed, and the loaded network parameters are not updated. Next, dilated convolutions with different dilation rates are built in three columns at the back end. Finally, the results processed by the dilated convolutions are cascaded and dimension-reduced to output the crowd density map training, realizing the crowd density estimation task.

[0047] The specific implementation steps are as follows:

[0048] Step1.1 Compress the picture to 1 / 8 of the original picture;

[0049] Step1.2 Add a dimension to the compressed picture so that it can correspond to the network output dimension;

[0050] Step2.1 Use the improved ResNet34 network to extract the shallow features of the picture;

[0051] Step2.1.1 Replace the 7×7 convolution in the first layer of ResNet34 with a 3×3 convolution to convert the 3 channels of the input picture into 64 channels, and remove the second pooling layer and the last fully connected layer;

[0052] Step2.1.2 After three changes in the number of channels in the residual network, the network output picture becomes 1 / 8 of the original picture size. The residual module can effectively prevent information loss and network degradation due to the excessive depth of the network;

[0053] A residual block can be expressed as:

[0054]

[0055] The residual block is divided into two parts: the direct mapping part and the residual part, where is the direct mapping part, is the residual part;

[0056] Step2.2 Directly call the originally trained parameters of ResNet34, and do not update the parameters of the residual network during the training process

[0057] The training uses a hybrid loss function, the Euclidean distance, to measure the difference between the true density map and the density map generated by the network. The Euclidean calculation formula:

[0058]

[0059] Among them, represents the loss of the crowd density estimation task, S represents the number of images in a training batch, i.e., the batch size, represents the estimated density value of the i-th training sample x i with parameter θ, represents the ground-truth density value;

[0060] Step3.1 Construct a three-array dilated convolutional network to process the output of the improved ResNet34 network. The convolutional kernels of each column of the constructed network are all 3×3, and the number of network channels are all [512, 512, 512, 256, 128, 64];

[0061] Step3.2 The dilation rates of the three columns of dilated convolutional networks are 1, 2, and 4 respectively, so that the network can capture more scale detail information;

[0062] Step4.1 Concatenate the output results of the three columns of dilated convolutions;

[0063] Step4.2 Use 1×1 convolution to change the number of channels of the output result to 1 and output the final density map;

[0064] Step5 Use the test set to verify the algorithm effect. The verification metrics use the Mean Absolute Error (MAE) and the Mean Square Error (MSE) as evaluation metrics. MAE can reflect the accuracy of the algorithm, and MSE reflects the robustness of the algorithm. The formulas are as follows:

[0065]

[0066]

[0067] Among them, M is the total number of test pictures, Q(i) represents the predicted number of people in the test set, and P(i) represents the true number of people marked in the test set.

Claims

1. A method for estimating crowd density by fusing features of three arrays based on a residual network, characterized in that, it includes the following steps: Step 1: Data preprocessing, resizing the dataset to a size that can be trained by the network according to requirements; Step 2: Construct an improved ResNet34 network model, and use this model to extract features from the test image. Using the method of transfer learning, without updating the network parameters, the specific process is as follows: Step 2.1 Use the improved ResNet34 network to extract shallow features of the picture; Step 2.1.1 Replace the 7×7 convolution in the first layer of ResNet34 with a 3×3 convolution to convert the 3 channels of the input picture into 64 channels, and remove the second pooling layer and the last fully connected layer; Step 2.1.2 After three changes in the number of channels in the residual network, the network output picture becomes 1 / 8 of the original picture size. The residual module can effectively prevent information loss and network degradation caused by the network being too deep; A residual block can be expressed as: x l+1 = x l + F(x l ; W l ) The residual block is divided into two parts, the direct mapping part and the residual part, where x l is the direct mapping part, F(x l ; W l ) is the residual part; Step 2.2 Directly call the originally trained parameters of ResNet34, and do not update the parameters of the residual network during the training process; The training uses a hybrid loss function, the Euclidean distance, to measure the difference between the real density map and the density map generated by the network. The Euclidean calculation formula: Among them, L e represents the loss of the crowd density estimation task, S represents the number of images in a training batch, i.e., batchsize, and G(X i, θ) represents the estimated density value of the i-th training sample X i , with the parameter being θ, and D i GT represents the true density value; Step 3: Construct a three-array dilated convolution network to process the output of the improved ResNet34 network, extract more multi-scale information of the image, and continuously update the parameter weights of the dilated convolution network during the training process. The specific process is as follows: Step 3.1 Construct a three-array dilated convolution network to process the output of the improved ResNet34 network. The convolution kernels of each column of the constructed network are all 3×3, and the number of network channels is [512, 512, 512, 256, 128, 64]; Step 3.2 The dilation rates of the three columns of dilated convolution networks are 1, 2, and 4 respectively, so that the network can capture more scale detail information; Step 4: Concatenate the output results of the three columns of dilated convolution to obtain a density prediction map; Step 5: After the network is trained, input the test set for testing to check the network effect.

2. The method for estimating crowd density by fusing features of three arrays based on a residual network according to claim 1, characterized in that, the specific process in Step 1 is as follows: Step 1.1 Compress the picture to 1 / 8 of the original picture; Step 1.2 Add a dimension to the compressed picture so that it can correspond to the network output dimension.

3. The method for estimating crowd density by fusing features of three arrays based on a residual network according to claim 1, characterized in that, the specific process in Step 4 is as follows: Step 4.1 Concatenate the output results of the three columns of dilated convolution; Step 4.2 Use a 1×1 convolution to change the number of channels of the output result to 1, and output the final density map.

4. The method for estimating crowd density by fusing features of three arrays based on a residual network according to claim 1, characterized in that, the specific process in Step 5 is as follows: Use the test set to verify the algorithm's effectiveness. The evaluation metrics are the Mean Absolute Error (MAE) and the Mean Square Error (MSE); MAE can reflect the accuracy of the algorithm, while MSE reflects the robustness of the algorithm. The formulas are as follows: Among them, M is the total number of test images, Q(i) represents the predicted number of people in the test set, and P(i) represents the actual number of people marked in the test set.

Citation Information

Patent Citations

  • Dense crowd counting method and device

    CN109241895A

  • Crowd density estimation method based on multi-scale convolutional neural network

    CN110020606A