A crowd counting method and device based on multi-level feature fusion
By adopting a multi-level feature fusion network in crowd counting, extracting and fusing details and semantic features of different levels, the problem of degradation in the counting effect of the prior art under crowd density or occlusion is solved, and higher counting accuracy and network performance are achieved.
Patent Information
- Application Number
- CN202210005553.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-05
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-01-05
AI Technical Summary
The existing crowd counting method has reduced the counting effect in the case of crowd dense or occlusion, and the convolutional neural network-based method has lost detailed features due to multiple convolution operations, so the accuracy of identifying small targets is low.
A multi-level feature fusion network is used to extract detailed features and semantic features of different levels through the front-end and back-end parts, and fuse them to varying degrees to generate a predicted population density map.
It improves the accuracy of crowd counting and can better identify large and small pedestrians, thereby improving network counting performance.
Smart Images

Figure CN114359833B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of crowd counting, and in particular to a crowd counting method and device based on multi-level feature fusion. Background Art
[0002] Crowd counting is mainly used to estimate the number of people and the distribution of people. It can usually be divided into traditional crowd counting methods and methods based on convolutional neural networks. Traditional crowd counting methods can be further divided into detection-based crowd counting methods and regression-based crowd counting methods. In the detection-based crowd counting method, the main approach is to first extract the features of pedestrians through the crowd image, and then use the extracted features to train the classifier, so that the pedestrians in the crowd image can be identified, and then the pedestrians are counted to estimate the number of people in the crowd image. This method will have a good effect in scenes with extremely sparse crowds, but once the number of pedestrians in the scene increases and the pedestrians' body parts are blocked, the counting effect will drop sharply. In regression-based crowd counting, the image features are directly mapped to the number of people through the learning of the regression model to obtain the number of people in the image. The crowd counting method based on convolutional neural networks is currently the mainstream method for crowd counting. Although the crowd counting method based on convolutional neural networks has achieved good results, there is still the problem of background interference. In the MCNN proposed by Zhang et al., three columns of convolution kernels are used to extract pedestrian features of three different scales: large, medium, and small. They are then fused to obtain the final crowd features to generate a crowd density map. In the Switch-CNN proposed by Sam et al., the crowd image is first divided into 9 parts of 3*3, and then these image blocks are divided into three categories according to the density through a switch network, and are respectively put into the corresponding network for training and learning. However, due to the repeated use of convolution operations in the above method, many detailed features are lost, making the network's recognition accuracy for small targets relatively low. Summary of the invention
[0003] In order to overcome the shortcomings of the existing technology, the present invention provides a crowd counting method and device based on multi-level feature fusion. The present invention combines the advantages of different levels of detail features into different levels of semantic features, effectively improving the accuracy of crowd counting.
[0004] In order to achieve the above object, the present invention adopts the following technical solutions:
[0005] According to a first aspect of the present invention, a crowd counting method based on multi-level feature fusion is provided, comprising the following steps:
[0006] Data preprocessing: Generate a real crowd density map based on the images in the public dataset for crowd counting and the corresponding two-dimensional coordinate marks of the center points of human heads. The real crowd density map is used for network training.
[0007] Constructing a multi-level feature fusion network, inputting a crowd image for which pedestrian quantity estimation is required into the multi-level feature fusion network, and outputting a predicted crowd density map;
[0008] The multi-level feature fusion network is divided into two parts, a front-end and a back-end. The front-end is used to extract detail features, and the front-end is divided into three detail feature extraction layers of different levels according to the network depth to extract detail features of different levels. The back-end is used to extract semantic features, and the back-end is divided into three semantic feature extraction layers of different levels according to the network depth to extract semantic features of different levels. Detail features of different levels are fused into semantic features of different levels to different degrees, and finally a predicted crowd density map is obtained through a convolution operation.
[0009] The front-end part is based on the VGG16 network and is divided into three different levels of detail feature extraction layers according to the network depth. The first level detail feature extraction layer is the 1st to 4th layers of the VGG16 network, and the first level detail feature A1 is extracted. The second level detail feature extraction layer is the 1st to 7th layers of the VGG16 network, and the second level detail feature A2 is extracted. The third level detail feature extraction layer is the 1st to 10th layers of the VGG16 network, and the third level detail feature A3 is extracted.
[0010] The back-end part is composed of 3 fusion layers and 6 residual blocks (RB), and is divided into three different levels of semantic feature extraction layers according to the network depth. The first-level semantic feature extraction layer is composed of a residual block followed by the front-end network, and the first-level semantic feature B1 is extracted. Then, the first-level detail feature A1 extracted by the first-level detail feature extraction layer is fused with the first-level semantic feature B1 extracted by the first-level semantic feature extraction layer through the first fusion layer to obtain the fused feature H1. The fusion layer is a channel splicing operation. The H1 process is as follows:
[0011] H1=C cat (A1,B1)
[0012] Among them C cat Represents the channel splicing operation; the second-level semantic feature extraction layer is composed of the first fusion layer followed by two residual blocks, which extracts the second-level semantic features. Then, the second-level semantic features, the second-level detail features, and the third-level detail features are fused through the second fusion layer to obtain the fused feature H2. The H2 process is as follows:
[0013] H2=C cat(A2,A3,C RB (H1)
[0014] Among them C RB Represents the residual block feature mapping relationship; the third-level semantic feature extraction layer is composed of the second fusion layer followed by three residual blocks, and the third-level semantic features are extracted. Then, the third-level semantic features, the first-level detail features, and the second-level detail features are fused through the third fusion layer to obtain the fused feature H3. The H3 process is as follows:
[0015] H3=C cat (A1,A2,C RB (H2)
[0016] The fused feature H3 is passed through a 1×1 convolutional layer to generate a predicted crowd density map;
[0017] The predicted crowd density map and the real crowd density map are compared by using the Euclidean loss function to obtain the network loss size, and the network parameters are continuously updated by using back propagation until suitable network parameters are obtained, and the final network model is obtained by training;
[0018] The crowd image for which the number of pedestrians needs to be estimated is input into the trained network model to obtain a crowd density map, which is then integrated pixel by pixel, that is, the values of all pixels in the map are added up to obtain an estimate of the number of pedestrians in the image, that is, the predicted total number of people.
[0019] Furthermore, the real crowd density map is expressed as:
[0020]
[0021] Among them, N represents the number of people marked in the crowd image, x i represents the 2D coordinate position of the center point of the head of the i-th person, x represents the position of other pixels in the crowd image except the head, δ(x i ) represents the impulse function, The standard deviation is σ i The adaptive Gaussian kernel of Where β is the weight parameter, Represents x i The average distance between the heads and their nearest neighbors.
[0022] Furthermore, the method also includes a data enhancement step, which is: first, the images in the data set are grayscaled, and then 9 sub-sample images with a size of one-quarter of the original image are randomly cropped from each grayscale image, and they are randomly rotated and flipped as training sets.
[0023] Furthermore, the network parameters are continuously updated by back propagation to train the final network model, and the specific steps include:
[0024] Based on the pytorch deep learning framework training, the network training error is obtained using the Euclidean loss function, which is defined as:
[0025]
[0026] Among them, M represents the size of the training set used, θ represents the network parameters, and X k represents the kth image input, D k , D(X k ; θ) represent the image X k The real crowd density map and the predicted crowd density map generated after network training;
[0027] Update the network model parameters through back propagation and gradient descent;
[0028] Set the appropriate size of training rounds. When the training rounds reach the set size, the training ends and the final network model parameters are obtained.
[0029] According to a second aspect of the present invention, there is provided a crowd counting device based on multi-level feature fusion, comprising a data processing module, a multi-level feature fusion network construction module, a network training module and a pedestrian number counting module;
[0030] The data processing module is used for data processing: generating a real crowd density map according to the images in the public data set for crowd counting and the corresponding two-dimensional coordinate marks of the center points of human heads, wherein the real crowd density map is used for network training;
[0031] The multi-level feature fusion network construction module is used to construct a multi-level feature fusion network, input a crowd image for which pedestrian quantity estimation is required into the multi-level feature fusion network, and output a predicted crowd density map;
[0032] The multi-level feature fusion network is divided into two parts, a front-end and a back-end. The front-end is used to extract detail features, and the front-end is divided into three detail feature extraction layers of different levels according to the network depth to extract detail features of different levels. The back-end is used to extract semantic features, and the back-end is divided into three semantic feature extraction layers of different levels according to the network depth to extract semantic features of different levels. Detail features of different levels are fused into semantic features of different levels to different degrees, and finally a predicted crowd density map is obtained through a convolution operation.
[0033] The network training module is used to obtain a network error by applying a Euclidean loss function to the real crowd density map and the predicted crowd density map, and then update the network parameters by back propagation to obtain a final network model through training;
[0034] The pedestrian counting module is used to input the crowd image for which pedestrian number estimation is required into the final network model, output a crowd density map, and then perform pixel-by-pixel integration on the map, that is, add up the values of all pixel points in the map to obtain an estimated value of the number of pedestrians in the image, that is, the predicted total number of people.
[0035] Furthermore, the front end of the multi-level feature fusion network is based on the VGG16 network, and is divided into three different levels of detail feature extraction layers according to the network depth. The first level detail feature extraction layer is the 1st to 4th layers of the VGG16 network, and the first level detail feature A1 is extracted. The second level detail feature extraction layer is the 1st to 7th layers of the VGG16 network, and the second level detail feature A2 is extracted. The third level detail feature extraction layer is the 1st to 10th layers of the VGG16 network, and the third level detail feature A3 is extracted.
[0036] The back-end part is composed of 3 fusion layers and 6 residual blocks (RB), and is divided into three different levels of semantic feature extraction layers according to the network depth. The first-level semantic feature extraction layer is composed of a residual block followed by the front-end network, and the first-level semantic feature B1 is extracted. Then, the first-level detail feature A1 extracted by the first-level detail feature extraction layer is fused with the first-level semantic feature B1 extracted by the first-level semantic feature extraction layer through the first fusion layer to obtain the fused feature H1. The fusion layer is a channel splicing operation. The H1 process is as follows:
[0037] H1=C cat (A1,B1)
[0038] Among them C cat Represents the channel splicing operation; the second-level semantic feature extraction layer is composed of the first fusion layer followed by two residual blocks, which extracts the second-level semantic features. Then, the second-level semantic features, the second-level detail features, and the third-level detail features are fused through the second fusion layer to obtain the fused feature H2. The H2 process is as follows:
[0039] H2=C cat (A2,A3,C RB (H1)
[0040] Among them C RBRepresents the residual block feature mapping relationship; the third-level semantic feature extraction layer is composed of the second fusion layer followed by three residual blocks, and the third-level semantic features are extracted. Then, the third-level semantic features, the first-level detail features, and the second-level detail features are fused through the third fusion layer to obtain the fused feature H3. The H3 process is as follows:
[0041] H3=C cat (A1,A2,C RB (H2)
[0042] The fused feature H3 is passed through a 1×1 convolutional layer to generate a predicted crowd density map.
[0043] According to a third aspect of the present invention, there is provided a computer device, comprising a memory and a processor, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the processor executes the steps of the above-mentioned crowd counting method based on multi-level feature fusion.
[0044] According to a fourth aspect of the present invention, there is provided a storage medium storing computer-readable instructions, which, when executed by one or more processors, causes the one or more processors to perform the steps of the above-mentioned crowd counting method based on multi-level feature fusion.
[0045] The beneficial effects of the present invention are as follows: since the detail features of the shallow network are conducive to identifying small-scale pedestrian features, the present invention divides the front-end network into different levels according to the network depth to extract detail features of different degrees, thereby better identifying small-scale pedestrians; since the semantic features of the deep network are conducive to identifying large-scale pedestrian information, the present invention divides the back-end network into different levels according to the network depth to extract semantic features of different degrees, and then merges the detail features of different levels with the semantic features of different levels, so that the network takes into account both the detail features that are conducive to the recognition of small-scale pedestrians and the semantic features that are conducive to the recognition of large-scale pedestrians, thereby better identifying large- and small-scale pedestrians, thereby improving the network counting performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 A schematic diagram of a multi-level feature fusion network structure provided by an embodiment of the present invention;
[0047] Figure 2 This is a flow chart of network training of the present invention. DETAILED DESCRIPTION
[0048] The present disclosure is further described below in conjunction with the accompanying drawings and embodiments.
[0049] Example 1
[0050] This embodiment provides a crowd counting method based on multi-level feature fusion, including the following steps:
[0051] S1: Data preprocessing
[0052] Download the crowd counting dataset used by the public. The images in the crowd counting dataset are all RGB three-channel color images, and the heads contained in each image are annotated with two-dimensional coordinates. The corresponding real crowd density map is generated for each image in the dataset through the image and head annotation information. This map is used for network training;
[0053] The expression for generating the real crowd density map is:
[0054]
[0055] Among them, N represents the number of people marked in the crowd image, x i represents the 2D coordinate position of the center point of the head of the i-th person, x represents the position of other pixels in the crowd image except the head, δ(x i ) represents the impulse function, The standard deviation is σ i The adaptive Gaussian kernel of Where β is the weight parameter, usually β = 0.3, Represents x i The average distance between the heads and their nearest neighbors.
[0056] In this embodiment, data enhancement is performed in the following manner:
[0057] First, the images in the dataset are grayscaled, and then nine sub-sample images with a quarter of the original size are randomly cropped from each grayscale image. They are randomly rotated and flipped as training sets.
[0058] S2: construct a multi-level feature fusion network, input a crowd image for which pedestrian quantity estimation is required into the multi-level feature fusion network, and output a predicted crowd density map;
[0059] like Figure 1 As shown, The construction process of the multi-level feature fusion network is as follows:
[0060] The multi-level feature fusion network divides the network into two parts: the front-end and the back-end. The front-end is used to extract detail features, and is divided into three different levels of detail feature extraction layers according to the network depth to extract detail features of different levels. The back-end is used to extract semantic features, and is divided into three different levels of semantic feature extraction layers according to the network depth to extract semantic features of different levels. Detail features of different levels are fused into semantic features of different levels to varying degrees, and finally the predicted crowd density map is obtained through convolution operations.
[0061] The front-end part is based on the VGG16 network and is divided into three different levels of detail feature extraction layers according to the network depth. The first level detail feature extraction layer is the 1st to 4th layers of the VGG16 network, and the first level detail feature A1 is extracted. The second level detail feature extraction layer is the 1st to 7th layers of the VGG16 network, and the second level detail feature A2 is extracted. The third level detail feature extraction layer is the 1st to 10th layers of the VGG16 network, and the third level detail feature A3 is extracted.
[0062] The backend part consists of 3 fusion layers and 6 residual blocks (RB), and is divided into three different levels of semantic feature extraction layers according to the network depth. The first-level semantic feature extraction layer is composed of a residual block followed by the front-end network, and the first-level semantic feature B1 is extracted. Then, the first-level detail feature A1 extracted by the first-level detail feature extraction layer is fused with the first-level semantic feature B1 extracted by the first-level semantic feature extraction layer through the first fusion layer to obtain the fused feature H1. The fusion layer is a channel splicing operation. The H1 process is as follows:
[0063] H1=C cat (A1,B1)
[0064] Among them C cat Represents the channel splicing operation; the second-level semantic feature extraction layer is composed of the first fusion layer followed by two residual blocks, which extracts the second-level semantic features. Then, the second-level semantic features, the second-level detail features, and the third-level detail features are fused through the second fusion layer to obtain the fused feature H2. The H2 process is as follows:
[0065] H2=C cat (A2,A3,C RR (H1)
[0066] Among them C RBRepresents the residual block feature mapping relationship; the third-level semantic feature extraction layer is composed of the second fusion layer followed by three residual blocks, and the third-level semantic features are extracted. Then, the third-level semantic features, the first-level detail features, and the second-level detail features are fused through the third fusion layer to obtain the fused feature H3. The H3 process is as follows:
[0067] H3=C cat (A1,A2,C RB (H2)
[0068] The fused feature H3 is passed through a 1×1 convolutional layer to generate a predicted crowd density map;
[0069] The predicted crowd density map and the real crowd density map are compared by using the Euclidean loss function to obtain the network loss size, and the network parameters are continuously updated by using back propagation until suitable network parameters are obtained, and the final network model is obtained by training;
[0070] S3: The predicted crowd density map and the actual crowd density map are compared by the Euclidean loss function to obtain the network loss size, and the network parameters are continuously updated by back propagation until the appropriate network parameters are obtained, and the final network model is obtained by training; in this embodiment, the network training process is as follows:
[0071] Based on the pytorch deep learning framework training, the optimizer used in the training is Adam, and the Euclidean loss function is used to obtain the network training error. The Euclidean loss function is defined as:
[0072]
[0073] Among them, M represents the size of the training set used, θ represents the network parameters, and X k represents the kth image input, D k , D(X k ; θ) represent the image X k The real crowd density map and the predicted crowd density map generated after network training;
[0074] Update the network model parameters through back propagation and gradient descent;
[0075] Set the training rounds to 2000 and the batch size to 2. When the training rounds reach the set value, the network training ends and the optimal parameters of the network model in this training are obtained. Figure 2 For the training process.
[0076] S4: Input the crowd image whose number of pedestrians needs to be estimated into the trained network model to obtain a crowd density map, and perform pixel-by-pixel integration on the map, that is, add up the values of all pixel points in the map to obtain an estimated value of the number of pedestrians in the image, that is, the predicted total number of people.
[0077] At the same time, the mean absolute error (MAE) and root mean square error (MSE) are used to evaluate the network effect. The calculation formula is as follows:
[0078]
[0079]
[0080] Where S represents the number of images in a training batch, represents the true value of the number of pedestrians in the yth image, C y It represents the estimated value of the number of pedestrians in the image obtained after the yth image passes through the network. The smaller the MAE or MSE value, the better the network effect.
[0081] As can be seen from Table 1, the present invention effectively improves the accuracy of crowd counting by integrating detail features of different levels into different semantic features to different degrees.
[0082] Table 1 Network effect evaluation table
[0083]
[0084] Example 2
[0085] This embodiment provides a crowd counting device based on multi-level feature fusion, which includes a data processing module, a multi-level feature fusion network construction module, a network training module and a pedestrian counting module;
[0086] The data processing module is used for data processing: generating a real crowd density map based on the images in the public data set for crowd counting and the corresponding two-dimensional coordinate marks of the center points of the human heads. The real crowd density map is used for network training.
[0087] The multi-level feature fusion network construction module is used to construct a multi-level feature fusion network, input the crowd image that needs to be estimated into the multi-level feature fusion network, and output the predicted crowd density map;
[0088] The multi-level feature fusion network is divided into two parts: the front-end and the back-end. The front-end is used to extract detail features, and is divided into three detail feature extraction layers of different levels according to the network depth to extract detail features of different levels. The back-end is used to extract semantic features, and is divided into three semantic feature extraction layers of different levels according to the network depth to extract semantic features of different levels. Detail features of different levels are fused into semantic features of different levels to different degrees, and finally the predicted crowd density map is obtained through convolution operation.
[0089] The network training module is used to obtain the network error by using the Euclidean loss function for the real crowd density map and the predicted crowd density map, and then update the network parameters by back propagation to obtain the final network model through training;
[0090] The pedestrian counting module is used to input the crowd image for which pedestrian number estimation is required into the final network model, and output the crowd density map, which is then integrated pixel by pixel, that is, the values of all pixels in the map are added up to obtain the estimated value of the number of pedestrians in the image, that is, the predicted total number of people.
[0091] In this embodiment, the front end of the multi-level feature fusion network is based on the VGG16 network, and is divided into three different levels of detail feature extraction layers according to the network depth. The first level detail feature extraction layer is the 1st to 4th layers of the VGG16 network, and the first level detail feature A1 is extracted. The second level detail feature extraction layer is the 1st to 7th layers of the VGG16 network, and the second level detail feature A2 is extracted. The third level detail feature extraction layer is the 1st to 10th layers of the VGG16 network, and the third level detail feature A3 is extracted.
[0092] The backend part consists of 3 fusion layers and 6 residual blocks (RB), and is divided into three different levels of semantic feature extraction layers according to the network depth. The first-level semantic feature extraction layer is composed of a residual block followed by the front-end network, and the first-level semantic feature B1 is extracted. Then, the first-level detail feature A1 extracted by the first-level detail feature extraction layer is fused with the first-level semantic feature B1 extracted by the first-level semantic feature extraction layer through the first fusion layer to obtain the fused feature H1. The fusion layer is a channel splicing operation. The H1 process is as follows:
[0093] H1=C cat (A1,B1)
[0094] Among them C catRepresents the channel splicing operation; the second-level semantic feature extraction layer is composed of the first fusion layer followed by two residual blocks, which extracts the second-level semantic features. Then, the second-level semantic features, the second-level detail features, and the third-level detail features are fused through the second fusion layer to obtain the fused feature H2. The H2 process is as follows:
[0095] H2=C cat (A2,A3,C RB (H1)
[0096] Among them C RB Represents the residual block feature mapping relationship; the third-level semantic feature extraction layer is composed of the second fusion layer followed by three residual blocks, and the third-level semantic features are extracted. Then, the third-level semantic features, the first-level detail features, and the second-level detail features are fused through the third fusion layer to obtain the fused feature H3. The H3 process is as follows:
[0097] H3=C cat (A1,A2,C RB (H2)
[0098] The fused feature H3 is passed through a 1×1 convolutional layer to generate a predicted crowd density map.
[0099] In one embodiment, a computer device is proposed, including a memory and a processor, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the processor executes the steps in the crowd counting method based on multi-level feature fusion in the above embodiments.
[0100] In one embodiment, a storage medium storing computer-readable instructions is proposed, and when the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps in the crowd counting method based on multi-level feature fusion in the above embodiments. The storage medium may be a non-volatile storage medium.
[0101] A person of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, and the storage medium may include: a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, etc.
[0102] The above description is merely a preferred embodiment of one or more embodiments of the present specification and is not intended to limit one or more embodiments of the present specification. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of one or more embodiments of the present specification shall be included in the scope of protection of one or more embodiments of the present specification.
Claims
1. A crowd counting method based on multi-level feature fusion, characterized in that: The following steps are involved: Data preprocessing: Generate a real crowd density map based on the images in the public dataset for crowd counting and the corresponding two-dimensional coordinate marks of the center points of human heads. The real crowd density map is used for network training. Constructing a multi-level feature fusion network, inputting a crowd image for which pedestrian quantity estimation is required into the multi-level feature fusion network, and outputting a predicted crowd density map; The multi-level feature fusion network is divided into two parts, a front-end and a back-end. The front-end is used to extract detail features, and the front-end is divided into three detail feature extraction layers of different levels according to the network depth to extract detail features of different levels. The back-end is used to extract semantic features, and the back-end is divided into three semantic feature extraction layers of different levels according to the network depth to extract semantic features of different levels. Detail features of different levels are fused into semantic features of different levels to different degrees, and finally a predicted crowd density map is obtained through a convolution operation. The front-end part is based on the VGG16 network and is divided into three different levels of detail feature extraction layers according to the network depth. The first level detail feature extraction layer is the 1st to 4th layers of the VGG16 network, and the first level detail feature A1 is extracted. The second level detail feature extraction layer is the 1st to 7th layers of the VGG16 network, and the second level detail feature A2 is extracted. The third level detail feature extraction layer is the 1st to 10th layers of the VGG16 network, and the third level detail feature A3 is extracted. The backend part is composed of 3 fusion layers and 6 residual blocks, and is divided into three different levels of semantic feature extraction layers according to the network depth. The first level semantic feature extraction layer is composed of a residual block followed by the front-end network, and the first level semantic feature B1 is extracted. Then, the first level detail feature A1 extracted by the first level detail feature extraction layer is fused with the first level semantic feature B1 extracted by the first level semantic feature extraction layer through the first fusion layer to obtain the fused feature H1. The fusion layer is a channel splicing operation. The H1 process is as follows: H1=C cat (A1,B1) Among them C cat Represents the channel splicing operation; the second-level semantic feature extraction layer is composed of the first fusion layer followed by two residual blocks, which extracts the second-level semantic features. Then, the second-level semantic features, the second-level detail features, and the third-level detail features are fused through the second fusion layer to obtain the fused feature H2. The H2 process is as follows: <h2 style=";text-align:left;direction:ltr">H2 = C<h2 style=";text-align:left;direction:ltr"> cat <h2 style=";text-align:left;direction:ltr"> (A2, A3, C<h2 style=";text-align:left;direction:ltr"> RB <h2 style=";text-align:left;direction:ltr"> (H1)) Among them C RB Represents the residual block feature mapping relationship; The third-level semantic feature extraction layer is composed of the second fusion layer followed by three residual blocks. The third-level semantic features are extracted. Then, the third-level semantic features, the first-level detail features, and the second-level detail features are fused through the third fusion layer to obtain the fused feature H3. The H3 process is as follows: <h2 style=";text-align:left;direction:ltr">H3=c<h2 style=";text-align:left;direction:ltr"> cat <h2 style=";text-align:left;direction:ltr"> (A1, A2, C<h2 style=";text-align:left;direction:ltr"> RB <h2 style=";text-align:left;direction:ltr"> (H2)) The fused feature H3 is passed through a 1×1 convolutional layer to generate a predicted crowd density map; The predicted crowd density map and the real crowd density map are compared by using the Euclidean loss function to obtain the network loss size, and the network parameters are continuously updated by using back propagation until suitable network parameters are obtained, and the final network model is obtained by training; The crowd image for which the number of pedestrians needs to be estimated is input into the trained network model to obtain a crowd density map, which is then integrated pixel by pixel, that is, the values of all pixels in the map are added up to obtain an estimate of the number of pedestrians in the image, that is, the predicted total number of people.
2. The crowd counting method based on multi-level feature fusion according to claim 1 is characterized in that: The real population density map is expressed as: Among them, N represents the number of people marked in the crowd image, x i represents the 2D coordinate position of the center point of the head of the i-th person, x represents the position of other pixels in the crowd image except the head, δ(x i ) represents the impulse function, The standard deviation is σ i The adaptive Gaussian kernel of Where β is the weight parameter, Represents x i The average distance between the heads and their nearest neighbors.
3. The crowd counting method based on multi-level feature fusion according to claim 1 is characterized in that: The method also includes a data enhancement step, specifically: firstly, grayscale the images in the data set, then randomly crop 9 sub-sample images of one-quarter the size of the original image from each grayscale image, and randomly rotate and flip them as training sets.
4. The crowd counting method based on multi-level feature fusion according to claim 1 is characterized in that: The network parameters are continuously updated by back propagation to train the final network model. The specific steps include: Based on the pytorch deep learning framework training, the network training error is obtained using the Euclidean loss function, which is defined as: Among them, M represents the size of the training set used, θ represents the network parameters, and X k represents the kth image input, D k , D(X k ; θ) represent the image X k The real crowd density map and the predicted crowd density map generated after network training; Update the network model parameters through back propagation and gradient descent; Set the appropriate size of training rounds. When the training rounds reach the set size, the training ends and the final network model parameters are obtained.
5. A crowd counting device based on multi-level feature fusion, characterized in that: It includes data processing module, multi-level feature fusion network construction module, network training module and pedestrian counting module; The data processing module is used for data processing: generating a real crowd density map according to the images in the public data set for crowd counting and the corresponding two-dimensional coordinate marks of the center points of human heads, wherein the real crowd density map is used for network training; The multi-level feature fusion network construction module is used to construct a multi-level feature fusion network, input a crowd image for which pedestrian quantity estimation is required into the multi-level feature fusion network, and output a predicted crowd density map; The multi-level feature fusion network is divided into two parts, a front-end and a back-end. The front-end is used to extract detail features, and the front-end is divided into three detail feature extraction layers of different levels according to the network depth to extract detail features of different levels. The back-end is used to extract semantic features, and the back-end is divided into three semantic feature extraction layers of different levels according to the network depth to extract semantic features of different levels. Detail features of different levels are fused into semantic features of different levels to different degrees, and finally a predicted crowd density map is obtained through a convolution operation. The network training module is used to obtain a network error by applying a Euclidean loss function to the real crowd density map and the predicted crowd density map, and then update the network parameters by back propagation to obtain a final network model through training; The pedestrian counting module is used to input a crowd image for which pedestrian number estimation is required into the final network model, output a crowd density map, and then perform pixel-by-pixel integration on the map, i.e., sum the values of all pixels in the map to obtain an estimated value of the number of pedestrians in the image, i.e., the predicted total number of people; The front end of the multi-level feature fusion network is based on the VGG16 network, and is divided into three different levels of detail feature extraction layers according to the network depth. The first level detail feature extraction layer is the 1st to 4th layers of the VGG16 network, and the first level detail feature A1 is extracted. The second level detail feature extraction layer is the 1st to 7th layers of the VGG16 network, and the second level detail feature A2 is extracted. The third level detail feature extraction layer is the 1st to 10th layers of the VGG16 network, and the third level detail feature A3 is extracted. The backend part is composed of 3 fusion layers and 6 residual blocks, and is divided into three different levels of semantic feature extraction layers according to the network depth. The first level semantic feature extraction layer is composed of a residual block followed by the front-end network, and the first level semantic feature B1 is extracted. Then, the first level detail feature A1 extracted by the first level detail feature extraction layer is fused with the first level semantic feature B1 extracted by the first level semantic feature extraction layer through the first fusion layer to obtain the fused feature H1. The fusion layer is a channel splicing operation. The H1 process is as follows: H1=C cat (A1,B1) Among them C cat Represents the channel splicing operation; the second-level semantic feature extraction layer is composed of the first fusion layer followed by two residual blocks, which extracts the second-level semantic features. Then, the second-level semantic features, the second-level detail features, and the third-level detail features are fused through the second fusion layer to obtain the fused feature H2. The H2 process is as follows: <h2 style=";text-align:left;direction:ltr">H2 = C<h2 style=";text-align:left;direction:ltr"> cat <h2 style=";text-align:left;direction:ltr"> (A2, A3, C<h2 style=";text-align:left;direction:ltr"> RB <h2 style=";text-align:left;direction:ltr"> (H1)) Among them C RB Represents the residual block feature mapping relationship; The third-level semantic feature extraction layer is composed of the second fusion layer followed by three residual blocks. The third-level semantic features are extracted. Then, the third-level semantic features, the first-level detail features, and the second-level detail features are fused through the third fusion layer to obtain the fused feature H3. The H3 process is as follows: <h2 style=";text-align:left;direction:ltr">H3=C<h2 style=";text-align:left;direction:ltr"> cat <h2 style=";text-align:left;direction:ltr"> (A1, A2, C<h2 style=";text-align:left;direction:ltr"> RB <h2 style=";text-align:left;direction:ltr"> (H2)) The fused feature H3 is passed through a 1×1 convolutional layer to generate a predicted crowd density map.
6. A computer device comprising a memory and a processor, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the processor executes the steps of the crowd counting method based on multi-level feature fusion as described in any one of claims 1 to 4.
7. A storage medium storing computer-readable instructions, wherein when the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the crowd counting method based on multi-level feature fusion as described in any one of claims 1 to 4.