Asymmetric bilateral network crowd counting method based on scale and background awareness
By proposing asymmetric bilateral network based on scale and background perception in population counting, the problems of scale change and background noise processing difficulties in the prior art are solved, and higher counting accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202111553027.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-12-17
Smart Images

Figure CN114241417B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to an asymmetric bilateral network crowd counting method based on scale and background perception. Background Art
[0002] Crowd counting is a basic task in many public security surveillance systems, which aims to estimate the number of people in a still image. Many studies have been devoted to solving this problem and have made some progress. Traditional methods solve the crowd counting problem by using hand-crafted features to detect each person and predict the number of people through regression or density estimation. The performance of these methods is usually low due to the insufficient semantic representation of hand-crafted features. Recently, CNN-based methods have dominated crowd counting thanks to the powerful feature representation of CNN. According to the different types of network structures, CNN-based crowd counting models can be divided into two categories: single-column based methods and multi-column based methods.
[0003] However, there are still some challenges that hinder the computer vision community from designing models that can perform accurate and robust crowd counting, such as occlusion, complex background, scale variation, non-uniform distribution, perspective distortion, rotation, illumination variation, and weather variation. Scale variation is the most dominant problem in crowd counting models. In the era of deep learning, researchers have devoted themselves to integrating semantic features of different scales to solve this problem. The main algorithmic frameworks can be divided into three categories: 1) Multi-column. A "multi-column" structure is adopted, in which each branch has a different filter kernel size to handle a specific scale. 2) Non-standard convolution. Some non-standard convolution operations, such as dilated convolution or deformed convolution, are used to model multi-scale information. 3) Feature Pyramid Network (FPN). It is assumed that features at different levels can capture different scale information, and FPN is used to fuse multi-level features. In order to eliminate the noise caused by background clutter, semantic segmentation or visual attention operations are two common methods to suppress the response of background regions. These methods guide the network to focus on individual instances through mask images. However, all the above methods usually require millions of additional tobe-learned parameters (e.g., learnable multiple columns, learnable segmentation / attention branches), which leads to higher computational burden. Summary of the invention
[0004] In view of the problems existing in the above-mentioned prior art, the present invention provides an asymmetric bilateral network based on scale and background perception for crowd counting and a method for using the asymmetric bilateral network for crowd counting, which can handle scale changes and background noise in a unified framework. The technical solution is as follows:
[0005] In a first aspect, the present application provides a crowd counting method based on an asymmetric bilateral network of scale and background perception, characterized in that the bilateral network has a first network and a second network forming an asymmetric bilateral network, and the crowd counting method includes:
[0006] Based on the deep feature data of the crowd image to be analyzed, the first network is input to obtain the scale perception feature of the crowd image;
[0007] Inputting shallow feature data of the crowd image to be analyzed into the second network to obtain background perception features of the crowd image;
[0008] The scale-aware features and the background-aware features are fused, and the background-aware features in the scale-aware features are suppressed by using an attention mechanism to obtain suppressed scale-aware features;
[0009] A density map is generated by a first regression algorithm based on the suppressed post-scale perceptual features.
[0010] In an embodiment of the present application, data that can characterize the background noise in the image is obtained through shallow feature data of the crowd image to be analyzed, which is used to suppress and remove the background noise of the image in the future. The first network and the second network use features from different semantic layers respectively. The two asymmetric branches of the first network and the second network have different structures. The first network is a densely connected stacked dilated convolution (DCSDC) subnetwork, and each dilated convolution layer has a different dilation rate, which depends on a deep feature data and can handle scale changes. Another branch, the second network, is a parameter-free densely connected stacked pool (DCSP) subnetwork, and each pooling layer has a different pool kernel and step size. It depends on shallow features and can fuse features with multiple receptive fields to reduce the impact of background noise. The outputs of the two networks are fused through the attention mechanism to generate a final density map.
[0011] In one embodiment, the first network includes multiple dilated convolutional layers with the same kernel size and different dilation rates, and the multiple dilated convolutional layers are cascaded in a densely connected manner.
[0012] In one embodiment, the number of dilated convolutional layers in the first network is 3.
[0013] In one embodiment, the dense connection mode includes:
[0014] Assume that the deep feature data of the crowd image to be analyzed is D I , the nonlinear functions of the three dilated convolutional layers are h1(·), h2(·), and h3(·), then the output of the first dilated convolutional layer is:
[0015] H 1 =[D1,h1(D1)];
[0016] The output of the second dilated convolutional layer is:
[0017] H 2 =[D1,h2(H 1 ),H 1 ];
[0018] The output of the third dilated convolutional layer is the scale-aware feature f of the crowd image. s for:
[0019] f s =[D1,h3(H 2 ),H 1 ,H 2 ].
[0020] In one embodiment, the kernel size of the dilated convolutional layer in the first network is 3×3, and the dilation rates are 1, 2 and 3 respectively.
[0021] In the embodiment of the present application, the output feature map sizes of each dilated convolutional layer in the first network are
[0022] where c i is the number of channels, i = 1, 2, ..., 8, d j is the number of steps, j = 1, 2, ..., 5.
[0023] In one embodiment, the second network comprises a plurality of maximum pooling layers with different pooling kernels and different step sizes, and the plurality of maximum pooling layers are cascaded in a densely connected manner.
[0024] In one embodiment, the maximum pooling layer in the second network has 3 layers.
[0025] In one embodiment, the dense connection mode includes:
[0026] Assume that the shallow feature data of the crowd image to be analyzed is S I , the nonlinear functions of the three maximum pooling layers are g1(·), g2(·), and g3(·), then the output of the first maximum pooling layer is:
[0027] G 1 =g1(S1);
[0028] The output of the second max pooling layer is:
[0029] G 2 =[g2(G 1 ),G 1 ];
[0030] The output of the third maximum pooling layer is the background perception feature f of the crowd image b for:
[0031] f b =[g3(G 2 ),G 1 ,G 2 ].
[0032] In one embodiment, in the second network:
[0033] The pooling kernel of the first and second maximum pooling layers is 2×2, and the stride is 2;
[0034] The pooling kernel of the third maximum pooling layer is 3×3 and the stride is 1.
[0035] In the embodiment of the present application, the output feature map sizes of each maximum pooling layer in the second network are
[0036] where c k is the number of channels, k = 1, 2, ..., 8, d j is the number of steps, j = 1, 2, ..., 5.
[0037] In one embodiment, the deep feature data and shallow feature data of the crowd image to be analyzed are obtained through a CNN module, and the acquisition method includes:
[0038] The crowd image to be analyzed is input into the CNN module, and the shallow feature data is output based on the previous CNN layer. Based on the output of deep feature data from the subsequent CNN layers Among them, d1 and d2 are the output strides, c1 and c2 are the channels of the feature map.
[0039] In one embodiment, the density map is generated by a first regression algorithm based on the suppressed post-scale perception feature, and the first regression algorithm is implemented by a density map regression head module, and the density map regression head module includes:
[0040] There are three cascaded convolutional layers, and each convolutional layer is connected to an activation layer. The convolution kernel sizes of the convolutional layers are 3×3, 3×3, and 1×1 respectively, and the activation functions of the activation layers are all Relu functions.
[0041] In the embodiment of the present application, the input feature map size of the density map regression head module is where c s is the number of channels of the scale-aware feature map, and the output feature map size of the density map regression head module is
[0042] In one embodiment, the asymmetric bilateral network, during training, further comprises:
[0043] The background perception features of the crowd image output by the second network are used to generate a foreground mask image through a second regression algorithm, and the second regression algorithm is implemented by a foreground mask regression head module, and the foreground mask regression head module includes:
[0044] There are three cascaded convolutional layers, and each convolutional layer is connected to an activation layer. The convolution kernel sizes of the convolutional layers are 3×3, 3×3, and 1×1, respectively, and the activation functions of the activation layers are Relu function, Relu function, and Sigmoid function, respectively.
[0045] In the embodiment of the present application, the input feature map size of the foreground mask regression head module is where c b is the number of channels of the scale-aware feature map, and the output feature map size of the foreground mask regression head module is
[0046] In one embodiment, the loss function used in training the asymmetric bilateral network is:
[0047] Loss = L c +λ1L ot +λ2L tv +λ3L b ,in:
[0048]
[0049]
[0050]
[0051] L b (F i P ,F i GT )=-F i GT logF i P +(F i GT -1)log(1-F i P );
[0052] Among them, D i GT D is the labeled data of the density map during asymmetric bilateral network training. i P The predicted output data of the density map output during the training of the asymmetric bilateral network;
[0053] F i GTis the annotation data of the foreground mask image during the training of the asymmetric bilateral network, F i P is the predicted output data of the foreground mask image during asymmetric bilateral network training, D i GT , D i P 、F i GT 、F i P The subscript i in represents the i-th image;
[0054] L c is the counting loss, L ot is the optimal transmission loss, L tv is the total change loss, L b Split loss for BCE;
[0055] ||*||1 represents the norm of the vector, α * and β * is the solution of the Monge-Kantarovich optimal transmission formula, and λ1, λ2, and λ3 are hyperparameters.
[0056] In a second aspect, the present application provides an asymmetric bilateral network based on scale and context perception for crowd counting, comprising: a first network and a second network forming an asymmetric bilateral network,
[0057] The input of the first network is used to receive deep feature data of the crowd image to be analyzed;
[0058] The input of the second network is used to receive shallow feature data of the crowd image to be analyzed;
[0059] The first network includes a plurality of dilated convolutional layers with the same kernel size and different dilation rates, wherein the plurality of dilated convolutional layers are cascaded in a densely connected manner, and the second network includes a plurality of maximum pooling layers with different pooling kernels and different step sizes, wherein the plurality of maximum pooling layers are cascaded in a densely connected manner;
[0060] The output of the first network and the output of the second network are connected with an attention mechanism unit for fusing scale and background features to generate a density map.
[0061] In one embodiment, the number of layers of the dilated convolutional layer is 3, the number of layers of the maximum pooling layer is 3, and the dense connection mode of the first network includes:
[0062] The output end of the deep feature data is simultaneously connected to the input end of the first dilated convolutional layer, the output end of the first dilated convolutional layer, the output end of the second dilated convolutional layer, and the output end of the third dilated convolutional layer.
[0063] The output end of the first dilated convolutional layer is simultaneously connected to the input end of the second dilated convolutional layer, the output end of the second dilated convolutional layer, and the output end of the third dilated convolutional layer.
[0064] The output end of the second dilated convolutional layer is simultaneously connected to the input end and the output end of the third dilated convolutional layer;
[0065] The dense connection mode of the second network includes:
[0066] The output end of the shallow feature data is connected to the first maximum pooling layer;
[0067] The output of the first maximum pooling layer is simultaneously connected to the input end of the second maximum pooling layer, the output end of the second maximum pooling layer, and the output end of the third maximum pooling layer;
[0068] The output end of the second maximum pooling layer is simultaneously connected to the input end of the third maximum pooling layer and the output end of the third maximum pooling layer.
[0069] In one embodiment, the asymmetric bilateral network also includes: a CNN module, the CNN module includes a first output terminal and a second output terminal, the first output terminal is output from the first CNN layer in the CNN module, the second output terminal is output from the last CNN layer in the CNN module, the first output terminal is connected to the input terminal of the second network, and the second output terminal is connected to the input terminal of the first network.
[0070] In one embodiment, the asymmetric bilateral network further includes: a density map regression head module connected to the output end of the attention mechanism unit, the density map regression head module including:
[0071] There are three cascaded convolutional layers, and each convolutional layer is connected to an activation layer. The convolution kernel sizes of the convolutional layers are 3×3, 3×3, and 1×1 respectively, and the activation functions of the activation layers are all Relu functions.
[0072] In one embodiment, during training, the asymmetric bilateral network further comprises: a foreground mask regression head module connected to the output end of the second network, the foreground mask regression head module comprising:
[0073] There are three cascaded convolutional layers, and each convolutional layer is connected to an activation layer. The convolution kernel sizes of the convolutional layers are 3×3, 3×3, and 1×1, respectively, and the activation functions of the activation layers are Relu function, Relu function, and Sigmoid function, respectively.
[0074] In one embodiment, the loss function used in training the asymmetric bilateral network is:
[0075] Loss = L c +λ1L ot +λ2L tv +λ3L b ,in:
[0076] L c (D i P ,D i GT )=|||D i P ||1-|||D i GT ||1||;
[0077]
[0078]
[0079] L b (F i P ,F i GT )=-F i GT logF i P +(F i GT -1)log(1-F i P );
[0080] Among them, D i GT D is the labeled data of the density map during asymmetric bilateral network training. i P The predicted output data of the density map output during the training of the asymmetric bilateral network;
[0081] F i GT is the annotation data of the foreground mask image during the training of the asymmetric bilateral network, F i P is the predicted output data of the foreground mask image during asymmetric bilateral network training, D i GT , D i P 、F i GT 、F i P The subscript i in represents the i-th image;
[0082] L c is the counting loss, L otis the optimal transmission loss, L tv is the total change loss, L b Split loss for BCE;
[0083] ||*1 represents the norm of the vector, α * and β * is the solution of the Monge-Kantarovich optimal transmission formula, and λ1, λ2, and λ3 are hyperparameters.
[0084] The crowd counting method of the present invention based on an asymmetric bilateral network with scale and background perception has the following beneficial effects:
[0085] (1) A novel asymmetric bilateral network is proposed to handle scale variation and background noise in a unified framework. A novel first network (DCSDC sub-network) is proposed to extract multi-scale information based on a deep feature layer. A novel second network (DCSP sub-network) is proposed to fuse features from multiple receptive fields to reduce the impact of background noise without any additional learnable parameters.
[0086] (2) In the present invention, data that can characterize the background noise in the image is obtained through the shallow feature data of the crowd image to be analyzed, which is used to suppress and remove the background noise of the image in the subsequent process. The first network and the second network use features from different semantic layers respectively. The two asymmetric branches of the first network and the second network have different structures. The first network is a densely connected stacked dilated convolution (DCSDC) subnetwork, each dilated convolution layer has a different dilation rate, and relies on a deep feature data to handle scale changes. The other branch, the second network, is a parameter-free densely connected stacked pooling (DCSP) subnetwork, each pooling layer has a different pool kernel and step size, which relies on shallow features and can fuse features with multiple receptive fields to reduce the impact of background noise. The outputs of the two networks are fused through the attention mechanism to generate the final density map. A full comparison experiment was conducted on multiple widely used large-scale crowd counting datasets with multiple crowd counting algorithms in the prior art. The experimental results show that the method proposed in this application can effectively improve the accuracy of dense crowd counting algorithms. BRIEF DESCRIPTION OF THE DRAWINGS
[0087] Figure 1 is a flow chart of an asymmetric bilateral network crowd counting method based on scale and context perception in an embodiment of the present application;
[0088] Figure 2 is a flow chart of an asymmetric bilateral network based on scale and background perception during training in an embodiment of the present application;
[0089] Figure 3-1is a structural diagram of an asymmetric bilateral network SBAB based on scale and context perception in an embodiment of the present application; Figure 3-2 is the structure of the density map regression head module in the SBAB network in the embodiment of the present application, Figure 3-3 It is the structure of the foreground mask regression head module in the SBAB network in the embodiment of the present application;
[0090] Figure 4 are some crowd image samples in the dataset used in the experiment;
[0091] Figure 5 It is the change of loss function value at different training times in the process of training SBAB network model using part A of ShanghaiTech dataset in the experiment;
[0092] Figure 6 It is the comparison result of the density map generated by the existing DM Count crowd counting algorithm and the SBAB network of this application in the experiment;
[0093] Figure 7 It is the MAE and MSE performance comparison result of the DCSP in this application and the SP in the prior art on the ShanghaiTech dataset part A dataset. DETAILED DESCRIPTION
[0094] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be interpreted as limiting the present invention.
[0095] It will be understood by those skilled in the art that, unless expressly stated, the singular forms "one", "said", and "the" used herein may also include plural forms. It should be further understood that the term "comprising" used in the specification of the present application refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be an intermediate element. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The term "and / or" used herein includes all or any unit and all combinations of one or more associated listed items.
[0096] For the crowd counting task, the embodiment of the present application proposes an asymmetric bilateral network based on scale and background perception, the asymmetric bilateral network includes: a first network and a second network forming the asymmetric bilateral network,
[0097] The input of the first network is used to receive the deep feature data of the crowd image to be analyzed, and the first network is used to output the scale perception feature f of the crowd image to be analyzed. s ;
[0098] The input of the second network is used to receive shallow feature data of the crowd image to be analyzed, and the second network is used to output background perception features f of the crowd image to be analyzed. b ;
[0099] The first network includes a plurality of dilated convolutional layers with the same kernel size and different dilation rates, wherein the plurality of dilated convolutional layers are cascaded in a densely connected manner, and the second network includes a plurality of maximum pooling layers with different pooling kernels and different step sizes, wherein the plurality of maximum pooling layers are cascaded in a densely connected manner;
[0100] The output of the first network and the output of the second network are connected with an attention mechanism unit for fusing scale and background features to generate a density map. The attention mechanism unit adopts an element-based multiple attention mechanism, which refines features based on features. Specifically, the mechanism is based on the background perception feature f of the crowd image. b Refine the scale-aware features f s , the formula is as follows: According to this formula, f s Features in the background area will be suppressed.
[0101] Among them, the number of dilated convolutional layers of the first network is 3. The three dilated convolutional layers are recorded in order as the first dilated convolutional layer, the second dilated convolutional layer, and the third dilated convolutional layer. The dense connection method of the three dilated convolutional layers is:
[0102] The output end of the deep feature data is simultaneously connected to the input end of the first dilated convolutional layer, the output end of the first dilated convolutional layer, the output end of the second dilated convolutional layer, and the output end of the third dilated convolutional layer.
[0103] The output of the first dilated convolutional layer is connected to the input of the second dilated convolutional layer, the output of the second dilated convolutional layer, and the output of the third dilated convolutional layer.
[0104] The output end of the second dilated convolutional layer is connected to the input end and output end of the third dilated convolutional layer at the same time;
[0105] On this basis, it is assumed that the deep feature data of the crowd image to be analyzed is D I, the nonlinear functions of the three dilated convolutional layers are h1(·), h2(·), and h3(·), then the output of the first dilated convolutional layer is:
[0106] H 1 =[D1,h1(D1)];
[0107] The output of the second dilated convolutional layer is:
[0108] H 2 =[D1,h2(H 1 ),H 1 ];
[0109] The output of the third dilated convolutional layer is the scale-aware feature f of the crowd image. s for:
[0110] f s =[D1,h3(H 2 ),H 1 ,H 2 ].
[0111] That is, the first network (DCSDC) implements a nonlinear mapping function H(·) based on the deep feature D I is input and generates a scale-aware feature f s , that is, f s =H(D I ).
[0112] Among them, the number of maximum pooling layers in the second network is 3. The three maximum pooling layers are recorded as the first maximum pooling layer, the second maximum pooling layer, and the third maximum pooling layer in order. Then the dense connection method of the three dilated convolutional layers is:
[0113] The output end of the shallow feature data is connected to the first maximum pooling layer;
[0114] The output of the first maximum pooling layer is simultaneously connected to the input end of the second maximum pooling layer, the output end of the second maximum pooling layer, and the output end of the third maximum pooling layer;
[0115] The output end of the second maximum pooling layer is connected to the input end and the output end of the third maximum pooling layer at the same time.
[0116] On this basis, it is assumed that the shallow feature data of the crowd image to be analyzed is S I , the nonlinear functions of the three maximum pooling layers are g1(·), g2(·), and g3(·), then the output of the first maximum pooling layer is:
[0117] G 1 =g1(S1);
[0118] The output of the second max pooling layer is:
[0119] G 2 =[g2(G 1 ),G 1 ];
[0120] The output of the third maximum pooling layer is the background perception feature f of the crowd image b for:
[0121] f b =[g3(G 2 ),G 1 ,G 2 ].
[0122] That is, the second network (DCSP) implements a nonlinear mapping function G(·) based on the shallow feature S I is input and generates a background-aware feature f b , that is, f b =G(S I ).
[0123] In the embodiment of the present application, the first network and the second network respectively use features from different semantic layers, and the two asymmetric branches of the first network and the second network have different structures. The first network is recorded as a densely connected stacked dilated convolution (DCSDC) subnetwork, and each dilated convolution layer has a different dilation rate, which depends on a deep feature data and can handle scale changes. Another branch, the second network, is recorded as a parameter-free densely connected stacked pool (DCSP) subnetwork, and each pooling layer has a different pool kernel and step size. It depends on shallow features and can fuse features with multiple receptive fields to reduce the impact of background noise. Scale changes and background noise are processed simultaneously in a unified framework. In the DCSDC subnetwork and the DCSP subnetwork, dense connections are used to fuse features with multiple receptive fields, and shallow features are used to suppress background noise, thereby improving the accuracy of crowd counting.
[0124] When the asymmetric bilateral network is used to perform a crowd counting task, a method for processing a crowd image to be analyzed by the first network and the second network includes the following steps:
[0125] Based on the deep feature data of the crowd image to be analyzed, the first network is input to obtain the scale perception feature f of the crowd image. s ;
[0126] Based on the shallow feature data of the crowd image to be analyzed, the second network is input to obtain the background perception feature f of the crowd image. b ;
[0127] The scale-aware features and the background-aware features are fused, and the background-aware features in the scale-aware features are suppressed by using an attention mechanism to obtain suppressed scale-aware features;
[0128] A density map is generated by a first regression algorithm based on the suppressed post-scale perceptual features.
[0129] Furthermore, the kernel size of the dilated convolutional layer in the first network is 3×3, and the dilation rates are 1, 2, and 3 respectively. The output feature map sizes of each dilated convolutional layer in the first network are:
[0130] where c i is the number of channels, i = 1, 2, ..., 8, d j is the number of steps, j = 1, 2, ..., 5, H is the image height, and W is the image width.
[0131] Correspondingly, since the feature dimensions of each convolutional layer should remain consistent in the first network DCSDC, the corresponding pooling sizes are also set to 1, 2, and 3 respectively.
[0132] The pooling kernels of the first and second maximum pooling layers in the second network are 2×2 with a step size of 2; the pooling kernel of the third maximum pooling layer is 3×3 with a step size of 1.
[0133] The output feature map sizes of each maximum pooling layer in the second network are:
[0134] where c k is the number of channels, k = 1, 2, ..., 8, d j is the number of steps, j = 1, 2, ..., 5.
[0135] Furthermore, the asymmetric bilateral network for crowd counting in the present application also includes: a CNN module, the CNN module includes an input end, a first output end and a second output end, the input end is used to receive a crowd image to be analyzed, the first output end is output from a sequentially earlier CNN layer in the CNN module, the second output end is output from a sequentially later CNN layer in the CNN module, the first output end is connected to the input end of the second network, and the second output end is connected to the input end of the first network.
[0136] When using the asymmetric bilateral network to perform crowd counting tasks, the crowd image to be analyzed is input into the CNN module, and the shallow feature data is output based on the sequentially preceding CNN layer. Based on the output of deep feature data from the subsequent CNN layers Among them, d1 and d2 are the output strides, c1 and c2 are the channels of the feature map.
[0137] Furthermore, the asymmetric bilateral network for crowd counting in the present application further includes: a density map regression head module connected to the output end of the attention mechanism unit, for implementing the first regression algorithm, the density map regression head module includes:
[0138] There are three cascaded convolutional layers, and each convolutional layer is connected to an activation layer. The convolution kernel sizes of the convolutional layers are 3×3, 3×3, and 1×1 respectively, and the activation functions of the activation layers are all Relu functions.
[0139] The input feature map size of the density map regression head module is where c s is the number of channels of the scale-aware feature map, and the output feature map size of the density map regression head module is
[0140] Furthermore, the asymmetric bilateral network for crowd counting in the present application further includes, during training, a foreground mask regression head module connected to the output end of the second network, for implementing a second regression algorithm, wherein the foreground mask regression head module includes:
[0141] There are three cascaded convolutional layers, and each convolutional layer is connected to an activation layer. The convolution kernel sizes of the convolutional layers are 3×3, 3×3, and 1×1, respectively, and the activation functions of the activation layers are Relu function, Relu function, and Sigmoid function, respectively.
[0142] The input feature map size of the foreground mask regression head module is where c b is the number of channels of the scale-aware feature map, and the output feature map size of the foreground mask regression head module is
[0143] Furthermore, the loss function used in the training of the asymmetric bilateral network for crowd counting in this application is:
[0144] Loss = L c +λ1L ot +λ2L tv +λ3L b ,in:
[0145] L c (D i P ,D i GT )=|||D i P ||1-|||D i GT ||1||;
[0146]
[0147]
[0148] L b (F i P ,F i GT )=-F i GT logF i P +(F i GT -1)log(1-F i P );
[0149] Among them, D i GT D is the labeled data of the density map during asymmetric bilateral network training. i P The predicted output data of the density map output during the training of the asymmetric bilateral network;
[0150] F i GT is the annotation data of the foreground mask image during the training of the asymmetric bilateral network, F i P is the predicted output data of the foreground mask image during asymmetric bilateral network training, D i GT , D i P 、F i GT 、F i P The subscript i in represents the i-th image;
[0151] L c is the counting loss, L ot is the optimal transmission loss, L tv is the total change loss, L b Split loss for BCE;
[0152] ||*||1 represents the norm of the vector, α * and β * is the solution of the Monge-Kantarovich optimal transmission formula, and λ1, λ2, and λ3 are hyperparameters.
[0153] In the embodiment of the present application, the method of generating a density map in "Distribution matching for crowd counting" is followed, and no Gaussian function is required to preprocess the labeled data D of the density map. i GT, let δ(x) represent the number of people marked at position x in the image, then the labeled data D of the density map is i GT for:
[0154] D i GT =δ(x);
[0155] The annotation data F of the foreground mask image i GT for:
[0156] That is, the position x where there is a person in the crowd image to be analyzed is recorded as the foreground area;
[0157] In order to improve the generalization performance model of the algorithm, inspired by "Distribution matching for crowdcounting", the counting loss, the best transfer loss and the total change loss are used to measure the labeled data D of the density map. i GT And the density map predicts the output data D i P The difference, counting loss L c for:
[0158] L c (D i P ,D i GT )=|||D i P ||1-|||D i GT ||1||;
[0159] Optimum transmission loss L ot for:
[0160]
[0161] Total change loss L tv for:
[0162]
[0163] In order to measure the foreground mask annotation data F i GT And the predicted output data F of the foreground mask image i P The difference between the BCE segmentation loss L b for:
[0164] L b (F i P ,Fi GT )=-F i GT logF i P +(F i GT -1)log(1-F i P );
[0165] The final loss is the weighted sum of the above four losses:
[0166] Loss = L c +λ1L ot +λ2L tv +λ3L b , where the hyperparameters λ1, λ2, and λ3 are used to adjust the weight ratios of different components.
[0167] Based on the scale- and background-aware asymmetric bilateral network including the CNN module, the first network, the second network, the attention mechanism unit, the density map regression head module, and the foreground mask regression head module, denoted as the SBAB network, the processing process of the crowd image to be analyzed includes the following steps:
[0168] The crowd image to be analyzed is input into the CNN module, and the output of the convolutional layer in the latter order of the CNN module is connected to the first network as the second output end, providing the first network with deep feature data of the crowd image, and the output of the convolutional layer in the former order of the CNN module is connected to the second network as the first output end, providing the first network with shallow feature data of the crowd image;
[0169] The first network is a densely connected stacked dilated convolutional subnetwork (DCSDC), which extracts scale-aware features f through three densely connected dilated convolutional layers for the received deep feature data. s The second network is a densely connected stacked pooling subnetwork (DCSP). For the shallow feature data received, the background perception feature f is extracted through a densely connected 3-layer maximum pooling layer. b ;
[0170] The output of the first network f s and the output of the second network f b , using the formula through the attention mechanism unit Implement an element-based multi-attention mechanism to exploit the background-aware features of crowd images. b Suppressing scale-aware features f s The characteristics of the background area in the image are used to obtain the suppressed post-scale perception characteristics;
[0171] Based on the obtained suppressed scale-aware features, a density map is generated through the density map regression head module.
[0172] Of course, the training process of this scale- and background-aware asymmetric bilateral network also includes:
[0173] The output of the second network is connected to the foreground mask regression head module to realize the background perception feature f output by the second network b The background mask image is generated by the second regression algorithm.
[0174] Based on the above scale- and background-aware asymmetric bilateral network, a crowd counting experiment is conducted. The experimental process is as follows:
[0175] A. Experimental environment
[0176] 1) Dataset: The experiments are conducted and evaluated on three common crowd counting benchmarks, including ShanghaiTech large-scale crowd counting dataset, UCF-QNRF large-scale crowd counting dataset, and NWPU-Crowd large-scale crowd counting dataset. Some data samples are as follows: Figure 4 In these datasets, each person in a group image is annotated with a point near the center of the head.
[0177] The ShanghaiTech dataset consists of 1198 annotated crowd images and 330165 annotated crowd images. It is divided into two separate parts: Part A and Part B. Part A contains 482 images collected from the Internet and is divided into training and testing subsets consisting of 300 and 182 images. Part B contains 716 images collected from busy streets and is divided into training and testing subsets consisting of 400 and 316 images.
[0178] The UCF-QNRF dataset consists of 1535 images (1201 for training and 334 for testing) containing more than 1.25 million labeled instances.
[0179] The NWPU-Crowd dataset consists of 5109 images with a total of 213375 annotated heads with various lighting scenarios and density ranges. The dataset is randomly split into training, validation, and testing subsets, containing 3109, 500, and 1500 images respectively.
[0180] 2) Evaluation metrics: Mean absolute error (MAE) and mean square error (MSE) are used to evaluate the performance of different crowd counting methods.
[0181] The calculation formula of MAE is:
[0182]
[0183] The calculation formula of MSE is as follows:
[0184]
[0185] Where N is the number of images in the test set, C i , C i GT are the predicted output value and labeled data of the crowd count of the i-th image respectively.
[0186] B. Implementation Details
[0187] We set the hyperparameters following the methods in the state-of-the-art DM Count crowd counting algorithm. All empirical studies are conducted on an Nvidia GTX 3070 GPU in Pytork. During training, the input size is 256×256 for the ShanghaiTech dataset part A, 384×384 for the NWPU-Crowd dataset part, and 512×512 for the ShanghaiTech dataset part B and UCF-QNRF parts, all of which are randomly cropped from the original images. During testing, the input images of the network remain at the original resolution. The AdamW optimizer is used to train all models, and the learning rate and weight decay are set to 0.00001 and 0.0001, respectively. The output strides and are set to 4 and 8, respectively. The hyperparameters λ1, λ2, and λ3 in the loss function are set to 0.1, 0.01, and 1, respectively.
[0188] C. Ablation experiment (using lightweight network MobileNet-V2 to implement backbone network CNN module)
[0189] 1) Convergence: Figure 5 The SBAB network model proposed in this application is trained using the ShanghaiTech dataset A part data set. During the training process, the loss function value comparison results of different training times are shown. We can observe that during the entire training process, the loss value decreases almost monotonically and converges smoothly. The loss function value of the SBAB network becomes stable after 800 stages.
[0190] 2) Impact of different components: Compared with the prior art DM Count crowd counting algorithm, the SBAB network proposed in this application includes a DCSDC subnetwork and a DCSP subnetwork. In order to verify the impact of the DCSDC subnetwork and the DCSP subnetwork on the SBAB network for crowd counting, the crowd counting effects of two variant networks of the SBAB network on crowd images are analyzed respectively. The first SBAB variant network is the original SBAB network without DCSDC, and the second SBAB variant network is the original SBAB network without the DCSP subnetwork. Table 1 shows the performance comparison of the complete SBAB network, the first SBAB variant network, and the second SBAB variant network in crowd counting on the above three data sets.
[0191]
[0192] Table 1 Performance comparison of different variants of SBAB networks
[0193] As can be seen from Table 1, the complete SBAB network performs best on all datasets, which indicates that both the DCSDC sub-network and the DCSP sub-network help improve the final crowd counting accuracy. At the same time, it can be seen that the second SBAB variant network performs better than the first SBAB variant network in terms of crowd counting, which indicates the importance of the DCSDC sub-network for the SBAB network to learn scale-aware features.
[0194] During the experiment, the effect of the SBAB network proposed in this application on crowd counting was qualitatively evaluated through a pre-set density map, such as Figure 6 As shown in the figure, it can be seen that the density map generated by the prior art DM Count crowd counting algorithm and the density map generated by the SBAB network of the present application have obvious differences in the following two aspects: (1) robustness to noise, as shown in the second and fourth rows. (2) robustness to scaling changes, as shown in the first and third rows.
[0195] In addition, the main difference between the DCSP in this application and the SP introduced in the prior art "S. Huang, X. Li, Z.-Q. Cheng, Z. Zhang, and A. Hauptmann, "Stacked pooling for boosting scale invariance of crowdcounting," in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2020, pp. 2578–2582." is that in the network framework in this application, features from different pool kernels are fused by dense connections. Figure 7 The MAE and MSE performance comparison of the DCSP in this application and the SP in the prior art on the ShanghaiTech dataset part A is shown. From the figure, it can be seen that the performance of DCSP is better than SP. Compared with the DCSP in this application that uses dense connections to fuse features with multiple receptive fields, the channel averaging operation in SP loses a lot of details.
[0196] D. Comparison with existing technologies
[0197] To verify the effectiveness of the SBAB network in this application, we compared the proposed method with 13 state-of-the-art methods in the experiment, namely SFCN, CAN, Bayes+, S-DCNet, SANet+SPANet, SDANet, ADSCNet, ASNet, AMRNet, AMSNet, DM Count, P2PNet and SFANet. The backbone network CNN module of the SBAB network in this application adopts the pre-trained VGG-19.
[0198] Table 2 is the algorithm performance comparison results based on MAE and MSE indicators of the SBAB network proposed in this application and other 13 existing algorithms on the ShanghaiTech dataset A part dataset.
[0199] Table 3 is the algorithm performance comparison results based on MAE and MSE indicators between the SBAB network proposed in this application and other existing algorithms on the NWPU-Crowd dataset.
[0200]
[0201] Table 2 Algorithm comparison results on the ShanghaiTech Part-A dataset
[0202]
[0203]
[0204] Table 3 Algorithm comparison results on the NWPU-Crowd dataset
[0205] From Table 2 and Table 3 we can see that:
[0206] (1) Despite the simple structures of DCSDC and DCSP, the SBAB network proposed in this application can achieve better counting performance than the existing crowd counting algorithms in both small and large data sets. In particular, when the same backbone network is used as a feature extractor, the crowd counting effect of the SBAB network is significantly better than that of DM Count.
[0207] (2) The performance of the asymmetric bilateral network in the SBAB network of this application is better than the commonly used bilateral structure (such as SFANet). This further proves that the features of different layers in the backbone network CNN module have different abilities to model background noise.
[0208] The present invention is not limited to the above-mentioned specific implementation modes. Various changes made by ordinary technicians in this field based on the above-mentioned concepts without creative work are all within the protection scope of the present invention.
Claims
1. A crowd counting method based on scale and context-aware asymmetric bilateral network, characterized in that: The bilateral network has a first network and a second network forming an asymmetric bilateral network, The first network includes a plurality of dilated convolutional layers having the same kernel size and different dilation rates, and the plurality of dilated convolutional layers are cascaded in a densely connected manner; The multiple dilated convolutional layers are densely connected, including: Assume that the deep feature data of the crowd image to be analyzed is D I , the nonlinear functions of the three dilated convolutional layers are h1(·), h2(·), and h3(·), then the output of the first dilated convolutional layer is: H 1 = [D1,h1(D1)]; The output of the second dilated convolutional layer is: H 2 =[D1,h2(H 1 ),H 1 ]; The output of the third dilated convolutional layer is the scale-aware feature f of the crowd image. s for: f s =[D1,h3(H 2 ),H 1 ,H 2 ]; The second network includes a plurality of maximum pooling layers with different pooling cores and different step sizes, and the plurality of maximum pooling layers are cascaded in a densely connected manner; The multiple maximum pooling layers are densely connected, including: Assume that the shallow feature data of the crowd image to be analyzed is S I , the nonlinear functions of the three maximum pooling layers are g1(·), g2(·), and g3(·), then the output of the first maximum pooling layer is: G 1 =g1(S1); The output of the second max pooling layer is: G 2 =[g2(G 1 ),G 1 ]; The output of the third maximum pooling layer is the background perception feature f of the crowd image b for: f b =[g3(G 2 ),G 1 ,G 2 ]; The crowd counting method comprises: Based on the deep feature data of the crowd image to be analyzed, the first network is input to obtain the scale perception feature of the crowd image; Inputting shallow feature data of the crowd image to be analyzed into the second network to obtain background perception features of the crowd image; The scale-aware features and the background-aware features are fused, and the background-aware features in the scale-aware features are suppressed by using an attention mechanism to obtain suppressed scale-aware features; A density map is generated by a first regression algorithm based on the suppressed post-scale perceptual features.
2. The crowd counting method based on scale and context-aware asymmetric bilateral network according to claim 1 is characterized in that: The number of dilated convolutional layers in the first network is 3.
3. The crowd counting method based on scale and context-aware asymmetric bilateral network according to claim 2, characterized in that: The kernel size of the dilated convolutional layer in the first network is 3×3, and the dilation rates are 1, 2 and 3 respectively.
4. The crowd counting method based on scale and context-aware asymmetric bilateral network according to claim 1, characterized in that: The maximum pooling layer in the second network has 3 layers.
5. The crowd counting method based on scale and context-aware asymmetric bilateral network according to claim 4, characterized in that: The second network The pooling kernel of the first and second maximum pooling layers is 2×2, and the stride is 2; The pooling kernel of the third maximum pooling layer is 3×3 and the stride is 1.
6. The crowd counting method based on scale and context-aware asymmetric bilateral network according to claim 1, characterized in that: The deep feature data and shallow feature data of the crowd image to be analyzed are obtained through the CNN module, and the acquisition includes: The crowd image to be analyzed is input into the CNN module, and the shallow feature data is output based on the previous CNN layer. Based on the output of deep feature data from the subsequent CNN layers Among them, d1 and d2 are the output strides, c1 and c2 are the channels of the feature map.
7. The crowd counting method based on scale and context-aware asymmetric bilateral network according to claim 1, characterized in that: The density map is generated by a first regression algorithm based on the suppressed post-scale perception feature, and the first regression algorithm is implemented by a density map regression head module, and the density map regression head module includes: There are three cascaded convolutional layers, and each convolutional layer is connected to an activation layer. The convolution kernel sizes of the convolutional layers are 3×3, 3×3, and 1×1 respectively, and the activation functions of the activation layers are all Relu functions.
8. The crowd counting method based on scale and context-aware asymmetric bilateral network according to claim 1, characterized in that: The asymmetric bilateral network, during training, further includes: The background perception features of the crowd image output by the second network are used to generate a foreground mask image through a second regression algorithm, and the second regression algorithm is implemented by a foreground mask regression head module, and the foreground mask regression head module includes: There are three cascaded convolutional layers, and each convolutional layer is connected to an activation layer. The convolution kernel sizes of the convolutional layers are 3×3, 3×3, and 1×1, respectively, and the activation functions of the activation layers are Relu function, Relu function, and Sigmoid function, respectively.
9. The crowd counting method based on scale and context-aware asymmetric bilateral network according to claim 8, characterized in that: The loss function used in the training of the asymmetric bilateral network is: Loss = L c +λ1L ot +λ2L tv +λ3L b ,in: L b (F i P ,F i GT )=-F i GT logF i P +(F i GT -1)log(1-F i P ); in, It is the labeled data of density map during asymmetric bilateral network training. The predicted output data of the density map output during the training of the asymmetric bilateral network; F i GT is the annotation data of the foreground mask image during the training of the asymmetric bilateral network, F i P It is the predicted output data of the foreground mask image during the training of the asymmetric bilateral network. F i GT 、F i P The subscript i in represents the i-th image; L c is the counting loss, L ot is the optimal transmission loss, L tv is the total change loss, L b Split loss for BCE; *1 represents the norm of the vector, α * and β * is the solution of the Monge-Kantarovich optimal transmission formula, and λ1, λ2, and λ3 are hyperparameters.
10. An asymmetric bilateral network based on scale and context perception for crowd counting, used to execute the crowd counting method based on an asymmetric bilateral network based on scale and context perception according to any one of claims 1 to 9, characterized in that: include: The first network and the second network form an asymmetric bilateral network, The input of the first network is used to receive deep feature data of the crowd image to be analyzed; The input of the second network is used to receive shallow feature data of the crowd image to be analyzed; The first network includes a plurality of dilated convolutional layers with the same kernel size and different dilation rates, and the plurality of dilated convolutional layers are cascaded in a densely connected manner; the second network includes a plurality of maximum pooling layers with different pooling kernels and different step sizes, and the plurality of maximum pooling layers are cascaded in a densely connected manner; The output of the first network and the output of the second network are connected with an attention mechanism unit for fusing scale and background features to generate a density map.
11. The scale- and context-aware asymmetric bilateral network for crowd counting according to claim 10, characterized in that: The number of layers of the dilated convolutional layer is 3, the number of layers of the maximum pooling layer is 3, and the dense connection mode of the first network includes: The output end of the deep feature data is simultaneously connected to the input end of the first dilated convolutional layer, the output end of the first dilated convolutional layer, the output end of the second dilated convolutional layer, and the output end of the third dilated convolutional layer. The output end of the first dilated convolutional layer is simultaneously connected to the input end of the second dilated convolutional layer, the output end of the second dilated convolutional layer, and the output end of the third dilated convolutional layer. The output end of the second dilated convolutional layer is simultaneously connected to the input end and the output end of the third dilated convolutional layer; The dense connection mode of the second network includes: The output end of the shallow feature data is connected to the first maximum pooling layer; The output of the first maximum pooling layer is simultaneously connected to the input end of the second maximum pooling layer, the output end of the second maximum pooling layer, and the output end of the third maximum pooling layer; The output end of the second maximum pooling layer is simultaneously connected to the input end of the third maximum pooling layer and the output end of the third maximum pooling layer.
12. The scale- and context-aware asymmetric bilateral network for crowd counting according to claim 11, characterized in that: The asymmetric bilateral network also includes: a CNN module, the CNN module includes a first output end and a second output end, the first output end is output from the first CNN layer in the CNN module, the second output end is output from the later CNN layer in the CNN module, the first output end is connected to the input end of the second network, and the second output end is connected to the input end of the first network.
13. The scale- and context-aware asymmetric bilateral network for crowd counting according to claim 12, characterized in that: The asymmetric bilateral network further includes: a density map regression head module connected to the output end of the attention mechanism unit, and the density map regression head module includes: There are three cascaded convolutional layers, and each convolutional layer is connected to an activation layer. The convolution kernel sizes of the convolutional layers are 3×3, 3×3, and 1×1 respectively, and the activation functions of the activation layers are all Relu functions.
14. The scale- and context-aware asymmetric bilateral network for crowd counting according to claim 13, characterized in that: During training, the asymmetric bilateral network further includes: a foreground mask regression head module connected to the output end of the second network, wherein the foreground mask regression head module includes: There are three cascaded convolutional layers, and each convolutional layer is connected to an activation layer. The convolution kernel sizes of the convolutional layers are 3×3, 3×3, and 1×1, respectively, and the activation functions of the activation layers are Relu function, Relu function, and Sigmoid function, respectively.
15. The scale- and context-aware asymmetric bilateral network for crowd counting according to claim 14, characterized in that: The loss function used in the training of the asymmetric bilateral network is: Loss = L c +λ1L ot +λ2L tv +λ3L b ,in: L b (F i P ,F i GT )=-F i GT logF i P +(F i GT -1)log(1-F i P ); in, It is the labeled data of density map during asymmetric bilateral network training. The predicted output data of the density map output during the training of the asymmetric bilateral network; F i GT is the annotation data of the foreground mask image during the training of the asymmetric bilateral network, F i P It is the predicted output data of the foreground mask image during the training of the asymmetric bilateral network. F i GT 、F i P The subscript i in represents the i-th image; L c is the counting loss, L ot is the optimal transmission loss, L tv is the total change loss, L b Split loss for BCE; *1 represents the norm of the vector, α * and β * is the solution of the Monge-Kantarovich optimal transmission formula, and λ1, λ2, and λ3 are hyperparameters.
Citation Information
Patent Citations
Video crowd counting system and method
CN111860162A
Cross-view image generation method based on asymmetric convolutional network and attention mechanism
CN112884893A