A crowd counting method based on multi-scale spatial guided perception aggregation network
Through the crowd counting method based on a multi-scale spatially guided perceptual aggregation network, the problems of scale changes, insufficient perception, complex background interference and poor feature fusion in crowd count are solved, and more accurate and robust crowd counting results are achieved.
Patent Information
- Application Number
- CN202210451241.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-24
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-04-24
AI Technical Summary
The prior art has problems in population counting with scale changes, insufficient perception, complex background interference and poor dissemination of feature fusion information.
A population counting method based on multi-scale spatially guided perceptual aggregation network (MGANet) is proposed. The multi-scale feature extraction network, spatially guided network and attention fusion network are extracted through multi-scale feature, capture multi-scale information, establish spatial context relationships, and fusion feature information, and optimize the counting results through adaptive scale loss function.
It improves the accuracy and robustness of population counting, alleviates the problems of scale changes and complex background interference, and enhances the ability to capture multi-scale and spatial context information.
Smart Images

Figure CN114694102B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a crowd counting method based on a multi-scale space guided perception aggregation network. Background Art
[0002] With the increasing concentration of urban population around the world, crowd counting and recognition technology based on computer vision plays an important role in public safety, abnormal event detection, urban traffic management, etc. In sparse scenes containing single or multiple targets in the image, crowd counting and recognition can be easily and accurately performed through target positioning detection technology. Due to the gradual development of deep learning, methods based on convolutional neural networks (CNNs) have achieved remarkable success in tasks such as image classification, pedestrian detection, and speech recognition. Therefore, researchers introduced CNN into the field of crowd counting, and achieved good crowd density estimation results in learning the mapping between images and density maps. In recent years, researchers have designed a variety of CNN-based crowd counting algorithms to overcome challenges such as scale changes, non-uniform distribution, occlusion, and complex background. These algorithms mainly include single-branch network models, multi-branch network models, attention mechanisms, and feature fusion methods.
[0003] The existing technologies have the following four problems in crowd recognition and counting:
[0004] 1. Since the scale change problem in the crowd counting task will affect the counting results, many solutions based on the traditional VGG16 and ResNet101 models still have defects.
[0005] 2. In crowd counting, many scenarios need to consider the positional relationship between spatial objects. However, capturing detailed features is not enough, and there is a problem of insufficient perception in this task.
[0006] 3. Complex background interference is also an important factor affecting the accuracy of crowd counting. Previous models solved the above problem through self-attention mechanism or image segmentation, but the model complexity is high and the computational cost required is high.
[0007] 4. When fusing features, simple fusion will weaken the effectiveness of information dissemination. In the field of crowd counting, it is especially important to pay attention to spatial position, high-level semantic information and background interference.
[0008] Therefore, in view of the defects of the existing crowd counting methods, the present invention proposes a crowd counting method based on a multi-scale spatial guided perception aggregation network (MGANet) to solve the above problems existing in the prior art. Summary of the invention
[0009] In view of the shortcomings of the prior art, the present invention proposes a crowd counting method based on a multi-scale spatial guided perception aggregation network (MGANet), which uses a reasonable and efficient guidance method to aggregate multi-scale information of the adaptively captured spatial environment to improve the accuracy and robustness of counting.
[0010] In order to solve the above technical problems, the technical solution of the present invention is:
[0011] A crowd counting method based on a multi-scale spatial guided perception aggregation network comprises the following steps:
[0012] S1. Establish a multi-scale feature extraction network
[0013] The multi-scale feature extraction network is based on the Inception-v3 model, removing the first two maximum pooling layers and the fully connected layer of Inception-v3, and retaining all convolutional layers in the Inception-v3 model.
[0014] The network layers included in the multi-scale feature extraction network are: five convolutional layers, three Inception-A, one Inception-D, four Inception-B, one Inception-E and two Inception-C;
[0015] S2. Input images of any resolution into the multi-scale feature extraction network
[0016] Given an image input I, the output features are represented by mapping:
[0017] x=F MFEN
[0018] Where x represents the multi-scale features captured by the image, and F MFEN represents a multi-scale feature extraction network;
[0019] S3, inputting the multi-scale features captured by the multi-scale feature extraction network into the spatial guidance network, and outputting the context-guided perception features and the guided perception map;
[0020] S4, passing the context-guided perception features and the guided perception map to the attention fusion network, finally outputting the density map and constructing the density map training set;
[0021] S5. Establish an adaptive scale loss function and perform adaptive training through the density map training set;
[0022] S6. Take the image to be predicted as input, repeat steps S2-S5, and output the calculation result of the crowd in the image to be predicted.
[0023] Preferably, the spatial guidance network includes a spatial context network and a guidance perception network.
[0024] Preferably, the step S3 specifically includes the following:
[0025] S3-1, the spatial context network encodes the remote region through the one-dimensional kernel, receives the depth information of different dimensions, and obtains the context-guided perception features;
[0026] S3-2. Obtain a guided perception map through a guided perception network.
[0027] Preferably, the step S3-1 includes the following sub-steps:
[0028] S3-1-1. The multi-scale feature x obtained by the multi-scale feature extraction network is extracted by sliding the strip window horizontally and vertically along the spatial dimension, where x∈R C×H×W , where R represents the real number domain, C is the number of spatial channels, H and W are the spatial height and width respectively, and the window sizes are (H, 1) and (1, W) respectively. The average value of all row features and the maximum value of all column features are obtained. Therefore, the output after horizontal merging is y h ∈R H×C , expressed as: The output after vertical merging is y w ∈R C×W , expressed as: Among them, i and j represent rows and columns respectively. Through the above operations, the strip area is encoded to capture the local detail information of the distant dense and refined image and collect the remote context information;
[0029] S3-1-2, y h and w Input into a one-dimensional convolutional layer with a kernel size of 3 to integrate the current position and its adjacent features to obtain semantic features y of different dimensions hc and wc .
[0030] S3-1-3, due to hc Position in and wc Position in They are independent of each other and are constructed by element-wise multiplication. and wc The relationship between and The correlation between features at different positions is represented by M, which is the output correlation feature map:
[0031]
[0032] Where mul[] represents element-wise multiplication;
[0033] S3-1-4, input the correlation feature map obtained in S3-1-3 into the cascade module, perform channel-level combination in the cascade module to increase the number of image features, integrate the information of its feature map, and use 1×1 convolution in space to process the merged features to obtain a similarity feature map, y h The different positions of the reshaped features are calculated with the similarity feature map to obtain the context-guided perception features. The calculated feature representation is shown in the following formula:
[0034]
[0035] Where Conv[] represents a convolution operation.
[0036] Preferably, the step S3-2 includes the following sub-steps:
[0037] S3-2-1. The multi-scale feature obtained by the multi-scale feature extraction network is represented as x∈R C×H×W , assuming x c Represented as the feature map corresponding to each channel, expressed as Since it is necessary to capture richer semantic information, global average pooling and global maximum pooling operations are first used to aggregate different spatial information, where l c1 is the global average pooling, l c2 It is the global maximum pooling, as shown below:
[0038] l c1 =Global AvgPool(x c ),l c2 =Global MaxPool(x c );
[0039] S3-2-2. Guide the perception network to adopt strategies to learn cross-channel interactions to preserve the precise correspondence between channel layers and weights. The specific cross-channel interaction method is as follows: Perform a 1D convolution operation with a kernel size of k, and then perform a superposition operation as shown in the formula:
[0040]
[0041] Among them, 1ConvD represents 1D convolution operation, Add represents the addition of corresponding features, and l c1 and l c2 The values of k in the branches are 3 and 5 respectively;
[0042] S3-2-3. At the channel level, channel standardization is used to achieve lateral suppression of channel types. Here, The channel normalization formula is as follows:
[0043]
[0044] S3-2-4, the calculation result of the above formula is used to obtain the guided perception map through the tanh activation function as shown in the following formula:
[0045]
[0046] Preferably, in the step S4, the size of the predicted density map is H×C×W, where the context-guided perceptual feature and the guided perceptual map are H×C×W, and the context-guided perceptual feature is denoted by f 1 , the guided perception map is denoted as f 2 , the attention fusion network will f 1 and f 2 Input to a 3×3 convolutional layer containing ReLU and BN, outputting f 1c and f 2c , and then f 1c and f 2c The combined output is then used as input to a 3×3 convolutional layer and a 1×1 convolutional layer. Global average pooling and a 1×1 convolutional layer are used for operation. The result after the operation is then predicted by softmax to fuse the attention weights α and β. α and β are then respectively compared with f 1c and f 2c Perform pixel-level dot multiplication and finally combine f 1 and f 2 Perform the sum operation and the result is output as f 3 , the calculation formula is as follows:
[0047] f 3 =Add(f 1 +f 2 +α·f 1c +β·f 2c )
[0048] Finally, the obtained f 3 Input to upsampling and convolution operations to obtain the final estimated prediction density map.
[0049] Preferably, in step S4, the real density map is obtained by marking the coordinates of the center positions of the heads in the crowd area in the real image and generating the real crowd density map through Gaussian smoothing operation. The real density map generation process can be specifically divided into the following steps:
[0050] Assume that any pixel x in a crowd image i There is a head mark at the position, which is represented by a unit impulse function δ(xxi ), then the marking of all head positions is as follows:
[0051]
[0052] Where x is a coordinate of any image, and N is the total number of heads marked in the image.
[0053] The generated H k (x) and the normalized two-dimensional Gaussian kernel G σ (x) Perform convolution operation to obtain the real crowd density map. The specific formula is as follows:
[0054] F k (x) = H k (x)*G σ (x)
[0055] Where σ represents the standard deviation, G σ (x) Normalized 2D Gaussian kernel, where the Gaussian kernel size is a fixed Gaussian kernel of 15×15.
[0056] Preferably, step S5 includes the following sub-steps:
[0057] S5-1. Calculate the Euclidean distance between the overall predicted density map and the true density map and use it as the Euclidean distance loss function. The formula is as follows:
[0058]
[0059] Where θ represents the parameters of the crowd counting network, M represents the number of training set samples, and X t represents the t-th input image, D(X t ; θ) represents the predicted density map, D t gt represents the true density map;
[0060] S5-2. Based on the above Euclidean distance loss function, an adaptive regional loss function is established;
[0061] S5-3. By weighting the Euclidean distance loss function and the adaptive region loss function, an adaptive scale loss function is obtained and trained.
[0062] Preferably, in step S5-2, the method for establishing the adaptive loss function is as follows:
[0063] S5-2-1. Using the center point of the predicted density map as the base center for cropping, crop a portion with a width of W from the predicted density map (H, W). i , high is H i The sub-area D(X t ;θ)i , and calculate the average density of unit pixels in the sub-region, the number of generated sub-regions is n, the scale factor is λ, W i and H i The representation is as follows:
[0064] W i =W×λ i-1 ,H i =H×λ i-1 ; i∈[1,n]
[0065] Then sub-area I i The average density per unit pixel is shown in the following formula:
[0066] AD i =∑D(X t ;θ) i / W i ×H i ;
[0067] S5-2-2. Sort the average density of the sub-regions in ascending order to obtain {AD 1 ,AD 2 ,...,AD i ,AD n}, select the AD with the largest average density d Then, the corresponding area is recorded as the difficult area D(X i ;θ) d , scale the area with the largest average density, and the scaling factor is as follows:
[0068] r=AD d / AD 1
[0069] Therefore, the scaled area is as follows:
[0070] D(X t ;θ) dr = r × D (X t ;θ) d ; D(X t ;θ) dr ∈D(X t ;θ);
[0071] S5-2-3. For the real density map, we do not directly crop and scale it. Instead, we first scale the binary head position map, and then calculate the corresponding real density map to reduce the gap between density areas at different levels. The real density map of the difficult area after scaling is The final adaptive region loss function is as follows:
[0072]
[0073] Here D(X t ;θ) dr Prediction density map showing the difficult area after scaling, True density map showing the difficult area after scaling.
[0074] Preferably, in step S5-3, by c and L dr The two losses are weighted, and the target loss function of the training is also the adaptive scale loss function L all , the formula is as follows:
[0075] L all =L c +μL dr
[0076] Where μ represents the weight, which is an adjustable hyperparameter, and the value of μ is set to 0.5.
[0077] The present invention has the following characteristics and beneficial effects:
[0078] 1. Use the Inception module to alleviate the scale variation problem in crowd counting, which has been verified in our work. Inspired by the Inception module, by using convolution kernels of different sizes, the width and scale adaptability of the network are improved, thereby improving the ability to capture multi-scale information. The MFEN network uses Inception-v3, a stack of multiple Inception modules, to improve the expressiveness of our network.
[0079] 2. Introduce the spatial context network (SCN) to establish regional dependencies between different spatial dimensions, capture remote context perception information of different spatial dimensions, and establish an information association model between different positions to enhance the receptive field. SCN encodes remote regions through a one-dimensional kernel and receives deep information of different dimensions. Due to the characteristics of its kernel, it only establishes dependencies on remote key information and traverses detailed element information of the entire scene. It uses the correlation between different positions and the dependency between channels to achieve a larger receptive field and improve the ability to analyze complex scenes. Considering the insufficient capture of spatial semantic information and the imperfect feature aggregation, the spatial context network avoids the establishment of unnecessary connections at distant locations, which is conducive to accurately aggregating relevant features of spatial dimensions, improving the situation where the target distribution in some areas is relatively messy, improving the ability to explore deep structured information, and expanding the receptive field.
[0080] 3. Guided Perception Network (GPN) avoids the side effects of prediction channels caused by the large number of parameters and dimensionality reduction in the fully connected layer, while modeling cross-channel dependencies, guiding the features of the intermediate layers, and reducing the existing semantic differences. In the process of guiding the perceptual weights and spatial context feature operations, more appropriate scale information and spatial positions are selected, which can alleviate background interference in complex scenes and emphasize objects in the area of interest.
[0081] 4. Attention Fusion Network (AFN) It not only allows each pixel to select appropriate context information in the aggregation stage, but also uses the guided perception network to guide multi-scale information and deep spatial context information. Compared with simple feature fusion, the AFN network proposed in this paper uses fused attention weights to adaptively allocate spatial context features and guided perception features. It can effectively guide multi-scale features, make up for the semantic information and spatial information between different levels, and improve resolution.
[0082] 5. The adaptive scale loss function determines the difficult area by learning the unit pixel density and re-constrains the dense area by scaling the area. Since dense crowds are usually more difficult to count than sparse crowds, we focus on areas with large errors and optimize the gaps between different levels of density to assist in the crowd counting task. BRIEF DESCRIPTION OF THE DRAWINGS
[0083] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0084] Figure 1 is a flow chart of a network (MGA Net) according to an embodiment of the present invention;
[0085] Figure 2 Inception V3 module in the embodiment of the present invention
[0086] Figure 3 It is a diagram of the architecture of a multi-scale feature extraction network (MFEN) in an embodiment of the present invention;
[0087] Figure 4 This is a diagram of the spatial context network (SCN) architecture in an embodiment of the present invention;
[0088] Figure 5 This is a diagram of the architecture of a guided perception network (GPN) in an embodiment of the present invention;
[0089] Figure 61 is an architecture diagram of an attention fusion network (AFN) in an embodiment of the present invention. DETAILED DESCRIPTION
[0090] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.
[0091] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate positions or positional relationships based on the positions or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second", and the like are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, features defined as "first", "second", and the like may explicitly or implicitly include one or more of the features. In the description of the present invention, unless otherwise specified, "multiple" means two or more.
[0092] In the description of the present invention, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood by specific circumstances.
[0093] The present invention provides a crowd counting method based on a multi-scale spatial guided perception aggregation network, such as Figure 1 As shown in the figure, it consists of three parts: multi-scale feature extraction network (MFEN), spatial guidance network (SGN) and attention fusion network (AFN), which implement the following steps:
[0094] S1. Establish a multi-scale feature extraction network
[0095] The multi-scale feature extraction network is based on the Inception-v3 model, removing the first two maximum pooling layers and the fully connected layer of Inception-v3, and retaining all convolutional layers in the Inception-v3 model.
[0096] The network layers included in the multi-scale feature extraction network are: five convolutional layers, three Inception-A, one Inception-D, four Inception-B, one Inception-E and two Inception-C;
[0097] Understandable, such as Figure 3 As shown in Figure 1, the multi-scale feature extraction network (MFEN) is based on the existing Inception-v3 network to further improve the network's expressiveness. At the end of the existing Inception-v3 model, the size of the feature map generated by convolution is 1 / 32 of the original input size. Its Inception module, such as Figure 2 As shown in the figure, the MFEN network makes the following modifications without changing the network parameters: the first two maximum pooling layers and the fully connected layer of Inception-v3 are deleted, and all convolutional layers in the Inception-v3 model are retained. The size of the MFEN network output x is 1 / 8 of the original input size.
[0098] In the above technical solution, by using convolution kernels of different sizes, the width and scale adaptability of the network are improved, thereby improving the ability to capture multi-scale information. The MFEN network uses Inception-v3, a stack of multiple Inception modules, to improve the expressiveness of the multi-scale feature extraction network.
[0099] S2. Input images of any resolution into the multi-scale feature extraction network
[0100] Given an image input I, the output features are represented by mapping:
[0101] x=F MFEN
[0102] Where x represents the multi-scale features captured by the image, and F MFEN represents a multi-scale feature extraction network;
[0103] S3. The multi-scale features captured by the multi-scale feature extraction network are input into the spatial guidance network, and the context-guided perception features and the guided perception map are output.
[0104] The spatial guidance network includes a spatial context network and a guidance perception network.
[0105] Specifically, the step S3 includes the following steps:
[0106] S3-1, the spatial context network encodes the remote region of the one-dimensional kernel, receives the depth information of different dimensions, and obtains the context-guided perception features, where the spatial context network, such as Figure 4 As shown,
[0107] Specifically, the step S3-1 includes the following sub-steps:
[0108] S3-1-1. The multi-scale feature x obtained by the multi-scale feature extraction network is extracted by sliding the strip window horizontally and vertically along the spatial dimension, where x∈R C×H×W , where R represents the real number domain, C is the number of spatial channels, H and W are the spatial height and width respectively, and the window sizes are (H, 1) and (1, W) respectively. The average value of all row features and the maximum value of all column features are obtained, so the output after horizontal merging is y h ∈R H×C , expressed as: The output after vertical merging is y w ∈R C×W , expressed as: Among them, i and j represent rows and columns respectively. Through the above operations, the strip area is encoded to capture the local detail information of the distant dense and refined image and collect the remote context information;
[0109] S3-1-2, y h and w Input into a one-dimensional convolutional layer with a kernel size of 3 to integrate the current position and its adjacent features to obtain semantic features y of different dimensions hc and wc .
[0110] S3-1-3, due to hc Position in and wc Position in They are independent of each other and are constructed by element-wise multiplication. and wc The relationship between and The correlation between features at different positions is represented by M, which is the output correlation feature map:
[0111]
[0112] Where mul[] represents element-wise multiplication;
[0113] It can be understood that in step S3-1-3, since only local information of the position itself is provided, no correlation is generated. To address this problem, an internal aggregation network is added to model the correlation between them.
[0114] S3-1-4, input the correlation feature map obtained in S3-1-3 into the cascade module, perform channel-level combination in the cascade module to increase the number of image features, integrate the information of its feature map, and use 1×1 convolution in space to process the merged features to obtain a similarity feature map, y h The different positions of the reshaped features are calculated with the similarity feature map to obtain the context-guided perception features, as shown in the following formula:
[0115]
[0116] Where Conv[] represents a convolution operation.
[0117] The above technical solution captures global remote context information and improves the ability to utilize inter-channel dependencies.
[0118] S3-2, obtain a guided perception graph through a guided perception network, where the guided perception network is as follows: Figure 5 shown.
[0119] Specifically, the step S3-2 includes the following sub-steps:
[0120] S3-2-1. The multi-scale feature obtained by the multi-scale feature extraction network is represented as x∈R C×H×W , assuming x c Represented as the feature map corresponding to each channel, expressed as Since it is necessary to capture richer semantic information, global average pooling and global maximum pooling operations are first used to aggregate different spatial information, where l c1 is the global average pooling, l c2 It is the global maximum pooling, as shown below:
[0121] l c1 =Global AvgPool(x c ),l c2 =Global MaxPool(x c );
[0122] S3-2-2. Guide the perception network to adopt strategies to learn cross-channel interactions to preserve the precise correspondence between channel layers and weights. The specific cross-channel interaction method is as follows: Perform a 1D convolution operation with a kernel size of k, and then perform a superposition operation as shown in the formula:
[0123]
[0124] Among them, 1ConvD represents 1D convolution operation, Add represents the addition of corresponding features, and l c1 and l c2The values of k in the branches are 3 and 5 respectively;
[0125] S3-2-3. At the channel level, channel standardization is used to achieve lateral suppression of channel types. Here, The channel normalization formula is as follows:
[0126]
[0127] Channel normalization is used in this step to make the relationship between channels more competitive and achieve lateral suppression of channel types.
[0128] S3-2-4, the calculation result of the above formula is used to obtain the guided perception map through the tanh activation function as shown in the following formula:
[0129]
[0130] S4, the context-guided perception features and the guided perception map are passed to the attention fusion network, and finally the density map is output and the density map training set is constructed, where the attention fusion network is as follows Figure 6 As shown;
[0131] Specifically, in step S4, the acquisition of the prediction density map, the size of the context-guided perceptual feature and the guided perceptual map is H×C×W, where the context-guided perceptual feature is denoted by f 1 , the guided perception map is denoted as f 2 , the attention fusion network will f 1 and f 2 Input to a 3×3 convolutional layer containing ReLU and BN, outputting f 1c and f 2c , and then f 1c and f 2c The combined output is then used as input to a 3×3 convolutional layer and a 1×1 convolutional layer. Global average pooling and a 1×1 convolutional layer are used for operation. The result after the operation is then predicted by softmax to fuse the attention weights α and β. α and β are then respectively compared with f 1c and f 2c Perform pixel-level dot multiplication and finally combine f 1 and f 2 Perform the sum operation and the result is output as f 3 , the calculation formula is as follows:
[0132] f 3 =Add(f 1 +f 2 +α·f 1c +β·f 2c )
[0133] Finally, the obtained f 3 Input to upsampling and convolution operations to obtain the final estimated prediction density map.
[0134] Furthermore, the real density map is generated by marking the coordinates of the center positions of the heads in the crowd area in the real image and performing Gaussian smoothing operation to generate the real crowd density map. The real density map generation process can be specifically divided into the following steps:
[0135] Assume that any pixel x in a crowd image i There is a head mark at the position, which is represented by a unit impulse function δ(xx i ), then the marking of all head positions is as follows:
[0136]
[0137] Where x is a coordinate of any image, and N is the total number of heads marked in the image.
[0138] The generated H k (x) and the normalized two-dimensional Gaussian kernel G σ (x) Perform convolution operation to obtain the real crowd density map. The specific formula is as follows:
[0139] F k (x) = H k (x)*G σ (x)
[0140] Where σ represents the standard deviation, G σ (x) Normalized 2D Gaussian kernel, where the Gaussian kernel size is a fixed Gaussian kernel of 15×15.
[0141] S5. Establish an adaptive scale loss function and perform adaptive training through the density map training set;
[0142] Specifically, step S5 includes the following sub-steps:
[0143] S5-1. Calculate the Euclidean distance between the overall predicted density map and the true density map and use it as the Euclidean distance loss function L c , the formula is as follows:
[0144]
[0145] Where θ represents the parameters of the crowd counting network, M represents the number of training set samples, and X t represents the t-th input image, D(X t ; θ) represents the predicted density map, represents the true density map;
[0146] S5-2. Based on the above Euclidean distance loss function, an adaptive regional loss function is established;
[0147] In step S5-2, the method for establishing the adaptive loss function is as follows:
[0148] S5-2-1. Using the center point of the predicted density map as the base center for cropping, crop a portion with a width of W from the predicted density map (H, W). i , high is H i The sub-area D(X t ;θ) i , and calculate the average density of unit pixels in the sub-region, the number of generated sub-regions is n, the scale factor is λ, W i and H i The representation is as follows:
[0149] W i =W×λ i-1 ,H i =H×λ i-1 ; i∈[1,n]
[0150] Then sub-area I i The average density per unit pixel is shown in the following formula:
[0151] AD i =∑D(X t ;θ) i / W i ×H i ;
[0152] S5-2-2. Sort the average density of the sub-regions in ascending order to obtain {AD 1 ,AD 2 ,...,AD i ,AD n}, select the AD with the largest average density d Then, the corresponding area is recorded as the difficult area D(X i ;θ) d , scale the area with the largest average density, and the scaling factor is as follows:
[0153] r=AD d / AD 1
[0154] Therefore, the scaled area is as follows:
[0155] D(X t ;θ) dr = r × D (X t ;θ) d ; D(X t ;θ) dr∈D(X t ;θ);
[0156] S5-2-3. For the real density map, we do not directly crop and scale it. Instead, we first scale the binary head position map, and then calculate the corresponding real density map to reduce the gap between density areas at different levels. The real density map of the difficult area after scaling is The final adaptive regional loss function L dr , as shown in the formula below:
[0157]
[0158] Here D(X t ;θ) dr Prediction density map showing the difficult area after scaling, True density map showing the difficult area after scaling.
[0159] S5-3. By weighting the Euclidean distance loss function and the adaptive region loss function, an adaptive scale loss function is obtained and trained.
[0160] In step S5-3, by c and L dr The two losses are weighted, and the target loss function of the training is also the adaptive scale loss function L all , the formula is as follows:
[0161] L all =L c +μL dr
[0162] Where μ represents the weight, which is an adjustable hyperparameter, and the value of μ is set to 0.5.
[0163] It is understandable that the adaptive scale loss function determines the difficult area by learning the unit pixel density size and re-constrains the dense area by scaling the area. Since dense crowds are usually more difficult to count than sparse crowds, we focus on areas with large errors and optimize the gaps between different levels of density to assist in the crowd counting task.
[0164] S6. Take the image to be predicted as input, repeat steps S2-S5, and output the calculation result of the crowd in the image to be predicted.
[0165] The embodiments of the present invention are described in detail above with reference to the accompanying drawings, but the present invention is not limited to the described embodiments. For those skilled in the art, various changes, modifications, substitutions and variations of these embodiments including components are made without departing from the principles and spirit of the present invention, and still fall within the scope of protection of the present invention.
Claims
1. A crowd counting method based on multi-scale spatial guided perception aggregation network, It is characterized in that The steps include: S1. Establish a multi-scale feature extraction network The multi-scale feature extraction network is based on the Inception-v3 model, removing the first two maximum pooling layers and the fully connected layer of Inception-v3, and retaining all convolutional layers in the Inception-v3 model. The network layers included in the multi-scale feature extraction network are: five convolutional layers, three Inception-A, one Inception-D, four Inception-B, one Inception-E and two Inception-C; S2. Input images of any resolution into the multi-scale feature extraction network Given an image input I, the output features are represented by mapping: x=F MFEN Where x represents the multi-scale features captured by the image, and F MFEN represents a multi-scale feature extraction network; S3, inputting the multi-scale features captured by the multi-scale feature extraction network into the spatial guidance network, and outputting the context-guided perception features and the guided perception map; S4, passing the context-guided perception features and the guided perception map to the attention fusion network, finally outputting a density map, and constructing a density map training set, wherein the density map training set includes a predicted density map and a true density map; S5. Establish an adaptive scale loss function and perform adaptive training through the density map training set; S6. Take the image to be predicted as input, repeat steps S2-S5, and output the calculation result of the crowd in the image to be predicted.
2. According to claim 1, the crowd counting method based on multi-scale spatial guided perception aggregation network, It is characterized in that The spatial guidance network includes a spatial context network and a guidance perception network.
3. According to claim 2, the crowd counting method based on multi-scale spatial guided perception aggregation network, It is characterized in that The step S3 specifically includes the following: S3-1, the spatial context network encodes the remote region through the one-dimensional kernel, receives the depth information of different dimensions, and obtains the context-guided perception features; S3-2. Obtain a guided perception map through a guided perception network.
4. The crowd counting method based on multi-scale spatial guided perception aggregation network according to claim 3, It is characterized in that The step S3-1 includes the following sub-steps: S3-1-1. The multi-scale feature x obtained by the multi-scale feature extraction network is extracted by sliding the strip window horizontally and vertically along the spatial dimension, where x∈R C×H×W , where R represents the real number domain, C is the number of spatial channels, H and W are the spatial height and width respectively, and the window sizes are (H, 1) and (1, W) respectively. The average value of all row features and the maximum value of all column features are obtained. Therefore, the output after horizontal merging is y h ∈R H×C , expressed as: The output after vertical merging is y w ∈R C×W , expressed as: Among them, i and j represent rows and columns respectively. Through the above operations, the strip area is encoded to capture the local detail information of the distant dense and refined image and collect the remote context information; S3-1-2, y h and w Input into a one-dimensional convolutional layer with a kernel size of 3 to integrate the current position and its adjacent features to obtain semantic features y of different dimensions hc and wc ; S3-1-3, due to hc Position in and wc Position in They are independent of each other and are constructed by element-wise multiplication. and wc The relationship between and The correlation between features at different positions is represented by M, which is the output correlation feature map: Where mul[] represents element-wise multiplication; S3-1-4, input the correlation feature map obtained in S3-1-3 into the cascade module, perform channel-level combination in the cascade module to increase the number of image features, integrate the information of its feature map, and use 1×1 convolution in space to process the merged features to obtain a similarity feature map, y h The different positions of the reshaped features are calculated with the similarity feature map to obtain the context-guided perception features, as shown in the following formula: Where Conv[] represents the convolution operation.
5. According to claim 4, the crowd counting method based on multi-scale spatial guided perception aggregation network, It is characterized in that The step S3-2 includes the following sub-steps: S3-2-1. The multi-scale feature obtained by the multi-scale feature extraction network is represented as x∈R C×H×W , assuming x c Represented as the feature map corresponding to each channel, expressed as Since it is necessary to capture richer semantic information, global average pooling and global maximum pooling operations are first used to aggregate different spatial information, where l c1 is the global average pooling, l c2 It is the global maximum pooling, as shown below: l c1 =GlobalAvgPool(x c ),l c2 =Global MaxPool(x c ); S3-2-2. Guide the perception network to adopt strategies to learn cross-channel interactions to preserve the precise correspondence between channel layers and weights. The specific cross-channel interaction method is as follows: Perform a 1D convolution operation with a kernel size of k, and then perform a superposition operation as shown in the formula: Conv1D represents a 1D convolution operation, Add represents the addition of corresponding features, and l c1 and l c2 The values of k in the branches are 3 and 5 respectively; S3-2-3. At the channel level, channel standardization is used to achieve lateral suppression of channel types. Here, The channel normalization formula is as follows: S3-2-4, the calculation result of the above formula is used to obtain the guided perception map through the tanh activation function as shown in the following formula:
6. The crowd counting method based on multi-scale spatial guided perception aggregation network according to claim 5, It is characterized in that In step S4, the acquisition of the prediction density map, the size of the context-guided perceptual feature and the guided perceptual map is H×C×W, where the context-guided perceptual feature is denoted by f 1 , the guided perception map is denoted as f 2 , the attention fusion network will f 1 and f 2 Input to a 3×3 convolutional layer containing ReLU and BN, outputting f 1c and f 2c , and then f 1c and f 2c The combined output is then used as input to a 3×3 convolutional layer and a 1×1 convolutional layer. Global average pooling and a 1×1 convolutional layer are used for operation. The result after the operation is then predicted by softmax to fuse the attention weights α and β. α and β are then respectively compared with f 1c and f 2c Perform pixel-level dot multiplication and finally combine f 1 and f 2 Perform the sum operation and the result is output as f 3 , the calculation formula is as follows: f 3 =Add(f 1 +f 2 +α·f 1c +β·f 2c ) Finally, the obtained f 3 Input to upsampling and convolution operations to obtain the final estimated prediction density map.
7. The crowd counting method based on multi-scale spatial guided perception aggregation network according to claim 6, It is characterized in that In step S4, the real density map is obtained. The real density map is generated by marking the coordinates of the center positions of the heads in the crowd area in the real image and performing Gaussian smoothing operation. The real density map generation process can be specifically divided into the following steps: Assume that any pixel x in a crowd image i There is a head mark at the position, which is represented by a unit impulse function δ(xx i ), then the marking of all head positions is as follows: Where x is a coordinate of any image, and N is the total number of heads marked in the image; The generated H k (x) and the normalized two-dimensional Gaussian kernel G σ (x) Perform convolution operation to obtain the real density map. The specific formula is as follows: F k (x)=H k (x)*G σ (x) Where σ represents the standard deviation, G σ (x) Normalized 2D Gaussian kernel, where the Gaussian kernel size is a fixed Gaussian kernel of 15×15.
8. The crowd counting method based on multi-scale spatial guided perception aggregation network according to claim 7, It is characterized in that The step S5 includes the following sub-steps: S5-1. Calculate the Euclidean distance between the overall predicted density map and the true density map and use it as the Euclidean distance loss function. The formula is as follows: Where θ represents the parameters of the crowd counting network, M represents the number of training set samples, and X t represents the t-th input image, D(X t ; θ) represents the predicted density map, represents the true density map; S5-2. Based on the above Euclidean distance loss function, an adaptive regional loss function is established; S5-3. By weighting the Euclidean distance loss function and the adaptive region loss function, an adaptive scale loss function is obtained and trained.
9. The crowd counting method based on multi-scale spatial guided perception aggregation network according to claim 8, It is characterized in that In step S5-2, the method for establishing the adaptive loss function is as follows: S5-2-1. Using the center point of the predicted density map as the base center for cropping, crop a portion with a width of W from the predicted density map (H, W). i , high is H i The sub-area D(X t ;θ) i , and calculate the average density of unit pixels in the sub-region, the number of generated sub-regions is n, the scale factor is λ, W i and H i The representation is as follows: W i =W×λ i-1 ,H i =H×λ i-1 ;i∈[1,n] Then sub-area I i The average density per unit pixel is shown in the following formula: AD i =∑D(X t ;θ) i / W i ×H i ; S5-2-2. Sort the average density of the sub-regions in ascending order to obtain {AD 1 ,AD 2 ,...,AD i ,AD n }, select the AD with the largest average density d Then, the corresponding area is recorded as the difficult area D(X i ;θ) d , scale the area with the largest average density, and the scaling factor is as follows: r=AD d / AD 1 Therefore, the scaled area is as follows: D(X t (i) dr =r×D(X t (i) d ;D(X t (i) dr ∈D(X t ;i); S5-2-3. For the real density map, we do not directly crop and scale it. Instead, we first scale the binary head position map, and then calculate the corresponding real density map to reduce the gap between density areas at different levels. The real density map of the difficult area after scaling is The final adaptive region loss function is as follows: Here D(X t ;θ) dr Prediction density map showing the difficult area after scaling, True density map showing the difficult area after scaling.
10. The crowd counting method based on multi-scale spatial guided perception aggregation network according to claim 8, It is characterized in that In step S5-3, by c and L dr The two losses are weighted, and the target loss function of the training is also the adaptive scale loss function L all , the formula is as follows: L all =L c +μL dr Where μ represents the weight, which is an adjustable hyperparameter, and the value of μ is set to 0.5.
Citation Information
Patent Citations
Crowd counting method based on coding-decoding structure multi-scale convolutional neural network
CN111242036A
Multi-level attention scale perception crowd counting method
CN113283356A