A video crowd counting method based on the combination of attention and spatial transformation network

By combining attention and spatial transformation network in video crowd counting, using continuous frame flow inference and personnel conservation constraints, the problem of multi-frame information combination and spatial invariance in video crowd counting is solved, and the accuracy and robustness of counting are improved.

CN116385964BActive Publication Date: 2025-07-08HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310296472.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-24
Publication Date
2025-07-08
Estimated Expiration
2043-03-24

AI Technical Summary

Technical Problem

The prior art is difficult to effectively combine multi-frame information in video crowd counting to deal with spatial invariance problems, and there is a risk of gradient disappearance or explosion and perspective distortion.

Method used

Using an attention-based and spatial transformation network method, through the inference of people flow between continuous frames, combined with personnel conservation constraints and deep learning neural networks, the spatial transformation network and residual network are used to process posture changes and rotation problems in video frames, and solve perspective distortion through the channel-space attention mechanism.

Benefits of technology

It improves the accuracy and robustness of video crowd counting, comprehensively considers time dimension information, reduces the complexity of network training and the risk of gradient disappearance and explosion, and improves counting accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116385964B_ABST
    Figure CN116385964B_ABST
Patent Text Reader

Abstract

The present invention discloses a video crowd counting method based on the combination of attention and spatial transformation network. First, two consecutive images collected by a camera are obtained and input into the front-end encoder model of the network, and density maps of the current frame and the previous frame of the current frame are output. Secondly, the output density maps of the current frame and the previous frame of the current frame are subjected to feature fusion and transmitted to the back-end decoder model of the network to output a crowd flow map. Then, a loss function is established, and the network is trained with the obtained real density map training set. Finally, the video frame to be processed is used as input, and the network outputs the calculation result of the number of people in the video frame to be predicted. The present invention makes the perspective of solving the crowd counting problem more comprehensive, effectively alleviates the problem that deep networks are difficult to train and prone to gradient disappearance and explosion, and respectively learns the importance of channels and the importance of space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and specifically refers to a video crowd counting method based on the combination of attention and spatial transformation network. Background Art

[0002] Due to the amazing progress of machine learning, especially after the wide application of deep learning in the field of computer vision, the field of computer vision has also achieved unprecedented development in recent years. For example, great success has been achieved in fields such as object detection and face recognition. With the rise of smart cities in recent years, the proportion of urban surveillance cameras has become higher and higher. Coupled with the frequent occurrence of stampede events in recent years, people have to pay attention to the field of crowd counting. In sparse scenes where the image contains a single or multiple targets, crowd counting and recognition can be easily and accurately performed through target localization and detection technology. Due to the gradual development of deep learning, methods based on convolutional neural networks (CNNs) have achieved remarkable success in tasks such as image classification, pedestrian detection, and speech recognition. Therefore, researchers introduced CNNs into the field of crowd counting and obtained good crowd density estimation results in learning the mapping between images and density maps. In recent years, researchers have designed a variety of CNN-based crowd counting algorithms to overcome challenges such as scale variation, non-uniform distribution, occlusion, and complex backgrounds. These algorithms mainly include single-branch network models, multi-branch network models, attention mechanisms, and feature fusion methods.

[0003] The existing technologies mainly have the following three problems in crowd recognition and counting:

[0004] 1. In video crowd counting, the problem of how to combine the information of multiple frames in the video needs to be solved first. When adding multiple frames, the information in the time dimension can be considered. So the problem is how to consider the time information dimension.

[0005] 2. In the field of crowd counting, the problem of the most common spatial invariance needs to be solved, such as various pose changes of people in video frames, rotation of people caused by the movement of the camera, etc. In addition, as the network depth deepens, the network shows the problem of degradation. The gradient of backpropagation is prone to dispersion, and it may also lead to gradient disappearance or gradient explosion, which is also a common problem.

[0006] 3. Due to the problem of the shooting angle of the camera, the problem of perspective distortion is first shown in the captured picture, which has a greater impact on the accuracy of crowd counting. Summary of the Invention

[0007] In view of the deficiencies of the prior art, the present invention infers the density of people from the flow of people between consecutive frames, and proposes a video crowd counting method based on the combination of attention and spatial transformation network to regress the number of people, improving the accuracy and robustness of counting.

[0008] To solve the above technical problems, the technical solution of the present invention is as follows:

[0009] This is a video-based counting scheme that does not directly estimate crowd density from an image, but infers it from the so-called flow of people between two consecutive frames.

[0010] Specifically, two consecutive given images are used as inputs, and the output flow of people is f t-1,t . Where f t-1,t represents the amount of people's movement between two consecutive frames I t-1 and I t , which is also called the flow of people here. The so-called flow of people is a vector field that associates the pedestrian motion vectors with each point in the frame space.

[0011] Once the amount of movement f t-1 between two consecutive frames I t is predicted, the density map of this frame t-1 , f t at the j-th spatial position can be reconstructed by summing up all the contributions of the flow of people entering j from the neighboring positions of the previous frame, expressed as:

[0012]

[0013] where the neighboring positions of the j-th position are denoted as N(j), and the number of people moving from location i to location j between time t-1 and t is denoted as Summing up all the pixel values of the obtained density map gives the final number of people at time t:

[0014] Specifically, the flow of people is constructed by imposing a people conservation constraint, which means that people cannot appear or disappear between consecutive frames unless they are at the edge of a frame. Here, only the ground truth density maps and at consecutive time steps (t-1, t) are used to estimate the flow of people. Specifically, the constraint conditions can be expressed as the following two:

[0015]

[0016]

[0017] Among them, the first constraint conserves the people near position j within consecutive frame intervals, and the second constraint strengthens the spatio-temporal symmetry of the flow, that is, when time goes backward, people should move in the opposite direction.

[0018] Select a suitable regression flow function in the above learning framework. Here, a deep learning neural network is selected as this function, where I t-1 , I t refers to two consecutive frames I t-1 and I t . Given two consecutive input images, it outputs the flow as f t-1,t . Among them, the parameter θ is optimized by strengthening the constraints of the equation in S1-2 during the training process. Here, the method combining attention and spatial transformation network is used as where is divided into a front-end encoder model and a back-end decoder model.

[0019] A video crowd counting method based on the combination of attention and spatial transformation network includes the following steps:

[0020] S1. Obtain the current frame and the previous frame of the current frame collected by the camera to obtain two consecutive images.

[0021] S2. Input the current frame and the previous frame of the current frame into the front-end encoder model of the network, and output the density map of the current frame and the density map of the previous frame of the current frame. The encoder model consists of a spatial transformation network (STN), a residual network, an attention mechanism, etc., specifically as follows:

[0022] Preferably, the above steps are specifically divided into S2-1, S2-2 and S2-3:

[0023] S2-1. Input the current frame into the spatial transformation network (STN). The entire spatial transformation network includes three parts: a regression network, a grid generator, and a sampler.

[0024] The regression network is a network used to regress the transformation parameter θ. Its input U is a feature image, and then through a series of hidden network layers (fully connected or convolutional network, plus a regression layer), it outputs the transformation parameter. The form of θ can be diverse. If a 2D affine transformation needs to be implemented, it is the output of a 6D (2x3) vector, and the formula is as follows:

[0025] θ = f loc (U)

[0026] where θ is the transformation parameter, the input U is the feature image, and f locis a regression network.

[0027] The grid generator constructs a sampling grid based on the predicted transformation parameters, which is the output obtained by sampling and transforming points in a set of input images. In fact, the grid generator obtains a mapping relationship Γ θ , assuming that the coordinates of each pixel in the feature image U are the coordinates of each pixel in the output V are The spatial transformation function Γ θ is a two-dimensional affine transformation function, and A θ represents the matrix composed of the transformation parameters θ. Then and The corresponding relationship can be written as:

[0028]

[0029] The sampler uses the output of the grid generator and the input feature map as inputs simultaneously to produce an output, obtaining the result of the feature map after transformation. The formula is as follows:

[0030]

[0031] where is the gray value of a certain point in the c-th channel on the output feature map, is the gray value of the point (n, m) in the c-th channel on the input feature map. When or is greater than 1, the corresponding max() term will take 0. That is to say, only the gray values of the 4 surrounding points of (x i , y i ) determine the gray value of the target pixel point. And when and are smaller, the influence is greater (i.e., closer to the point (n, m)), and the weight is greater.

[0032] S2-2. Take the output of the spatial transformation network as the input and pass it to a forward feature extraction network F forward . This network model is used to extract the features of the current frame, and the output features are represented by the mapping as:

[0033] x = F forward

[0034] where x represents the features captured in the current frame image, and F forward represents the forward feature extraction network, which is composed of 9 convolutional layers and 5 max pooling layers.

[0035] S2-3. Take the features extracted by the previous feature extraction network as input and pass them into three consecutive channel attention residual networks (CBARM). CBARM consists of a 1D channel attention mechanism M c ∈R C×1×1 and a 2D spatial attention mechanism M s ∈R 1 ×H×W as well as a residual network. Denote the feature map input extracted by the previous feature extraction network as F∈R C×H×W . The entire attention mechanism process can be summarized as the following formula:

[0036]

[0037]

[0038] where denotes element-wise multiplication. When performing element-wise multiplication, the attention values are broadcast accordingly. The channel attention value is broadcast along the spatial dimension, and vice versa. F′ is the output along the channel attention network, and F″ is the final refined output.

[0039] Preferably, S2-3 includes the following sub-steps:

[0040] S2-3-1. First, aggregate the spatial information of a feature map through average pooling and max pooling operations to generate two different spatial context descriptors: and These two descriptors represent the features averaged pooled and max pooled respectively. Then, these two descriptors are fed forward into a shared network to generate the channel attention map M c ∈R C×1×1 . This shared network consists of a multi-layer perceptron (MLP) with one hidden layer. To reduce the parameter overhead, the activation size of the hidden layer is set to R C / r×1×1 , where r is the reduction rate. After each spatial context descriptor is processed by the shared network, the output feature vectors are fused using element-wise addition. The channel attention calculation method is as shown in the following formula:

[0041]

[0042] where σ is the sigmoid function, the weight sizes of the MLP are W0∈R C / r×C and W1∈R C / C×r , shared by the two inputs of the average pooled feature and the max pooled feature, and the Relu activation function is followed by W0, that is, the pooled feature needs to be processed by the Relu function before being input into the MLP.

[0043] S2-3-2. Use the channel attention map as the input and feed it into the spatial attention mechanism. The specific process is as described above. The spatial attention mechanism can generate a spatial attention map by leveraging the spatial correlation of features.

[0044] To calculate the spatial attention, first perform average pooling and max pooling operations along the channel dimension, and then concatenate them to generate an efficient feature descriptor. Performing pooling operations along the channel dimension has been proven to effectively highlight information-rich regions. For the concatenated feature descriptor, use a convolutional layer to generate a spatial attention map M s (F) ∈ R H×W , which encodes which regions are highlighted or suppressed.

[0045] As stated above, two pooling operations generate two two-dimensional 2D maps: and Each map represents the features averaged along the channels and the features max-pooled along the channels, respectively. Then, the two are concatenated and convolved by a standard convolutional layer to generate a 2D spatial attention map. The spatial attention calculation method is as follows:

[0046]

[0047] where σ is the sigmoid function, and f 7×7 represents a convolution operation with a 7×7 convolution kernel;

[0048] S2-3-3. Integrate the residual block ResBlock in the residual neural network ResNet into the channel-spatial attention mechanism CBAM to form the attention residual network (CBARM). One end of ResBlock is inserted at the front end of the channel attention mechanism, and the other end is inserted at the back end of the spatial attention mechanism. The formula is as follows:

[0049] y l = h(x l ) + F(x l , W l )

[0050] x l+1 = σ(y l )

[0051] where x l and x l+1 represent the input and output of the residual unit, that is, the input and output of the CBARM attention residual network. F is the learned residual, and W l is the weight. h(x l ) = x lrepresents the identity mapping, and σ is the Relu activation function.

[0052] Then the features learned from the shallow layer l to the deep layer L are:

[0053]

[0054] Among them, the deep layer L can be expressed as the sum of any shallower l layer and the residual part between them. W i is the weight, and L is the unit cumulative sum of the features of each residual block.

[0055] S3. Perform feature fusion in the concat form on the density map of the current frame output by the front-end encoder model in S2 and the density map of the previous frame of the current frame. Then the single output channel after concat is as follows, and the specific formula is as follows:

[0056]

[0057] where Z concat represents the single output channel after concat, X i corresponds to the channel of the density map of the output current frame, Y i corresponds to the channel of the density map of the previous frame of the current frame, * represents convolution, K represents the corresponding convolution kernel, i is the channel of the feature map, and c is the number of channels of the feature map.

[0058] S4. Use Z concat after feature fusion in S3 as the input and pass it to the back-end decoder model of the network. Finally, the human flow map is output. The decoder model consists of four dilated convolutions.

[0059] S5. Establish a loss function and train the network through the obtained real density map training set.

[0060] Preferably, step S5 is divided into the following steps:

[0061] S5-1. Obtaining the real density map. The real density map is generated by marking the coordinates of the center positions of human heads in the crowd area of the real image and performing Gaussian smoothing operation. The process of generating the real density map is as follows;

[0062] In each input t-th frame image I t is annotated with a set of s t two-dimensional points that represent the positions of human heads in the scene. Set the positions containing these points to 1 in the density map and set the positions not containing these points to 0 in the density map. Then convolve this density map with a Gaussian kernel with mean μ and standard deviation σ Perform convolution to obtain the corresponding ground truth density map The formula is as follows:

[0063]

[0064] where P j represents the center at position j.

[0065] S5-2. To strengthen the constraints of the two equations in S1-2, the combined loss function is used as the weighted term of the two loss functions, and the formula is expressed as:

[0066]

[0067]

[0068]

[0069] where is the ground truth density value, that is, the number of people at time t and position j. α is set to 1. During training, three consecutive frames are used to evaluate L combi .

[0070] S6. Take the video frame to be processed as the input, repeat steps S2 - S5, and output the calculation result of the number of people in the video frame to be predicted.

[0071] The present invention has the following beneficial effects:

[0072] 1. Different from the research on crowd counting based on images, in crowd counting in the video field, first, the problem of how to combine the information of multiple frames in the video needs to be solved. When adding multiple frames, the information in the time dimension can also be taken into account, making the perspective of solving the crowd counting problem more comprehensive. The method here is to consider the current frame and the frames before and after, and then separately consider the features of these frames through a forward encoder model, and then use concat to fuse the information of these multiple frames, so that the information in the time dimension can also be considered.

[0073] 2. Use the spatial transformation network to specifically design an explicit processing module for the network to handle the most common problem of spatial invariance in the field of crowd counting, such as various pose changes of people in the video frame, rotation of people caused by the movement of the camera, etc. Secondly, the network draws on CSRNet and uses it as the main network architecture. Specifically, the spatial transformation network is placed at the very front of the network architecture or after the image preprocessing stage. Integrating the spatial transformation network module into the convolutional neural network allows the network to automatically learn how to perform the transformation of the feature map, thus helping to reduce the overall cost in network training. The value output in the localization network indicates how to transform each training data.

[0074] 3. Use a residual network. Starting from features, better results are achieved through the effective utilization of features. The core idea is cross-layer connection. In the network, this residual network architecture is used, such that in a specified layer, each layer takes the previous layer as input, and at the same time, this layer is also passed as input to the subsequent layer to ensure the maximum utilization of features between layers. The residual network architecture is placed before and after each dilated convolution at the front end of the network. Such an architecture can well reuse the information of shallower layers and effectively alleviate the problems that deep neural networks are difficult to train and prone to vanishing and exploding gradients.

[0075] 4. Applying the channel-spatial attention mechanism can solve the perspective distortion problem in the field of crowd counting. In the channel attention mechanism, the model compresses features in the spatial dimension, turning each two-dimensional feature channel into a real number. This real number has a global receptive field to some extent. After obtaining the weights of each feature channel through this architecture, the weights are then applied to each original feature channel. Based on a specific task, the importance of different channels can be learned, and thus the information between different channels in the video frame can be utilized. And in the spatial attention mechanism, the channels themselves are dimensionally reduced. The maximum pooling and average pooling results are respectively obtained and then concatenated into a feature map, and then a convolutional layer is used for learning. These two mechanisms take into account both channels and space, and learn the importance of channels and space respectively. Description of the Drawings

[0076] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0077] Figure 1 It is the overall network flowchart of the embodiment of the present invention;

[0078] Figure 2 It is the encoder and decoder modules in the embodiment of the present invention;

[0079] Figure 3 It is the architecture diagram of the front feature extraction network in the embodiment of the present invention;

[0080] Figure 4 It is the architecture diagram of the spatial transformation network (STN) in the embodiment of the present invention;

[0081] Figure 5 It is the architecture diagram of the attention residual network (CBARM) in the embodiment of the present invention. Detailed Embodiments

[0082] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.

[0083] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the present invention. In addition, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, the meaning of "a plurality" is two or more.

[0084] In the description of the present invention, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", "connected" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood through specific situations.

[0085] The present invention relates to the field of computer vision technology, and specifically refers to a background technology of a video crowd counting method based on the combination of attention and spatial transformation network, as Figure 1 shown, which is specifically composed of an encoder and a decoder, and realizes the following steps:

[0086] Regression is used to estimate the crowd density of the pedestrian flow. This is a counting scheme based on video. It does not directly estimate the crowd density from the image, but infers it from the so-called pedestrian flow between two consecutive frames.

[0087] Specifically, two given consecutive images are used as inputs, and the output pedestrian flow is f t-1,t . Where f t-1,t represents the amount of people's movement between two consecutive frames I t-1 and I t , which is also called the pedestrian flow here. The so-called pedestrian flow is a vector field that associates the pedestrian motion vector with each point in the frame space.

[0088] Once people flow the momentum f t-1 between two consecutive frames I t and I t-1 , f t is predicted, the j-th spatial position of the density map of this frame can be reconstructed by summing up all the contributions of the incoming flows from neighboring positions of the previous frame into j, expressed as:

[0089]

[0090] where the adjacent positions of the j-th position are denoted as N(j), and the number of people moving from location i to location j between time t - 1 and t is denoted as Summing up all the pixel values of the obtained density map gives the final number of people at time t:

[0091] In particular, the human flow is constructed by imposing a conservation constraint on people, which means that people cannot appear or disappear between consecutive frames if not at the edge of a frame. Here, only the ground truth density maps and at consecutive time steps (t - 1, t) are used to estimate the human flow. Specifically, the constraint conditions can be expressed as the following two:

[0092]

[0093]

[0094] Among them, the first of the above constraint conditions conserves the people near position j within consecutive frame intervals, and the second constraint condition strengthens the spatio-temporal symmetry of the flow, that is, when time goes backward, people should move in the opposite direction.

[0095] In the above learning framework, a suitable function for the regression flow is selected. Here, a deep learning neural network is chosen as this function, where I t-1 , I t refers to two consecutive frames I t-1 and I t . Given two consecutive input images, it outputs the flow as f t-1,t . Among them, the parameter θ is optimized by strengthening the constraints of the equations in S1-2 during the training process. Here, the method of combining attention and spatial transformation network is used as where is divided into a front-end encoder model and a back-end decoder model, as Figure 2 shown.

[0096] S1. Obtain the current frame and the previous frame of the current frame collected by the camera to get two consecutive images.

[0097] S2. Input the current frame and the previous frame of the current frame into the front - end encoder model of the network respectively, as shown in Figure 3 . Output the density map of the current frame and the density map of the previous frame of the current frame. The encoder model consists of a Spatial Transformer Network (STN), a Residual Network, an attention mechanism, etc., as shown in Figure 4 specifically as follows:

[0098] Preferably, the above steps are specifically divided into S2 - 1, S2 - 2 and S2 - 3:

[0099] S2 - 1. Input the current frame into the Spatial Transformer Network (STN). The entire Spatial Transformer Network consists of three parts: a regression network, a grid generator, and a sampler.

[0100] The regression network is a network used to regress the transformation parameter θ. Its input U is a feature image, and then through a series of hidden network layers (fully connected or convolutional network, plus a regression layer), it outputs the spatial transformation parameter. The form of θ can be diverse. For example, to implement a 2D affine transformation, it is the output of a 6 - dimensional (2x3) vector, and the formula is as follows:

[0101] θ = f loc (U)

[0102] where θ is the transformation parameter, the input U is the feature image, and f loc is the regression network.

[0103] The grid generator constructs a sampling grid based on the predicted transformation parameter. It is the output obtained by sampling and transforming a set of points in the input image. The grid generator actually obtains a mapping relationship Γ θ . Assume that the coordinates of each pixel of the feature image U are and the coordinates of each pixel of the output V are The spatial transformation function Γ θ is a two - dimensional affine transformation function, and A θ represents the matrix composed of the transformation parameter θ. Then and The corresponding relationship can be written as:

[0104]

[0105] The sampler uses the output of the grid generator and the input feature map as inputs simultaneously to generate an output, obtaining the feature map result after the transformation of the feature map. The formula is as follows:

[0106]

[0107] Among them is the gray value of a certain point in the c-th channel of the output feature map is the gray value of the point (n, m) in the c-th channel of the input feature map. When or is greater than 1, the corresponding max() term will take 0, that is, only the gray values of the 4 points around (x i , y i ) determine the gray value of the target pixel point. And when and are smaller, the influence is greater (that is, closer to the point (n, m)), and the weight is greater

[0108] S2-2. Take the output of the spatial transformation network as the input and pass it to a forward feature extraction network F forward . This network model is used to extract the features of the current frame, and the output features are represented by mapping as follows:

[0109] x = F forward

[0110] Among them, x represents the features captured in the current frame image, and F forward represents the forward feature extraction network, which consists of 9 convolutional layers and 5 max pooling layers, as follows:

[0111] First, it passes through Conv1 (the first convolutional layer), the input size of which is a 224x224 RGB image, the convolutional kernel size is 3x3, the number of convolutional kernels is 64, and the stride is 1; then it passes through Conv2 (the second convolutional layer), the input of which is the feature image extracted by Conv1, with a size of 224x224x64, the convolutional kernel size is 3x3, and the number of convolutional kernels is 64; then it passes through a MaxPooling1 (the first pooling layer), and the size of the pooling layer MaxPooling1 is 2, and the stride is 2

[0112] Continue to pass through Conv3 (the 3rd convolutional layer), the input of which is the feature image output by MaxPooling1, with a size of 112x112x64, the convolutional kernel size is 3x3, the number of convolutional kernels is 128, and the stride is 1; pass through Conv4 (the 4th convolutional layer), the input of which is the feature image extracted by Conv3, with a size of 112x112x128, the convolutional kernel size is 3x3, and the number of convolutional kernels is 128, and the stride is 1

[0113] Then, through MaxPooling2 (the second pooling layer), the input is the feature image extracted by Conv4, with a size of 112x112x128. The size of the pooling layer MaxPooling2 is 2, and the stride is 2.

[0114] Then, through Conv5 (the fifth convolutional layer), the input is the feature image output by MaxPooling2, with a size of 56x56x128. The size of the convolutional kernel is 7x7, the number of convolutional kernels is 256, and the stride is 1.

[0115] MaxPooling3 (the third pooling layer), the input is the feature image extracted by Conv5, with a size of 56x56x256. The size of the pooling layer MaxPooling3 is 2, and the stride is 2.

[0116] Continue through Conv6 (the sixth convolutional layer), the input is the feature image output by MaxPooling3, with a size of 28x28x256. The size of the convolutional kernel is 5x5, the number of convolutional kernels is 512, and the stride is 1; through MaxPooling4 (the fourth pooling layer), the input is the feature image extracted by Conv6, with a size of 28x28x512. The size of the pooling layer MaxPooling4 is 2, and the stride is 2.

[0117] Continue through Conv7 (the seventh convolutional layer), the input is the feature image output by MaxPooling4, with a size of 14x14x512. The size of the convolutional kernel is 3x3, the number of convolutional kernels is 512, and the stride is 1; through Conv8 (the eighth convolutional layer), the input is the feature image extracted by Conv7, with a size of 14x14x512. The size of the convolutional kernel is 3x3, the number of convolutional kernels is 512, and the stride is 1; then through Conv9 (the ninth convolutional layer), the input is the feature image extracted by Conv8, with a size of 14x14x512. The size of the convolutional kernel is 3x3, the number of convolutional kernels is 512, and the stride is 1; finally, through a MaxPooling5 (the fifth pooling layer), the input is the feature image extracted by Conv9, with a size of 28x28x512. The size of the pooling layer MaxPooling5 is 2, and the stride is 2.

[0118] S2-3. As Figure 5 shown, take the features extracted by the previous feature extraction network as the input and pass them to three consecutive attention residual networks (CBARM). CBARM consists of a 1D channel attention mechanism M c ∈R C×1×1 and a 2D spatial attention mechanism M s ∈R 1×H×Wand a residual network. Denote the feature map input extracted by the previous feature extraction network as

[0119] F ∈ R C×H×W ,, the entire attention mechanism process can be summarized by the following formula:

[0120]

[0121]

[0122] where denotes element-wise multiplication. When performing element-wise multiplication, the attention values are broadcast accordingly. The channel attention values are broadcast along the spatial dimension, and vice versa. F′ is the output along the channel attention network, and F″ is the final refined output.

[0123] Preferably, S2-3 includes the following sub-steps:

[0124] S2-3-1: First, aggregate the spatial information of a feature map through average pooling and max pooling operations to generate two different spatial context descriptors: and These two descriptors respectively represent the features pooled by average pooling and max pooling. Then, these two descriptors are fed forward into a shared network to generate a channel attention map M c ∈ R C×1×1 . This shared network consists of a multi-layer perceptron (MLP) with one hidden layer. To reduce parameter overhead, the activation size of the hidden layer is set to R C / r×1×1 , where r is the reduction rate. After each spatial context descriptor is processed by the shared network, the output feature vectors are fused using element-wise addition. The channel attention calculation method is as shown in the following formula:

[0125]

[0126] where σ is the sigmoid function, and the weight sizes of the MLP are W0 ∈ R C / r×C and W1 ∈ R C / C×r , which are shared by the two inputs of the average-pooled feature and the max-pooled feature, and W0 follows the Relu activation function, that is, the pooled feature needs to be processed by the Relu function before being input into the MLP.

[0127] S2-3-2: Take the channel attention map as the input and pass it into the spatial attention mechanism. The specific process is as described above where the spatial attention mechanism can generate a spatial attention map by utilizing the spatial correlation of features. Different from channel attention, spatial attention focuses on the information-rich parts between different positions in the feature map, which is complementary to channel attention.

[0128] To calculate the spatial attention, first, average pooling and max pooling operations are performed along the channel dimension, and the two are concatenated to generate an efficient feature descriptor. Performing pooling operations along the channel dimension has been proven to effectively highlight information-rich regions. For the concatenated feature descriptor, a convolutional layer is used to generate a spatial attention map M s (F) ∈ R H×W , which encodes which regions are highlighted or suppressed.

[0129] Two 2D maps are generated through the two pooling operations: and Each map represents the features averaged and max-pooled along the channels respectively. Then the two are concatenated and convolved by a standard convolutional layer to generate a 2D spatial attention map. The spatial attention calculation method is as follows:

[0130]

[0131] where σ is the sigmoid function, and f 7×7 represents a convolution operation with a 7×7 convolutional kernel.

[0132] S2-3-3. Integrate the ResBlock in ResNet into the channel-spatial attention mechanism CBAM to form the attention residual network (CBARM). One end of the BesBlock is inserted at the front end of the channel attention mechanism, and the other end is inserted at the back end of the spatial attention mechanism. The specific positions are shown in the attached drawings. The formula is as follows:

[0133] y l = h(x l ) + F(x l , W l )

[0134] x l+1 = σ(y l )

[0135] where x l and x l+1 represent the input and output of the residual unit, that is, the input and output of the CBARM attention residual network. F is the learned residual, and h(x l ) = x l represents the identity mapping, and σ is the Relu activation function.

[0136] Then the features learned from the shallow layer l to the deep layer L are:

[0137]

[0138] Among them, the deep layer L can be expressed as the sum of any shallower layer l and the residual part between them. L is the unit cumulative sum of the features of each residual block.

[0139] S3. Perform feature fusion in the form of concat on the density map of the current frame output by the front-end encoder model in S2 and the density map of the previous frame of the current frame. Then, the single output channel after concat is as follows, and the specific formula is as follows:

[0140]

[0141] Among them, Z concat represents the single output channel after concat, X i corresponds to the channel of the density map of the current frame of the output, Y i corresponds to the channel of the density map of the previous frame of the current frame. * represents convolution, K represents the corresponding convolution kernel, i is the channel of the feature map, and c is the total number of channels of the feature map.

[0142] S4. Take Z concat after feature fusion in S3 as the input and pass it to the backend decoder model of the network. Finally, the human inflow map is output. The decoder model consists of four dilated convolutions, which are specifically as follows.

[0143] Preferably, S4 is divided into the following steps:

[0144] S4-1. Input the density map of the current frame output by the front-end encoder model and the density map of the previous frame of the current frame through feature fusion into a two-dimensional dilated convolution with a kernel size of 3, a filter number of 512, and a dilation rate of 2. The definition of the two-dimensional dilated convolution is as follows:

[0145]

[0146] where y(m,n) is the output after dilated convolution of the fused feature map input x(m,n) and the filter w(i,j). The length and width are M and N respectively. i and j represent rows and columns respectively. The parameter r is the dilation rate of the dilated convolution. Here, if r = 1, the dilated convolution becomes a normal convolution.

[0147] S4-2. Take the feature map y(m,n) output in S4-1 as the input and pass it to another dilated convolution, where the kernel size of the dilated convolution is 3, the filter number is 256, and the dilation rate is 2. The feature map output after the dilated convolution is p(m,n).

[0148] S4-3. Take the feature map p(m,n) output by S4-2 as the input and pass it into another dilated convolution, where the kernel size of the dilated convolution is 3, the number of filters is 128, and the dilation rate is 2. The feature map output after the dilated convolution is q(m,n).

[0149] S4-3. Take the feature map q(m,n) output by S4-3 as the input and pass it into another dilated convolution, where the kernel size of the dilated convolution is 3, the number of filters is 64, and the dilation rate is 2. The feature map output after the dilated convolution is s(m,n).

[0150] S5. Establish a loss function and train it using the obtained training set of true density maps.

[0151] Preferably, step S5 is divided into the following steps:

[0152] S5-1. Obtaining the true density map. The true density map is generated by marking the coordinates of the center positions of human heads in the crowd area of the true image and performing Gaussian smoothing operation. The generation process of the true density map is as follows.

[0153] In each input t-th frame image I t a set of s t two-dimensional points are annotated, which represent the positions of human heads in the scene. By convolving the image containing 1 at these positions and 0 at other positions with a Gaussian kernel having a mean μ and a standard deviation σ the corresponding ground truth density map is obtained The formula is as follows:

[0154]

[0155] where P j represents the center of position j;

[0156] S5-2. To strengthen the constraints of the two equations in S1-2, the combined loss function is used as the weighted term of the two loss functions, and the formula is expressed as:

[0157]

[0158]

[0159]

[0160] where is the ground truth density value, that is, the number of people at time t and position j, and α is set to 1. During training, three consecutive frames are used to evaluate L combi .

[0161] S6. Take the video frame to be processed as the input, repeat steps S2 - S5, and output the calculation result of the number of people in the video frame to be predicted.

[0162] Example:

[0163] 1. Evaluation Metrics

[0164] Similar to the mainstream crowd counting methods, here the Mean Absolute Error (MAE) and the Mean Square Error (MSE) are used as the metrics to evaluate the performance. Let T be the number of test images, and Z i and be the correct annotation number and the predicted number of the i-th image respectively. The calculation formulas of MAE and MSE are as follows:

[0165]

[0166]

[0167] 2. Experimental Datasets

[0168] Currently, the commonly used datasets for video crowd counting are relatively small. Here, a new large-scale crowd counting dataset, the Fudan-ShanghaiTech dataset (FDST), is used. FDST contains 100 videos captured from 13 different scenarios, consisting of more than 330,000 pedestrians, and it is the largest video crowd counting dataset. It is composed of two parts, A and B. Part A is obtained from the Internet, while part B is selected from street areas in Shanghai. The training set of the FDST dataset consists of 60 videos and 9,000 frame images, and the test set contains 40 videos and 6,000 frame images.

[0169] 3. Experimental Procedures and Results

[0170] Use the combined loss function L combi to train the proposed model that combines the attention and spatial transformation network Compare the data results obtained by this method with other methods that also use the FDST dataset. The counting accuracy has been significantly improved, which also proves the effectiveness of the proposed model in the field of video crowd counting. The experimental results are shown in Table 1:

[0171] Table 1

[0172]

[0173] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings, but the present invention is not limited to the described embodiments. For those skilled in the art, without departing from the principles and spirit of the present invention, various changes, modifications, substitutions, and variations can be made to these embodiments including components, and still fall within the protection scope of the present invention.

Claims

1. A video crowd counting method based on the combination of attention and spatial transformation network, characterized in that, It includes the following steps: S1. Obtain the current frame collected by the camera and the previous frame of the current frame to get two consecutive images; S2. Input the current frame and the previous frame image of the current frame into the network 's front-end encoding encoder model respectively, and output the density map of the current frame and the density map of the previous frame of the current frame, where the encoder model consists of a spatial transformation network STN, a residual network, and an attention mechanism; S3. Perform feature fusion in the form of concat on the density map of the current frame output by the front-end encoder model in S2 and the density map of the previous frame of the current frame. Then, the single output channel after concat is as follows: Among them, Z concat represents a single output channel after concat, X i corresponds to the channel of the density map of the current frame of the output, Y i corresponds to the channel of the density map of the previous frame of the current frame, * represents convolution, K i represents the corresponding convolution kernel, i is the channel of the feature map, and c is the total number of channels of the feature map; S4. Pass Z concat as input to the backend decoder model of the network to output the flow graph of people; The decoder model consists of four dilated convolutions; S5. Establish a loss function and train the network using the obtained true density map training set for training; S5-1. Obtaining the real density map. The real density map is generated by marking the coordinates of the center positions of human heads in the crowd area of the real image and performing Gaussian smoothing operation. The generation process of the real density map is as follows: In the t-th frame image I of each input t a set of s t two-dimensional points are annotated to represent the positions of human heads in the scene; the positions containing these points are set to 1 in the density map, and the positions not containing these points are set to 0 in the density map. Then, the density map is convolved with a Gaussian kernel having a mean μ and a standard deviation σ to obtain the corresponding ground-truth density map The formula is as follows: where P j represents the center at position j; S5-2. The loss function is: Among them, is the number of people at time t and position j, that is, the ground truth density map, α is the weight, is the number of people moving from location i to position j between time t-1 and t, and N(j) is the adjacent position of the j-th position; S6. Take the video frame to be processed as the input, repeat steps S2 - S5, and output the calculation result of the number of people in the video frame to be predicted.

2. The video crowd counting method based on the combination of attention and spatial transformation network according to claim 1, wherein The specific process of S2 is as follows: S2-1. Input the current frame into the Spatial Transformer Network (STN). The entire spatial transformer network includes three parts: a regression network, a grid generator, and a sampler; S2-2. Use the output of the spatial transformation network as the input and pass it to a forward feature extraction network F forward The output features are represented by mapping as follows: x = F forward Among them, x represents the feature captured in the current frame image, and F forward indicates that the previous feature extraction network is composed of multiple convolutional layers and max pooling layers; S2-3. Take the features extracted by the previous feature extraction network as input and pass them to three consecutive CBARM (Convolutional Block Attention Residual Module). CBARM consists of a one-dimensional channel attention mechanism M c ∈R C×1×1 , a two-dimensional spatial attention mechanism M s ∈R 1×H×W , and a residual network. Denote the feature map input extracted by the previous feature extraction network as F ∈ R C×H×W . The entire attention mechanism process is summarized as shown in the following formula: Among them represents bitwise multiplication. When performing bitwise multiplication, the attention value is broadcast. The channel attention value is broadcast along the spatial dimension, and vice versa; F ′ is the output along the channel attention network, F ″ is the refined output.

3. A method for video crowd counting based on the combination of attention and spatial transformation network according to claim 2, characterized in that, In S2-1, the regression network is a network that regresses the transformation parameter θ. The input U is the feature image, and then the transformation parameter θ is output after passing through the hidden network layer; The grid generator is a sampling grid constructed based on the predicted transformation parameter θ, which is the output obtained by sampling and transforming a set of points in the input image; What the grid generator obtains is a mapping relationship Assume that the coordinates of each pixel of the feature image U are The coordinates of each pixel of the output V are Spatial transformation function Is a two-dimensional affine transformation function, A θ Represents the matrix composed of the transformation parameters θ, then And The corresponding relationship is: The sampler uses the output of the grid generator and the input feature map as inputs at the same time to obtain the output feature map result after the feature map is transformed. The formula is as follows: where V i c is the gray value of a point on the c-th channel of the output feature map, is the gray value of the point (n, m) on the c-th channel of the input; when or is greater than 1, the corresponding max() term will take 0, that is, the gray values of the 4 surrounding points of (x i , y i ) determine the gray value of the target pixel point; and when and are smaller, the influence is greater and the weight is greater.

4. A video crowd counting method based on the combination of attention and spatial transformation network according to claim 3, characterized in that In the regression network, the hidden network layer consists of a fully connected or convolutional network, plus a regression layer.

5. A method for video crowd counting based on the combination of attention and spatial transformation network according to claim 2, characterized in that, S2-3 includes the following sub-steps: S2-3-1. First, aggregate the spatial information of a feature map through average pooling and max pooling operations to generate two different spatial context descriptors: and represent the features obtained by average pooling and max pooling respectively; Second, these two descriptors are fed forward into a network shared by both to generate a channel attention map M c ∈R C×1×1 , after each spatial context descriptor is processed by the shared network, the output feature vectors are fused using bitwise addition, and the channel attention calculation method is as shown in the formula: where σ is the sigmoid function, and the weight sizes of the MLP are W0 ∈ R C / r×C and W1 ∈ R C / C×r , which are shared by the two inputs of the average pooling feature and the max pooling feature, and W0 follows the Relu activation function, that is, the pooling feature is processed by the Relu function before being input into the MLP; S2-3-2. Use the channel attention map as the input and feed it into the spatial attention mechanism. The specific process is as described above. Generate a spatial attention map by leveraging the spatial correlation of features. Calculate the spatial attention. Perform average pooling operation and max pooling operation along the channel dimension, and concatenate the two to generate an efficient feature descriptor. For the concatenated feature descriptor, use a convolutional layer to generate a spatial attention map M s (F)∈R H×W , which encodes the regions that are highlighted or suppressed; Two two-dimensional 2D maps are generated through two pooling operations: and Each 2D map represents the features averaged along the channels and the features max-pooled, respectively. The two are concatenated and convolved by a standard convolutional layer to generate a 2D spatial attention map. The spatial attention calculation method is as follows: where σ is the sigmoid function, and f 7×7 represents a convolution operation with a convolution kernel size of 7×7; S2-3-3. Integrate the residual block ResBlock in the Residual Neural Network (ResNet) into the Channel-Spatial Attention Module (CBAM) to form CBARM; One end of the residual block ResBlock is inserted at the front end of the channel attention mechanism, and the other end is inserted at the back end of the spatial attention mechanism. The formula is as follows: y l = h(x l ) + F(x l , W l ) x l+1 = Φ(y l ) where x l and x l+1 represent the input and output of the residual module, W l is the weight, that is, the input and output of CBARM, F is the learned residual, h(x l ) = x l represents the identity mapping, and Φ is the Relu activation function; Then, the features learned from the shallow layer l to the deep layer L are: Among them, the deep layer L is expressed as the sum of any shallower l layer and the residual part between them, and W i is the weight, and L is the unit cumulative sum of the features of each residual block.

6. A video crowd counting method based on the combination of attention and spatial transformation network according to claim 5, characterized in that In S2-3-1, the shared network consists of a Multi-Layer Perceptron (MLP) with one hidden layer.

Citation Information

Patent Citations

  • Action video recognition method combining hybrid convolution residual network and attention

    CN112149504A

  • No-reference video quality evaluation method fusing spatio-temporal features

    CN112954312A