A crowd counting method based on feature interaction

By using the technical means of feature interaction and multi-scale attention module in the crowd counting method, the problems of scale change and simple attention mechanism are solved, and the accuracy and robustness of crowd density estimation are improved.

CN115272957BActive Publication Date: 2025-07-01YANSHAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210805244.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-08
Publication Date
2025-07-01
Estimated Expiration
2042-07-08

AI Technical Summary

Technical Problem

The population counting method based on density estimation has scale changes, which affects the accuracy of the counting results, and the attention mechanism is too simple, which affects the counting performance.

Method used

Using a crowd counting method based on feature interaction, a three-layer semantic feature map is extracted through a deep neural network model, and its input semantic interaction structure and multi-scale attention module are fused and processed to generate the main feature map for crowd density estimation.

Benefits of technology

The accuracy and robustness of population density estimation are improved, scale limitations and feature similarity problems in traditional methods are overcome, and high-quality population density maps are generated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115272957B_ABST
    Figure CN115272957B_ABST
Patent Text Reader

Abstract

The present invention discloses a crowd counting method based on feature interaction, belonging to the technical field of image processing, which includes the following steps: inputting the original image into a deep neural network model for feature extraction; sending three-layer semantic feature maps into a semantic interaction structure; respectively inputting the fused three-layer semantic feature maps into a multi-scale attention module; performing upsampling and channel adjustment on the scale-aware information features corresponding to the high-level semantic feature map and fusing them with the scale-aware information features corresponding to the middle-level semantic feature map; performing upsampling and channel adjustment on the fused features and fusing them with the scale-aware information features corresponding to the low-level semantic feature map; inputting the main feature map for crowd density estimation into the backend network of the deep neural network model to obtain a crowd density estimation map and a crowd counting result. The present invention can effectively improve the accuracy of crowd density estimation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a crowd counting method based on feature interaction. Background Art

[0002] Crowd counting is an important research content in the fields of computer vision and intelligent monitoring, and its purpose is to estimate the number of people in an image or video scene. It has a wide range of applications in fields such as security monitoring, traffic management, and urban planning. For example, during the epidemic, controlling the crowd density can reduce the probability of cluster transmission; in areas with a high concentration of people such as scenic spots, stadiums, and squares, warning messages can be issued to prevent stampede accidents, etc. In recent years, crowd counting methods based on convolutional neural networks have become the mainstream methods for crowd counting. The basic idea is to use a convolutional neural network to generate an estimated density map, assign a density value to each pixel, and the sum of the density values of the density map is recorded as the total number of people in the scene.

[0003] Currently, the scale change problem caused by differences in shooting distance and angle seriously affects the accuracy of the counting results. There will be drastic scale changes within the same crowd image or between different images, and this drastic scale change poses a huge challenge to the crowd density prediction based on convolutional neural networks. To address the above problems, the inventor proposed a crowd counting method and system based on density estimation in the invention patent "A Crowd Counting Method and System Based on Density Estimation" (CN113538402B), which integrates multi-layer semantic information and multi-scale information and achieves better counting results.

[0004] However, the crowd counting method and system based on density estimation have the following problems:

[0005] 1. Only simple fusion of multi-layer semantic information and multi-scale information is performed, with a simple structure, and the limitations of the network scale are not considered, which makes the semantic information and scale information extracted by this method insufficient.

[0006] 2. In dealing with the problem of feature similarity, the attention mechanism used by this method is too simple, and the importance of cross-dimensional information is not considered, affecting the counting performance.

[0007] To solve the above problems of the crowd counting method and system based on density estimation, it is necessary to further optimize the crowd counting method and system proposed by the present inventor. Summary of the Invention

[0008] The technical problem to be solved by the present invention is to provide a crowd counting method based on feature interaction, which can effectively solve the scale change problem in the crowd counting task, help generate a high-quality crowd density map, improve the counting performance of the multi-column network, have high accuracy and good robustness, and effectively improve the accuracy of crowd density estimation.

[0009] To solve the above technical problems, the technical solution adopted by the present invention is as follows:

[0010] A crowd counting method based on feature interaction, comprising the following steps:

[0011] Input the original image into a deep neural network model for feature extraction to obtain three layers of semantic feature maps, where the three layers of semantic feature maps include a low-level semantic feature map, a middle-level semantic feature map, and a high-level semantic feature map;

[0012] Send the three layers of semantic feature maps into a semantic interaction structure to obtain the corresponding three layers of fused semantic feature maps;

[0013] Input the fused three layers of semantic feature maps into a multi-scale attention module respectively to obtain the scale-aware information features of the corresponding semantic feature maps;

[0014] Upsample and adjust the channels of the scale-aware information features corresponding to the high-level semantic feature map and fuse them with the scale-aware information features corresponding to the middle-level semantic feature map;

[0015] Upsample and adjust the fused features and fuse them with the scale-aware information features corresponding to the low-level semantic feature map to obtain the main feature map for crowd density estimation;

[0016] Input the main feature map for crowd density estimation into the backend network of the deep neural network model to obtain the crowd density estimation map and the crowd counting result.

[0017] A further improvement of the technical solution of the present invention lies in: inputting the original image into a deep neural network model for feature extraction to obtain three layers of semantic feature maps, including the following steps:

[0018] Input the original image into a deep neural network model, where the deep neural network model includes two consecutive convolutional layers, a pooling layer, two convolutional layers, a pooling layer, three convolutional layers, and a pooling layer, to obtain a low-level semantic feature map; the number of channels of the feature maps generated by each convolutional layer is 64, 64, 128, 128, 256, 256, and 256 in sequence from the input to the output direction; the convolutional kernel size of the convolutional layer is 3*3; the stride of the pooling layer is 2;

[0019] Input the low-level semantic feature map into the deep neural network model, and successively pass through three convolutional layers and one pooling layer to obtain the middle-level semantic feature map; the number of channels of the feature map generated by each convolutional layer is 512; the size of the convolutional kernel of the convolutional layer is 3*3; the stride of the pooling layer is 2;

[0020] Input the middle-level semantic feature map into the deep neural network model, and pass through three convolutional layers to obtain the high-level semantic feature map; the number of channels of the feature map generated by each convolutional layer is 512; the size of the convolutional kernel of the convolutional layer is 3*3.

[0021] A further improvement of the technical solution of the present invention is that: sending the three-layer semantic feature maps into a semantic interaction structure, including the following steps:

[0022] Send the high-level semantic feature map into the semantic interaction structure to obtain the fused semantic feature map corresponding to the high-level semantic feature map;

[0023] Send the middle-level semantic feature map into the semantic interaction structure, and interact with the fused semantic feature map corresponding to the high-level semantic feature map to obtain the fused semantic feature map corresponding to the middle-level semantic feature map;

[0024] Send the low-level semantic feature map into the semantic interaction structure, and interact with the fused semantic feature map corresponding to the middle-level semantic feature map to obtain the fused semantic feature map corresponding to the low-level semantic feature map.

[0025] A further improvement of the technical solution of the present invention is that: the semantic interaction structure includes:

[0026] Upsample the high-level semantic feature map using bilinear interpolation;

[0027] Connect the result of upsampling the high-level semantic feature map with the middle-level semantic feature map in channels to obtain the intermediate feature corresponding to the middle-level semantic feature map;

[0028] Perform feature fusion on the intermediate feature through two 3*3 convolutions to obtain the fused semantic feature map of the middle-level semantic feature map.

[0029] Upsample the obtained fused semantic feature map of the middle-level semantic feature map using bilinear interpolation;

[0030] Connect the result of upsampling the middle-level semantic feature map with the low-level semantic feature map in channels to obtain the intermediate feature corresponding to the low-level semantic feature map;

[0031] Perform feature fusion on the intermediate feature through two 3*3 convolutions to obtain the fused semantic feature map of the low-level semantic feature map.

[0032] A further improvement of the technical solution of the present invention lies in that: the multi-scale attention module includes 4 branches with different receptive fields, an operation of concatenating the results of the 4 branches in the channel dimension, a convolutional layer, an additional global channel attention mechanism, and an element-wise multiplication operation; each branch sequentially includes a convolutional layer, a dilated convolutional layer, a multi-scale interaction structure, and a global spatial attention mechanism;

[0033] The global channel attention mechanism includes:

[0034] Perform dimension transposition and tiling operations on the input feature map in the three dimensions of channel, height, and width to obtain the feature map after the dimension transposition and tiling operations;

[0035] Use linear transformation to reduce the channel dimension of the feature map after the dimension transposition and tiling operations to 1 / 4 of the original, and use the ReLU activation function for non-linear transformation, and then use linear transformation to change the channel dimension to the same as the original feature map to amplify the dependence of cross-dimensional features on the channel dimension, obtaining the feature map processed by the multi-layer perceptron;

[0036] Perform dimension transposition and reshaping operations on the feature map processed by the multi-layer perceptron in the three dimensions of channel, height, and width to obtain the feature map after the dimension transposition and reshaping operations;

[0037] Perform Sigmoid function transformation on the feature map after the dimension transposition and reshaping operations, and perform element-wise multiplication operation with the original input feature map to obtain the output feature map;

[0038] The multi-scale interaction structure includes:

[0039] Connect the result of the interaction of the small receptive field feature map with the large receptive field feature map in the channel dimension to obtain an intermediate feature;

[0040] Use a 3×3 convolution to fuse the features of the intermediate feature to obtain a fused multi-scale interaction feature map;

[0041] The global spatial attention mechanism includes:

[0042] Pass the input feature map through a dilated convolutional layer with a convolutional kernel size of 3×3 and a dilation rate the same as that of the dilated convolution used to extract multi-scale features inside the branch to obtain a feature map with the number of channels reduced to 1 / 4 of the original;

[0043] Pass the feature map with the number of channels reduced to 1 / 4 of the original through a dilated convolutional layer with a convolutional kernel size of 3×3 and a dilation rate the same as that of the dilated convolution used to extract multi-scale features inside the branch to obtain a feature map with the same number of channels as the original input feature map;

[0044] Perform a Sigmoid function transformation on the feature map with the same number of channels as the original input feature map, and perform an element-wise multiplication operation with the original input feature map to obtain an output feature map;

[0045] The branch includes a first branch, a second branch, a third branch, and a fourth branch;

[0046] The feature map passing through the first branch includes:

[0047] Pass the feature map through a convolutional layer with a kernel size of 1*1 to obtain a feature map with the number of channels reduced to 1 / 4 of the original;

[0048] Pass the feature map with the number of channels reduced to 1 / 4 of the original through a dilated convolutional layer with a kernel size of 3*3 and a dilation rate of 1 to obtain a scale feature map with a receptive field of 3*3;

[0049] Feed the scale feature map with a receptive field of 3*3 into the multi-scale interaction structure to obtain a feature map after multi-scale interaction;

[0050] Feed the feature map after multi-scale interaction into the global spatial attention mechanism to obtain a feature map with new feature weights assigned;

[0051] The feature map passing through the second branch includes:

[0052] Pass the feature map through a convolutional layer with a kernel size of 1*1 to obtain a feature map with the number of channels reduced to 1 / 4 of the original;

[0053] Pass the feature map with the number of channels reduced to 1 / 4 of the original through a dilated convolutional layer with a kernel size of 3*3 and a dilation rate of 2 to obtain a scale feature map with a receptive field of 7*7;

[0054] Feed the scale feature map with a receptive field of 7*7 into the multi-scale interaction structure to obtain a feature map after multi-scale interaction;

[0055] Feed the feature map after multi-scale interaction into the global spatial attention mechanism to obtain a feature map with new feature weights assigned;

[0056] The feature map passing through the third branch includes:

[0057] Pass the feature map through a convolutional layer with a kernel size of 1*1 to obtain a feature map with the number of channels reduced to 1 / 4 of the original;

[0058] Pass the feature map with the number of channels reduced to 1 / 4 of the original through a dilated convolutional layer with a kernel size of 3*3 and a dilation rate of 3 to obtain a scale feature map with a receptive field of 11*11;

[0059] Feed the scale feature map with a receptive field of 11*11 into the multi-scale interaction structure to obtain the feature map after multi-scale interaction;

[0060] Feed the feature map after multi-scale interaction into the global spatial attention mechanism to obtain the feature map with new feature weights assigned;

[0061] Pass the feature map through the fourth branch, including:

[0062] Pass the feature map through a convolutional layer with a kernel size of 1*1 to obtain a feature map with the number of channels reduced to 1 / 4 of the original;

[0063] Pass the feature map with the number of channels reduced to 1 / 4 of the original through a dilated convolutional layer with a kernel size of 3*3 and a dilation rate of 4 to obtain a scale feature map with a receptive field of 15*15;

[0064] Feed the scale feature map with a receptive field of 15*15 into the multi-scale interaction structure to obtain the feature map after multi-scale interaction;

[0065] Feed the feature map after multi-scale interaction into the global spatial attention mechanism to obtain the feature map with new feature weights assigned.

[0066] A further improvement of the technical solution of the present invention is that the fused three-layer semantic feature maps are respectively input into the multi-scale attention module to obtain the scale-aware information features corresponding to the semantic feature maps, including the following steps:

[0067] Input the fused low-level semantic feature map into the four branches of the multi-scale attention module to obtain four-scale low-level semantic feature maps;

[0068] Connect the four-scale low-level semantic feature maps on the channel dimension, perform feature fusion using a 3*3 convolution, and then multiply the result by the feature information obtained from the global channel attention mechanism to obtain the scale-aware information features corresponding to the low-level semantic feature map;

[0069] Input the fused middle-level semantic feature map into the four branches of the multi-scale attention module to obtain four-scale middle-level semantic feature maps;

[0070] Connect the four-scale middle-level semantic feature maps on the channel dimension, perform feature fusion using a 3*3 convolution, and then multiply the result by the feature information obtained from the global channel attention mechanism to obtain the scale-aware information features corresponding to the middle-level semantic feature map;

[0071] Input the fused high-level semantic feature map into the four branches of the multi-scale attention module to obtain four-scale high-level semantic feature maps;

[0072] Connect the high-level semantic feature maps of the four scales in the channel dimension, perform feature fusion using a 3×3 convolution, and then multiply the result by the feature information obtained from the global channel attention mechanism to obtain the scale-aware information features corresponding to the high-level semantic feature maps.

[0073] A further improvement of the technical solution of the present invention is: upsample and adjust the channels of the scale-aware information features corresponding to the high-level semantic feature maps and fuse them with the scale-aware information features corresponding to the middle-level semantic feature maps, including the following steps:

[0074] Upsample the scale-aware information features corresponding to the high-level semantic feature maps using bilinear interpolation, and adjust the channels using a 1×1 convolution to obtain the first feature map;

[0075] Perform an element-wise addition operation on the first feature map and the scale-aware information features corresponding to the middle-level semantic feature maps to obtain the fused features.

[0076] A further improvement of the technical solution of the present invention is: upsample and adjust the channels of the fused features and fuse them with the scale-aware information features corresponding to the low-level semantic feature maps to obtain the main feature map for crowd density estimation, including the following steps:

[0077] Upsample the fused features using bilinear interpolation, and adjust the channels using a 1×1 convolution to obtain the second feature map;

[0078] Perform an element-wise addition operation on the second feature map and the scale-aware information features corresponding to the low-level semantic feature maps to obtain the main feature map for crowd density estimation.

[0079] A further improvement of the technical solution of the present invention is: input the main feature map for crowd density estimation into the backend network of the deep neural network model to obtain the crowd density estimation map and the crowd counting result, including the following steps:

[0080] Input the main feature map for crowd density estimation into two convolutional layers to obtain the crowd density estimation map and the crowd counting result; the number of channels of the feature maps generated by each convolutional layer is 64 and 1 in the order from input to output; the convolution size of both convolutional layers is 3×3.

[0081] Due to the adoption of the above technical solution, the technical progress achieved by the present invention is:

[0082] 1. The present invention extracts rich multi-scale information through the multi-scale attention module, that is, it improves the ability to extract multi-scale information and the sensitivity to valuable information by using the interaction structure and the attention mechanism, and overcomes the scale limitation and feature similarity problems in traditional multi-column networks.

[0083] 2. The present invention conducts interaction and fusion on semantic information at different levels of the backbone network through a semantic information fusion module, providing richer detailed features, enhancing the feature aggregation ability of the network, and improving the utilization efficiency of the backbone network. BRIEF DESCRIPTION OF THE DRAWINGS

[0084] Figure 1 is a schematic flowchart of the method for crowd counting based on feature interaction of the present invention;

[0085] Figure 2 is a schematic diagram of the overall structure of the deep neural network model of the present invention;

[0086] Figure 3 is a schematic diagram of estimating crowd density using the method for crowd counting based on feature interaction of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0087] The present invention will be further described in detail below with reference to the drawings and embodiments:

[0088] As Figure 1 shown, the method for crowd counting based on feature interaction includes the following steps:

[0089] Step 100: Input the original image into the deep neural network model for feature extraction to obtain low-level, middle-level, and high-level semantic feature maps. Thirteen convolutional layers and four max-pooling layers are involved in this process. Specifically, first, it passes through seven convolutional layers and three max-pooling layers, which are in sequence: two convolutional layers, one pooling layer, two convolutional layers, one pooling layer, three convolutional layers, and one pooling layer; the convolutional size of the convolutional layers is 3*3, and the number of channels of the generated feature maps is in sequence: 64, 64, 128, 128, 256, 256, and 256; the stride of the three pooling layers is 2. The generated low-level feature map continues to be input into the deep neural network model. Specifically, it includes three convolutional layers and one max-pooling layer, which are in sequence: three convolutional layers, one pooling layer; the convolutional size of the convolutional layers is 3*3, and the number of channels of the generated feature maps is in sequence: 512, 512, and 512; the stride of the pooling layer is 2. The generated middle-level feature map continues to be input into the deep neural network model. Specifically, it includes three convolutional layers; the convolutional size of the convolutional layers is 3*3, and the number of channels of the generated feature maps is in sequence: 512, 512, and 512; the stride of the pooling layer is 2. Finally, low-level, middle-level, and high-level semantic feature maps are obtained.

[0090] Step 200: Feed the low-level, middle-level, and high-level semantic feature maps into the semantic interaction structure to obtain the corresponding fused semantic feature maps. Specifically, feed the high-level semantic feature map into the semantic interaction structure to obtain the fused semantic feature map corresponding to the high-level semantic feature map; feed the middle-level semantic feature map into the semantic interaction structure to interact with the fused semantic feature map corresponding to the high-level semantic feature map to obtain the fused semantic feature map corresponding to the middle-level semantic feature map; feed the low-level semantic feature map into the semantic interaction structure to interact with the fused semantic feature map corresponding to the middle-level semantic feature map to obtain the fused semantic feature map corresponding to the low-level semantic feature map.

[0091] The following is a specific description of the semantic interaction structure:

[0092] The semantic interaction structure upsamples the high-level semantic feature map using bilinear interpolation, then concatenates the result of the upsampling with the low-level semantic feature map in the channel dimension to obtain intermediate features, and performs feature fusion through two 3×3 convolutions to obtain the fused semantic feature map.

[0093] Step 300: Feed the fused three-layer semantic feature maps into the multi-scale attention module respectively to obtain the scale-aware information features of the corresponding semantic feature maps. The multi-scale attention module includes 4 branches with different receptive fields, namely the first branch, the second branch, the third branch, and the fourth branch, and each branch can perceive information at different scales. Then, the results of the 4 branches are concatenated in the channel dimension, and after feature fusion using a 3×3 convolution, they are multiplied by the feature information obtained by the global channel attention mechanism to obtain the scale-aware information features of the corresponding semantic feature maps.

[0094] The following is a specific description of the global channel attention mechanism:

[0095] First, perform dimension transposition and flattening operations on the input feature map in the three dimensions of channel, height, and width to obtain the feature map after dimension transposition and flattening; then use a linear transformation to reduce the channel dimension of the feature map to 1 / 4 of the original, and perform a non-linear transformation using the ReLU activation function, and then use a linear transformation to change the channel dimension to the same as the original feature map to amplify the dependence of the cross-dimensional features on the channel dimension to obtain the feature map processed by the multi-layer perceptron; then perform dimension transposition and reshaping operations on the three dimensions of channel, height, and width to obtain the feature map after dimension transposition and reshaping, then perform a Sigmoid function transformation, and perform an element-wise multiplication operation with the original input feature map to obtain the output feature map.

[0096] The expression of the Sigmoid function is: Where z is each element of the operation result, here referring to the feature map after 1*1 convolution processing, and f(z) is the result after Sigmoid transformation of each element.

[0097] The multi-scale interaction structure is specifically described as follows:

[0098] First, connect the result of the interaction of the small receptive field feature map with the large receptive field feature map in the channel dimension to obtain intermediate features; then, perform feature fusion on the intermediate features using a 3*3 convolution to obtain the fused multi-scale interaction feature map.

[0099] The global spatial attention mechanism is specifically described as follows:

[0100] First, pass the input feature map through a dilated convolution layer with a convolution kernel size of 3*3 and a dilation rate the same as that of the dilated convolution used to extract multi-scale features inside the branch, to obtain a feature map with the number of channels reduced to 1 / 4 of the original. Then, pass it through a dilated convolution layer with a convolution kernel size of 3*3 and a dilation rate the same as that of the dilated convolution used to extract multi-scale features inside the branch, to obtain a feature map with the same number of channels as the original input feature map. Then, perform a Sigmoid function transformation and perform an element-wise multiplication operation with the original input feature map to obtain the output feature map.

[0101] The expression of the Sigmoid function is: Where z is each element of the operation result, here referring to the feature map after 1*1 convolution processing, and f(z) is the result after Sigmoid transformation of each element.

[0102] The structure of each branch is specifically described as follows:

[0103] The first branch sequentially includes a convolution layer, a dilated convolution layer, a multi-scale interaction structure, and a global spatial attention mechanism. First, pass through a convolution layer with a convolution kernel size of 1*1 to obtain a feature map with the number of channels reduced to 1 / 4 of the original. Then, pass through a dilated convolution layer with a convolution kernel size of 3*3 and a dilation rate of 1 to obtain a scale feature map with a receptive field of 3*3. Then, send it into the multi-scale interaction structure to obtain the feature map after multi-scale interaction. Then, send it into the global spatial attention mechanism to obtain the feature map with new feature weights assigned.

[0104] The second branch sequentially includes a convolutional layer, a dilated convolutional layer, a multi-scale interaction structure, and a global spatial attention mechanism. First, it passes through a convolutional layer with a kernel size of 1*1 to obtain a feature map with the number of channels reduced to 1 / 4 of the original. Then, it passes through a dilated convolutional layer with a kernel size of 3*3 and a dilation rate of 2 to obtain a scale feature map with a receptive field of 7*7. Then, it is fed into the multi-scale interaction structure to obtain a feature map after multi-scale interaction. Then, it is fed into the global spatial attention mechanism to obtain a feature map with new feature weights assigned.

[0105] The third branch sequentially includes a convolutional layer, a dilated convolutional layer, a multi-scale interaction structure, and a global spatial attention mechanism. First, it passes through a convolutional layer with a kernel size of 1*1 to obtain a feature map with the number of channels reduced to 1 / 4 of the original. Then, it passes through a dilated convolutional layer with a kernel size of 3*3 and a dilation rate of 3 to obtain a scale feature map with a receptive field of 11*11. Then, it is fed into the multi-scale interaction structure to obtain a feature map after multi-scale interaction. Then, it is fed into the global spatial attention mechanism to obtain a feature map with new feature weights assigned.

[0106] The fourth branch sequentially includes a convolutional layer, a dilated convolutional layer, a multi-scale interaction structure, and a global spatial attention mechanism. First, it passes through a convolutional layer with a kernel size of 1*1 to obtain a feature map with the number of channels reduced to 1 / 4 of the original. Then, it passes through a dilated convolutional layer with a kernel size of 3*3 and a dilation rate of 4 to obtain a scale feature map with a receptive field of 15*15. Then, it is fed into the multi-scale interaction structure to obtain a feature map after multi-scale interaction. Then, it is fed into the global spatial attention mechanism to obtain a feature map with new feature weights assigned.

[0107] Step 400: Upsample the scale-aware information feature corresponding to the high-level semantic feature map using bilinear interpolation, and use a 1*1 convolution to adjust the channels. Then, perform an element-wise addition operation with the scale-aware information feature corresponding to the middle-level semantic feature map to obtain a fused feature.

[0108] Step 500: Upsample the fused feature using bilinear interpolation, and use a 1*1 convolution to adjust the channels. Then, perform an element-wise addition operation with the scale-aware information feature corresponding to the low-level semantic feature map to obtain the main feature map for crowd density estimation.

[0109] Step 600: Input the obtained main feature map for crowd density estimation into the backend network of the deep neural network model to obtain the crowd density estimation map and the crowd count result corresponding to the image to be estimated. The backend network includes two convolutional layers. The number of channels of the feature maps generated by each convolutional layer is 64 and 1 respectively from the input to the output direction. The convolutional size of both convolutional layers is 3*3. Input the main feature map for crowd density estimation into the backend network of the deep neural network model. After passing through two convolutional layers in sequence, the crowd density estimation map and the crowd count result are obtained.

[0110] Embodiment

[0111] The solution of the present invention will be further described below in conjunction with specific embodiments of the present invention.

[0112] Step 1: Establish a deep neural network model for crowd counting. The overall structure is as Figure 2 shown, including the following steps:

[0113] 1.1) Establish a front-end network module. Arbitrarily input an image sample x to obtain low-level, middle-level, and high-level semantic feature maps. This stage includes thirteen convolutional operations and four max-pooling operations, which are two convolutional layers, one pooling layer, two convolutional layers, one pooling layer, three convolutional layers, one pooling layer, three convolutional layers, one pooling layer, and three convolutional layers in sequence. The convolutional kernel size of the convolutional layers is 3*3, the stride of the pooling layer is 2, and the number of channels of the feature maps generated by the convolutional layers are 64, 64, 128, 128, 256, 256, 256, 512, 512, 512, 512, 512, and 512 respectively. The low-level semantic feature map is obtained after the third pooling operation, the middle-level semantic feature map is obtained after the fourth pooling operation, and the high-level semantic feature map is obtained after the thirteenth convolutional operation.

[0114] 1.2) Establish a semantic information fusion module:

[0115] 1.2.1) Establish a semantic interaction structure. Receive the low-level, middle-level, and high-level semantic feature maps in 1.1) as inputs. After semantic interaction, obtain the corresponding fused semantic feature maps. This stage can be divided into three processes. First, the input high-level semantic feature map is used as the corresponding fused semantic feature map of the high-level semantic feature map. Secondly, a channel connection operation is performed between the input middle-level semantic feature map and the corresponding fused semantic feature map of the high-level semantic feature map after the upsampling operation, and then passed through two convolutional layers with a convolutional kernel size of 3*3 to obtain the corresponding fused semantic feature map of the middle-level semantic feature map. Finally, a channel connection operation is performed between the input low-level semantic feature map and the corresponding fused semantic feature map of the middle-level semantic feature map after the upsampling operation, and then passed through two convolutional layers with a convolutional kernel size of 3*3 to obtain the corresponding fused semantic feature map of the low-level semantic feature map.

[0116] 1.2.2) Establish a multi-scale attention module, which respectively receives the three fused semantic feature maps in 1.2.1) as inputs, extracts scale information, and obtains the scale-aware information features corresponding to the semantic feature maps. This stage includes four branches with different receptive fields, a channel connection operation, a 3×3 convolution, and an element-wise multiplication operation. Each of the four branches with different receptive fields first reduces the number of channels of the feature map to 1 / 4 of the original through a 1×1 convolution, and then passes through dilated convolutional layers with dilation rates of 1, 2, 3, and 4 respectively, with a convolutional kernel size of 3×3. Then, the scale feature maps with receptive fields of 3×3, 7×7, 11×11, and 15×15 obtained are sent into the multi-scale interaction structure. In the multi-scale interaction structure, the scale feature map with a receptive field of 3×3 in the input is used as the multi-scale interaction feature map corresponding to the receptive field of 3×3; the scale feature map with a receptive field of 7×7 in the input is connected to the multi-scale interaction feature map corresponding to the receptive field of 3×3 in the channel dimension to obtain an intermediate feature. After fusing the intermediate feature using a 3×3 convolution, the multi-scale interaction feature map corresponding to the receptive field of 7×7 is obtained; the scale feature map with a receptive field of 11×11 in the input is connected to the multi-scale interaction feature map corresponding to the receptive field of 7×7 in the channel dimension to obtain an intermediate feature. After fusing the intermediate feature using a 3×3 convolution, the multi-scale interaction feature map corresponding to the receptive field of 11×11 is obtained; the scale feature map with a receptive field of 15×15 in the input is connected to the multi-scale interaction feature map corresponding to the receptive field of 11×11 in the channel dimension to obtain an intermediate feature. After fusing the intermediate feature using a 3×3 convolution, the multi-scale interaction feature map corresponding to the receptive field of 15×15 is obtained. The multi-scale interaction feature maps corresponding to the four different receptive fields are all processed through the global spatial attention mechanism.In the global spatial attention mechanism, the receptive field of the input is 3×3. The corresponding multi-scale interaction feature map is passed through a dilated convolutional layer with a convolutional kernel size of 3×3 and a dilation rate of 1 to reduce the number of channels to 1 / 4 of the original. Then, it is passed through another dilated convolutional layer with a convolutional kernel size of 3×3 and a dilation rate of 1 to adjust the number of channels to the same as the original feature map. Then, a Sigmoid function transformation is performed, and an element-wise multiplication operation is carried out with the original feature map. The corresponding multi-scale interaction feature map with a receptive field of 7×7 is passed through a dilated convolutional layer with a convolutional kernel size of 3×3 and a dilation rate of 2 to reduce the number of channels to 1 / 4 of the original. Then, it is passed through another dilated convolutional layer with a convolutional kernel size of 3×3 and a dilation rate of 2 to adjust the number of channels to the same as the original feature map. Then, a Sigmoid function transformation is performed, and an element-wise multiplication operation is carried out with the original feature map. The corresponding multi-scale interaction feature map with a receptive field of 11×11 is passed through a dilated convolutional layer with a convolutional kernel size of 3×3 and a dilation rate of 3 to reduce the number of channels to 1 / 4 of the original. Then, it is passed through another dilated convolutional layer with a convolutional kernel size of 3×3 and a dilation rate of 3 to adjust the number of channels to the same as the original feature map. Then, a Sigmoid function transformation is performed, and an element-wise multiplication operation is carried out with the original feature map. The corresponding multi-scale interaction feature map with a receptive field of 15×15 is passed through a dilated convolutional layer with a convolutional kernel size of 3×3 and a dilation rate of 4 to reduce the number of channels to 1 / 4 of the original. Then, it is passed through another dilated convolutional layer with a convolutional kernel size of 3×3 and a dilation rate of 4 to adjust the number of channels to the same as the original feature map. Then, a Sigmoid function transformation is performed, and an element-wise multiplication operation is carried out with the original feature map. The calculation formula for the Sigmoid transformation of each element is as follows:.

[0117]

[0118] The results of the 4 branches are concatenated in the channel dimension, and after feature fusion using a 3×3 convolution, they are multiplied by the feature information obtained from the global channel attention mechanism to obtain the scale-aware information features of the corresponding semantic feature map. In the global channel attention mechanism, the fused semantic feature map is received as the input, and dimension transposition and tiling operations are performed in the 3 dimensions of channel, height, and width. Then, a linear transformation is used to reduce the channel dimension of the feature map to 1 / 4 of the original, and a ReLU activation function is used for non-linear transformation. Then, another linear transformation is used to change the channel dimension to the same as the original feature map, obtaining the feature map processed by the multi-layer perceptron. Then, dimension transposition and reshaping operations are performed in the 3 dimensions of channel, height, and width, and then a Sigmoid function transformation is performed, and an element-wise multiplication operation is carried out with the original input feature map to obtain the output feature information.

[0119] 1.2.3) Establish a feature fusion module that receives the scale-aware information features corresponding to the three-layer semantic feature maps in 1.2.2) as inputs, performs feature fusion, and obtains the main feature map for crowd density estimation. This stage can be divided into two processes. First, perform an upsampling operation on the scale-aware information features corresponding to the high-level semantic feature map, and use a 1*1 convolution to adjust the number of channels to 256. Then, perform an element-wise addition operation with the scale-aware information features corresponding to the middle-level semantic feature map to obtain the fused features. Perform an upsampling operation on the fused features, and use a 1*1 convolution to adjust the number of channels to 128. Then, perform an element-wise addition operation with the scale-aware information features corresponding to the low-level semantic feature map to obtain the main feature map for crowd density estimation.

[0120] 1.3) Establish a backend network module that receives the main feature map for crowd density estimation in 1.2.3) as an input, and obtains the crowd density estimation map and the crowd counting result corresponding to the input image sample x. This stage includes two convolution operations, and the sizes of the convolution kernels are both 3*3. The numbers of channels of the feature maps generated by the convolutional layers are 64 and 1 respectively, so as to obtain the crowd density estimation map and the crowd counting result. Use the Euclidean distance between the crowd density estimation map and the ground truth density map as the loss function. Calculate the absolute difference between the crowd density estimation map and the ground truth density map of each single image, and average the sum of the absolute differences of all images to obtain the result of the loss function. The calculation formula is as follows:

[0121]

[0122] where θ represents the parameters of the network model, N represents the number of training samples, X i represents the original image input to the network, G(X i ; θ) represents the estimated density map obtained after the original image passes through the network model, represents the ground truth density map.

[0123] After obtaining the crowd density estimation map and the crowd counting result corresponding to the image to be estimated each time, determine the error of the deep neural network model according to the loss function of the deep neural network model, backpropagate the error, adjust the parameters of the deep neural network model, and optimize the deep neural network model. The parameter Θ to be learned is updated using the Adam optimization algorithm in each optimization iteration until the weighted sum result L(Θ) converges to a smaller value, and the parameters and the trained model are saved.

[0124] Use the trained deep neural network model to perform crowd counting on any input image. Directly input any single image into the trained model, and the corresponding crowd density estimation map and crowd counting result can be obtained, as Figure 3 shown. Figure 3Schematic diagram of performing crowd counting using the crowd counting method based on feature interaction of the present invention.

[0125] In summary, the present invention extracts rich multi-scale information through a multi-scale attention module, improves the ability to extract multi-scale information and the sensitivity to valuable information, and overcomes the scale limitation and feature similarity problems in traditional multi-column networks; through a semantic information fusion module, the semantic information at different levels of the backbone network is interacted and fused to provide more abundant detailed features, enhance the feature aggregation ability of the network, and improve the utilization efficiency of the backbone network.

Claims

1. A crowd counting method based on feature interaction, characterized in that: Including the following steps: Input the original image into a deep neural network model for feature extraction to obtain three-layer semantic feature maps, where the three-layer semantic feature maps include a low-level semantic feature map, a middle-level semantic feature map, and a high-level semantic feature map, including the following steps: Input the original image into the deep neural network model. The deep neural network model includes two consecutive convolutional layers, a pooling layer, two convolutional layers, a pooling layer, three convolutional layers, and a pooling layer connected in sequence to obtain a low-level semantic feature map; the number of channels of the feature maps generated by each convolutional layer is 64, 64, 128, 128, 256, 256, and 256 in sequence from the input to the output direction; the convolutional kernel size of the convolutional layer is 3*3; the stride of the pooling layer is 2; Continue to input the low-level semantic feature map into the deep neural network model, and pass through three convolutional layers and a pooling layer in sequence to obtain a middle-level semantic feature map; the number of channels of the feature maps generated by each convolutional layer is 512; the convolutional kernel size of the convolutional layer is 3*3; the stride of the pooling layer is 2; Continue to input the middle-level semantic feature map into the deep neural network model, and pass through three convolutional layers to obtain a high-level semantic feature map; the number of channels of the feature maps generated by each convolutional layer is 512; the convolutional kernel size of the convolutional layer is 3*3; Send the three-layer semantic feature maps into a semantic interaction structure to obtain the corresponding three-layer fused semantic feature maps; The semantic interaction structure includes: Upsample the high-level semantic feature map using bilinear interpolation; Connect the result of the upsampled high-level semantic feature map with the middle-level semantic feature map in channels to obtain the intermediate feature corresponding to the middle-level semantic feature map; Perform feature fusion on the intermediate feature through two 3*3 convolutions to obtain the fused semantic feature map of the middle-level semantic feature map; Upsample the obtained fused semantic feature map of the middle-level semantic feature map using bilinear interpolation; Connect the result of the upsampled middle-level semantic feature map with the low-level semantic feature map in channels to obtain the intermediate feature corresponding to the low-level semantic feature map; Perform feature fusion on the intermediate feature through two 3*3 convolutions to obtain the fused semantic feature map of the low-level semantic feature map; Input the fused three-layer semantic feature maps into a multi-scale attention module respectively to obtain the scale-aware information features of the corresponding semantic feature maps; Upsample and adjust the channels of the scale-aware information features corresponding to the high-level semantic feature map and fuse them with the scale-aware information features corresponding to the middle-level semantic feature map; Upsample and adjust the fused features and fuse them with the scale-aware information features corresponding to the low-level semantic feature map to obtain the main feature map for crowd density estimation; Input the main feature map for crowd density estimation into the backend network of the deep neural network model to obtain a crowd density estimation map and a crowd counting result.

2. The method for crowd counting based on feature interaction according to claim 1, wherein: Sending the three-layer semantic feature maps into the semantic interaction structure includes the following steps: Send the high-level semantic feature map into the semantic interaction structure to obtain the fused semantic feature map corresponding to the high-level semantic feature map; Feed the middle - layer semantic feature map into the semantic interaction structure, interact with the semantic feature map after corresponding fusion with the high - layer semantic feature map, and obtain the semantic feature map after corresponding fusion of the middle - layer semantic feature map; Feed the low - layer semantic feature map into the semantic interaction structure, interact with the semantic feature map after corresponding fusion with the middle - layer semantic feature map, and obtain the semantic feature map after corresponding fusion of the low - layer semantic feature map.

3. A crowd counting method based on feature interaction according to claim 1, characterized in that: The multi - scale attention module includes 4 branches with different receptive fields, an operation of concatenating the results of the 4 branches in the channel dimension, a convolutional layer, an additional global channel attention mechanism, and an element - wise multiplication operation; each branch sequentially includes a convolutional layer, a dilated convolutional layer, a multi - scale interaction structure, and a global spatial attention mechanism; The global channel attention mechanism includes: Perform dimension transposition and tiling operations on the input feature map in the 3 dimensions of channel, height, and width to obtain the feature map after dimension transposition and tiling operations; Use linear transformation to reduce the channel dimension of the feature map after dimension transposition and tiling operations to 1 / 4 of the original, perform non - linear transformation using the ReLU activation function, and then use linear transformation to change the channel dimension to be the same as the original feature map to amplify the dependence of cross - dimensional features on the channel dimension, and obtain the feature map after being processed by a multi - layer perceptron; Perform dimension transposition and reshaping operations on the feature map after being processed by the multi - layer perceptron in the 3 dimensions of channel, height, and width to obtain the feature map after dimension transposition and reshaping operations; Perform Sigmoid function transformation on the feature map after dimension transposition and reshaping operations, and perform element - wise multiplication operation with the original input feature map to obtain the output feature map; The multi - scale interaction structure includes: Connect the result after interacting the small - receptive - field feature map with the large - receptive - field feature map in the channel dimension to obtain an intermediate feature; After fusing the intermediate feature using a 3×3 convolution, obtain the fused multi - scale interaction feature map; The global spatial attention mechanism includes: Pass the input feature map through a dilated convolutional layer with a convolutional kernel size of 3×3 and a dilation rate the same as that of the dilated convolution used to extract multi - scale features inside the branch, and obtain a feature map with the number of channels reduced to 1 / 4 of the original; Pass the feature map with the number of channels reduced to 1 / 4 of the original through a dilated convolutional layer with a convolutional kernel size of 3×3 and a dilation rate the same as that of the dilated convolution used to extract multi - scale features inside the branch, and obtain a feature map with the number of channels the same as the original input feature map; Perform Sigmoid function transformation on the feature map with the number of channels the same as the original input feature map, and perform element - wise multiplication operation with the original input feature map to obtain the output feature map; The branches include the first branch, the second branch, the third branch, and the fourth branch; The feature map passing through the first branch includes: Pass the feature map through a convolutional layer with a convolutional kernel size of 1×1 to obtain a feature map with the number of channels reduced to 1 / 4 of the original; The feature map with the number of channels reduced to 1 / 4 of the original passes through a dilated convolutional layer with a convolutional kernel size of 3*3 and a dilation rate of 1, resulting in a scale feature map with a receptive field of 3*3; The scale feature map with a receptive field of 3*3 is fed into a multi-scale interaction structure to obtain a feature map after multi-scale interaction; The feature map after multi-scale interaction is fed into a global spatial attention mechanism to obtain a feature map with new feature weights assigned; The feature map passes through a second branch, including: The feature map passes through a convolutional layer with a convolutional kernel size of 1*1, resulting in a feature map with the number of channels reduced to 1 / 4 of the original; The feature map with the number of channels reduced to 1 / 4 of the original passes through a dilated convolutional layer with a convolutional kernel size of 3*3 and a dilation rate of 2, resulting in a scale feature map with a receptive field of 7*7; The scale feature map with a receptive field of 7*7 is fed into a multi-scale interaction structure to obtain a feature map after multi-scale interaction; The feature map after multi-scale interaction is fed into a global spatial attention mechanism to obtain a feature map with new feature weights assigned; The feature map passes through a third branch, including: The feature map passes through a convolutional layer with a convolutional kernel size of 1*1, resulting in a feature map with the number of channels reduced to 1 / 4 of the original; The feature map with the number of channels reduced to 1 / 4 of the original passes through a dilated convolutional layer with a convolutional kernel size of 3*3 and a dilation rate of 3, resulting in a scale feature map with a receptive field of 11*11; The scale feature map with a receptive field of 11*11 is fed into a multi-scale interaction structure to obtain a feature map after multi-scale interaction; The feature map after multi-scale interaction is fed into a global spatial attention mechanism to obtain a feature map with new feature weights assigned; The feature map passes through a fourth branch, including: The feature map passes through a convolutional layer with a convolutional kernel size of 1*1, resulting in a feature map with the number of channels reduced to 1 / 4 of the original; The feature map with the number of channels reduced to 1 / 4 of the original passes through a dilated convolutional layer with a convolutional kernel size of 3*3 and a dilation rate of 4, resulting in a scale feature map with a receptive field of 15*15; The scale feature map with a receptive field of 15*15 is fed into a multi-scale interaction structure to obtain a feature map after multi-scale interaction; The feature map after multi-scale interaction is fed into a global spatial attention mechanism to obtain a feature map with new feature weights assigned.

4. A crowd counting method based on feature interaction according to claim 3, wherein: The three fused semantic feature maps are respectively input into a multi-scale attention module to obtain scale-aware information features corresponding to the semantic feature maps, including the following steps: The fused low-level semantic feature map is input into four branches of the multi-scale attention module to obtain low-level semantic feature maps of four scales; The four low-level semantic feature maps of the four scales are connected in channels, and after feature fusion using a 3*3 convolution and multiplied by the feature information obtained from the global channel attention mechanism, the scale-aware information features corresponding to the low-level semantic feature maps are obtained; The fused middle-level semantic feature map is input into four branches of the multi-scale attention module to obtain middle-level semantic feature maps of four scales; Connect the middle - level semantic feature maps of the four scales in the channel dimension, perform feature fusion using a 3×3 convolution, and then multiply the result by the feature information obtained from the global channel attention mechanism to obtain the scale - aware information features corresponding to the middle - level semantic feature maps; Input the fused high - level semantic feature maps into the four branches of the multi - scale attention module to obtain high - level semantic feature maps of four scales; Connect the high - level semantic feature maps of the four scales in the channel dimension, perform feature fusion using a 3×3 convolution, and then multiply the result by the feature information obtained from the global channel attention mechanism to obtain the scale - aware information features corresponding to the high - level semantic feature maps.

5. A crowd counting method based on feature interaction according to claim 1, characterized in that: Upsample and adjust the channels of the scale - aware information features corresponding to the high - level semantic feature maps and fuse them with the scale - aware information features corresponding to the middle - level semantic feature maps, including the following steps: Upsample the scale - aware information features corresponding to the high - level semantic feature maps using bilinear interpolation and adjust the channels using a 1×1 convolution to obtain the first feature map; Perform an element - wise addition operation on the first feature map and the scale - aware information features corresponding to the middle - level semantic feature maps to obtain the fused features.

6. The population counting method based on feature interaction according to claim 1, wherein: Upsample and adjust the channels of the fused features and fuse them with the scale - aware information features corresponding to the low - level semantic feature maps to obtain the main feature map for crowd density estimation, including the following steps: Upsample the fused features using bilinear interpolation and adjust the channels using a 1×1 convolution to obtain the second feature map; Perform an element - wise addition operation on the second feature map and the scale - aware information features corresponding to the low - level semantic feature maps to obtain the main feature map for crowd density estimation.

7. A crowd counting method based on feature interaction according to claim 1, characterized in that: Input the main feature map for crowd density estimation into the backend network of the deep neural network model to obtain the crowd density estimation map and the crowd counting result, including the following steps: Input the main feature map for crowd density estimation into two convolutional layers to obtain the crowd density estimation map and the crowd counting result; the number of channels of the feature maps generated by each convolutional layer is 64 and 1 in sequence from the input to the output direction; the convolution kernel sizes of the two convolutional layers are both 3×3.

Citation Information

Patent Citations

  • A population counting method and system based on density estimation

    CN113538402B

  • Dense crowd counting method based on multi-scale feature pyramid network

    CN113011329A