Super real-time crowd counting method based on stem network-encoder-decoder architecture
By adopting the stem network-encoder-decoder architecture in the crowd counting method, the problem of insufficient efficiency and real-time in the prior art is solved, and ultra-fast inference and efficient counting are realized, which is suitable for hyper-real-time crowd counting tasks.
Patent Information
- Application Number
- CN202211486208.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-24
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-11-24
AI Technical Summary
The existing crowd counting methods have shortcomings in efficiency and real-time performance, and cannot meet the needs of hyper-real-time crowd counting tasks.
The hyper-real-time crowd counting method based on the stem network-encoder-decoder architecture is adopted to achieve fast inference and efficient counting through the combination of multiple convolution operations and feature pyramid models.
While ensuring a certain counting accuracy, ultra-fast inference is achieved, improving the real-time performance and network efficiency of the method in crowd scenarios.
Smart Images

Figure CN115797860B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of pattern recognition, and in particular relates to a super real-time crowd counting method. Background Art
[0002] In recent years, with the increase in the number of people, stampedes often occur. In order to prevent large-scale stampedes, crowd management and control are particularly important. Therefore, accurately estimating the number of people from videos or images has become a very important function in intelligent monitoring. This function is called crowd counting, which is of great significance in population control, public safety management and urban planning.
[0003] In general, most existing methods pay more attention to the accuracy of the model and ignore the efficiency of the model. For example, Zhang et al. proposed a multi-column convolutional neural network (MCNN) for crowd counting in "Y. Zhang, D. Zhou, S. Chen, S. Gao, and Y. Ma, 'Single-image crowd counting via multi-column convolutional neural network, 'in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 589-597." and Yi et al. designed a scale-adaptive crowd counting network (SACCN) in "Q. Yi, Y. Liu, A. Jiang, J. Li, K. Mei, and M. Wang, 'Scale-Aware Network with Regional and Semantic Attentions for Crowd Counting under Cluttered Background, 'CoRR, vol. abs / 2101.01479, 2021." Although these networks have high accuracy, their efficiency is not reliable, so they are not suitable for real-time monitoring tasks. The real-time nature of the model means that it needs to have efficient and fast reasoning capabilities on most devices. However, the limitations of the above methods (the need for larger memory and complex calculations caused by a large number of parameters) hinder the application of this process in intelligent monitoring and embedded devices.
[0004] Recently, researchers have focused on lightweight networks with fewer parameters and applied them to crowd counting tasks. For example, Gao et al. proposed a lightweight crowd counting network (PCC-Net-light) in "J.Gao, Q.Wang, and X.Li,'PCC-Net: Perspective Crowd Countingvia Spatial Convolutional Network,'IEEE Transactions on Circuits and Systemsfor Video Technology,vol.30,no.10,pp.3486-3498,2020." and Liu et al. proposed an efficient structured knowledge transfer crowd counting method (1 / 4-CSRNet+SKT) in "L.Liu, J.Chen, H.Wu, T.Chen, G.Li, and L.Lin,'Efficient Crowd Counting via Structured Knowledge Transfer,'in MM 20:The28th ACM International Conference on Multimedia,Virtual Event / Seattle,2020,pp.2645-2654." However, although the above methods reduce the number of model parameters, there is no significant improvement in inference speed. Therefore, they are still not suitable for ultra-real-time crowd counting tasks. Summary of the invention
[0005] In order to overcome the shortcomings of the prior art, the present invention provides a super real-time crowd counting method based on a stem network-encoder-decoder architecture, which is oriented to crowd scenes and achieves the goal of super real-time crowd counting through the stem network-encoder-decoder architecture. Due to the adoption of a new network architecture and the special algorithm design based on the characteristics of the crowd counting task, a better crowd counting effect can be achieved in the end, and the real-time performance of the method in the crowd scene is improved. The method of the present invention can achieve fast reasoning in crowd counting tasks. Compared with previous crowd counting methods, this method is faster and more effective in crowd counting tasks.
[0006] The technical solution adopted by the present invention to solve the technical problem includes the following steps:
[0007] Step 1: Input the original image into the stem network for processing;
[0008] Step 1-1: The size of the original input image is 3×W×H. The resolution of the input image is changed to 1 / 4 of the original image through multiple convolution operations to obtain preliminary features.
[0009] Step 1-2: Input the preliminary features into the channel disorder information mixing module;
[0010] The deep feature channel is divided into two branches, one of which performs 1×1, 3×3, and 1×1 convolution operations in sequence, and the other is used as a marker; then, the deep features of the two branches are concatenated along the channel dimension; finally, the channel random information mixing module is input to randomly shuffle the order of deep features to achieve information mixing communication;
[0011] Step 2: Load the mixed features into the encoder for processing;
[0012] Step 2-1: Form a multi-scale feature branch by gradually downsampling, with the lowest resolution being
[0013] Step 2-2: After two more stages, each stage repeats the module group twice, and each module includes two conditional channel weighting modules CCW in series and a multi-branch local fusion module MLF;
[0014] The MLF module fuses the multi-scale features output by the CCW module through downsampling and summation operations. Specifically, in the MLF module, it is divided into two parts: 1) Layers with the same resolution, the number of channels and scales are the same, and multi-scale features are used directly; 2) Layers with different resolutions, the number of channels and scales are different, and channels are converted and downsampled from high resolution to low resolution through a series of 3×3 convolutions with a step size of 2. The number of branches is 3, corresponding to a resolution of 1 / 2 of the original image. times, the corresponding number of channels is: 36, 64, 96;
[0015] Step 3: Use the feature pyramid model FPN as the decoder and combine multi-scale features to process feature maps of different sizes;
[0016] The decoding process is carried out by upsampling and lateral connection. First, the low-resolution feature map is upsampled twice. Then, the sampling result is fused with the feature map of the same size generated by 1×1 convolution to locate the crowd. Finally, 3×3 convolution is used to eliminate the aliasing effect of upsampling.
[0017] The decoding process is iterated repeatedly until the final density map is created;
[0018] After the FPN decoding operation, the three feature maps of different scales output by FPN are combined, and then feature regression is performed through two 1×1 convolution operations to make the number of channels 1, and the final prediction map is output to achieve ultra-real-time crowd counting;
[0019] Step 4: Define the loss function;
[0020] The input image is converted into a density map and the mean square error loss is used; the formula is as follows:
[0021]
[0022] Where Loss is the mean square error loss, N represents the number of input images, represents the prediction density map, represents the ground truth map, j represents the jth input image, represents the Euclidean metric.
[0023] Preferably, the convolution operation in step 1-1 uses a serial stack of large-core convolution layers with convolution kernel sizes of 9×9, 7×7, and 5×5 to expand the receptive field during feature extraction and use downsampling to reduce the image resolution to
[0024] Preferably, the channel disorder information mixing module Shuffle Block is proposed by Ma et al. in the document “N.Ma, X.Zhang, H.Zheng, and J.Sun,'ShuffleNetV2:Practical Guidelines for EfficientCnn Architecture Design,'2018.”.
[0025] Preferably, the conditional channel weighting module CCW is proposed by Yu et al. in the document “C. Yu, B. Xiao, C. Gao, L. Yuan, L. Zhang, N. Sang, and J. Wang, 'Lite-HRNet: A Lightweight High-Resolution Network, 'in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2021, pp. 10440–10450.”;
[0026] Modifications are made based on the channel disorder information mixing module: element-by-element weighting operations are used to replace the 1×1 convolution in the channel disorder information mixing module to reduce complexity; specifically, the cross-resolution weight calculation function and the spatial weight calculation function are used to replace the first and second 1×1 convolutions for channel weighting. And there are n branches in the nth stage, and in the nth branch, the element weighting operation is as follows:
[0027] Y n =W n ⊙X n (2)
[0028] Where W n is the weight graph, X n is the input image, Y n is the output image and ⊙ is the element-wise multiplication operator.
[0029] Preferably, the FPN is proposed by Lin et al. in the document “T.Lin, P.Dollár, B.Ross, K.He, B.Hariharan, and J.Serge, 'Feature Pyramid Networks for Object Detection,' in Proc.IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp.936–944.”
[0030] The beneficial effects of the present invention are as follows:
[0031] The present invention can achieve ultra-fast reasoning while achieving a certain counting accuracy, and its ultra-real-time performance and network efficiency are superior to the current most advanced methods in crowd counting tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 The figure is a flow chart of the method of the present invention. DETAILED DESCRIPTION
[0033] The present invention is further described below in conjunction with the accompanying drawings and embodiments.
[0034] The present invention proposes a super real-time network method, which can quickly obtain the number of people from videos or images for low-power devices for crowd counting tasks. The network framework proposed by the present invention includes three main modules: stem network, encoder and decoder.
[0035] STEM Network:
[0036] 1. After inputting the initial image, it is early downsampled to reduce the size of the feature map, and the resolution is changed to 1 / 4 of the original image through a series of convolution operations.
[0037] 2. By using large kernel convolutions of sizes 9×9, 7×7, and 5×5, the receptive field is expanded to extract more detailed head features.
[0038] 3. Apply the extracted features to the channel shuffle information mixing module (Shuffle Block), that is, enhance the robustness of feature learning by shuffling the feature channels to achieve a mixed information effect.
[0039] Encoder:
[0040] After the stem network, the extracted features are divided into multi-scale feature branches through a step-by-step downsampling operation. It includes two stages and a series of modules, each of which has two conditional channel weighting (CCW) blocks and a multi-branch local fusion (MLF) block.
[0041] Decoder:
[0042] The Feature Pyramid Network (FPN) is used as the decoder to alleviate the problem of incomplete fusion of multi-branch features in the encoder, and then the final prediction map is output through two feature regression layers.
[0043] The ultra-real-time crowd counting method based on the "stem network-encoder-decoder" architecture is implemented in the following steps:
[0044] 1. In the stem network, the input original image size is 3×W×H. Use large kernel convolution and early downsampling operations to obtain preliminary features. For example, use 9, 7, 5 large kernel convolution to expand the receptive field and use downsampling to reduce the image resolution to
[0045] 2. Input the preliminary features into the channel shuffle information mixing module (Shuffle Block). The specific operation is as follows: First, divide the channel into two branches, one of which performs 1×1, 3×3, and 1×1 convolution operations in sequence, and the other is used as a marker. After that, the two branches are connected in series. Finally, the channel is shuffled to achieve information mixing communication. Among them, Shuffle Block was proposed by Ma et al. in the document "N.Ma,X.Zhang,H.Zheng,and J.Sun,'ShuffleNetV2:PracticalGuidelines for Efficient Cnn Architecture Design,'2018."
[0046] 3. Load the mixed features into the encoder. In order to reduce the number of parameters and floating points, this part is designed in two stages. At the same time, multi-scale feature branches are formed through step-by-step downsampling operations, with the lowest resolution of The two stages of the design contain a series of modules, each with two conditional channel weighting (CCW) blocks and a multi-branch local fusion (MLF) block, and the modules are repeated twice in each stage. The conditional channel weighting (CCW) block was proposed by Yu et al. in the paper "C.Yu, B.Xiao, C.Gao, L.Yuan, L.Zhang, N.Sang, and J.Wang, 'Lite-HRNet: A Lightweight High-Resolution Network, 'in Proc.IEEE Conference on Computer Vision and Pattern Recognition, 2021, pp.10440–10450." The CCW block is modified based on the channel out-of-order information mixing module: the 1×1 convolution in the channel out-of-order information mixing module is replaced by an element-by-element weighted operation to reduce complexity. Specifically, the first and second 1×1 convolutions are replaced by the cross-resolution weight calculation function and the spatial weight calculation function for channel weighting respectively. And there are n branches in the nth stage. In the nth branch, the element weighting operation is as follows:
[0047] Y n =W n ⊙X n
[0048] Where W n is the weight graph, X n is the input image, Y n is the output image and ⊙ is the element-wise multiplication operator.
[0049] The multi-branch local fusion (MLF) block fuses the multi-scale features output by the CCW block through downsampling and summation operations. Specifically, in the MLF block, it is divided into two parts: 1) Layers with the same resolution, the number of channels and scale are the same, and the features are used directly; 2) Layers with different resolutions, the number of channels and scales are different, and channel conversion and downsampling are performed from high resolution to low resolution (completed by serially connecting 3×3 convolutions with a step size of 2). The number of branches is 3, and the corresponding resolution is: The corresponding channel numbers are: 36, 64, 96.
[0050] 4. In this work, the encoder does not perform cross-fusion of feature maps, and the low-resolution feature map information is not fused with the high-resolution feature map information. Therefore, in order to alleviate the problem of incomplete information integration caused by local fusion in the encoder, the feature pyramid model (FPN) is used as the decoder to combine multi-scale features to process feature maps of different sizes. The decoding process is carried out in the form of upsampling and lateral connection. Specifically: first, the low-resolution feature map is upsampled twice; then the sampling result is fused with the feature map of the same size generated by 1×1 convolution to locate the details; finally, a 3×3 convolution is used to eliminate the aliasing effect of upsampling. At the same time, this process is iterative until the final density map is created. Among them, FPN was proposed by Lin et al. in the document "T.Lin, P.Dollár, B.Ross, K.He, B.Hariharan, and J.Serge,'Feature Pyramid Networks for Object Detection,'in Proc.IEEE Conference onComputer Vision and Pattern Recognition, 2017, pp.936–944."
[0051] Furthermore, after the FPN decoding operation, the three output feature maps are combined, and then feature regression is performed through two 1×1 convolution operations to make the number of channels 1, and the final prediction map is output.
[0052] 5. Definition of loss function. Convert the input image into a density map and use mean square error loss. The formula is as follows:
[0053]
[0054] Where Loss is the mean square error loss, N represents the number of input images, represents the prediction density map, represents the ground truth map, j represents the jth input image, represents the Euclidean metric. Specific embodiment:
[0056] The effects of the present invention can be further illustrated by the following experimental results.
[0057] 1. Experimental environment and settings
[0058] The present invention is to use i7-6900K @ 3.4GHz, 64GB RAM, 2 NVIDIA GTX1080Ti GPUs, running on Ubuntu 16.04.
[0059] In the experiment, three indicators are used to evaluate the performance of the model, namely, mean absolute error (MAE), mean square error (MSE), and mean normalized absolute error (NAE). The definitions are as follows:
[0060]
[0061]
[0062]
[0063] Where N is the number of images in the three datasets. is the actual quantity, is the predicted number of the i-th test image. In addition, MAE is an indicator for evaluating the accuracy of the calculated crowd number, while MSE shows the robustness of the estimated crowd number. NAE is an indicator for evaluating the impact of eliminating negative samples (avoiding zero denominator).
[0064] Three datasets are used in the experiment: UCF-QNRF, NWPU-Crowd, and ShanghaiTech.
[0065] The UCF-QNRF dataset was proposed by Idrees et al. in the paper "H. Idrees, M. Tayyab, K. Athrey, D. Zhang, S. Al-Máadeed, N. Rajpoot, and M. Shah, 'Composition Loss for Counting, Density Map Estimation and Localization in Dense Crowds,' in Proc. IEEE European Conference on Computer Vision, vol. 11206, 2018, pp. 544–559." It contains 1535 dense crowd images, of which 1201 images are used for training and 334 for testing. These images include the most diverse viewpoint settings, density and lighting changes, making this dataset more realistic and difficult. The NWPU-Crowd dataset was proposed by Wang et al. in the paper "Q. Wang, J. Gao, W. Lin, and X. Li, 'NWPU-Crowd: A Large-Scale Benchmark for Crowd Counting and Localization, 'IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 43, no. 6, pp. 2141–2149, 2021." It includes 5109 images, of which 3109, 500, and 2000 images are used for training, validation, and testing, respectively, with 2133375 annotated points and boxes. This is the latest and largest dataset, covering various lighting scenes with the largest density range (0 to 20033). The ShanghaiTech dataset was proposed by Zhang et al. in the paper "Y. Zhang, D. Zhou, S. Chen, S. Gao, and Y. Ma, 'Single-Image Crowd Counting via Multi-Column Convolutional Neural Network, 'in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 589–597." and consists of two parts: SHHA and SHHB. SHHA is 482 images randomly grabbed from the Internet. 182 of them are test sets and the rest are training sets. SHHB is 716 images taken from a busy commercial street. 316 of them are test sets and the rest are training sets.
[0066] 2. Experimental content
[0067] The method of the present invention is used to test the network on three data sets, and the error values on different data sets are calculated to measure the accuracy of the network. The network parameters, floating-point operations, inference time, and frame rate (FPS) are calculated to measure the performance of the network.
[0068] To demonstrate the effectiveness of the algorithm, Tables 1 and 2 list the performance comparisons of some network methods on the NWPU-Crowd and UCF-QNRF datasets and the ShanghaiTech (SHHA and SHHB) datasets, respectively. On the one hand, seven non-pre-training methods are compared on three datasets, namely the MCNN algorithm proposed in the paper "Y. Zhang, D. Zhou, S. Chen, S. Gao, and Y. Ma, 'Single-Image Crowd Counting via Multi-Column Convolutional Neural Network, 'in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 589–597."; the SANet algorithm described in detail in the paper "X. Cao, Z. Wang, Y. Zhao, and F. Su, 'Scale Aggregation Network for Accurate and Efficient Crowd Counting, 'in Proc. IEEE European Conference on Computer Vision, vol. 11209, 2018, pp. 757–773."; and the CNN-Based Cascaded Multi-Task Learningof High-Level Prior and Density Estimation for Crowd Counting,'in Proc.IEEE International Conference on Advanced Video and Signal Based Surveillance, 2017, pp.1–6. "The CMTL algorithm is described in detail in the literature "Z.Shen, Y.Xu, B.Ni, M.Wang, J.Hu, and X.Yang, 'Crowd Counting via Adversarial Cross-Scale Consistency Pursuit,' in Proc.IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp.5245–5254. "The ACSCP algorithm is described in detail in the literature "J.Gao, Q.Wang, and X.Li,'PCC-Net: Perspective Crowd Counting via Spatial Convolutional Network,'IEEE Transactions on Circuits and Systems for Video Technology,vol.30,no.10,pp.3486–3498,2020.' The PCC-Net-light algorithm is described in detail in the literature "DBSam andR.V.Babu,'Top-Down Feedback for Crowd Counting Convolutional Neural Network,'in Proc.AAAI Conference on Artificial Intelligence,2018,pp.7323–7330.' The TDF-CNN algorithm is described in detail in the literature "X.Shi,X.Li,C.Wu,S.Kong,J.Yang,and L.He,'A Real-Time Deep Network for Crowd Counting,'in IEEE International Conference on Acoustics, Speech and Signal Processing, 2020, pp.2328–2332.” On the other hand, compared with the pre-trained networks at the bottom of the list, namely the CSRNet algorithm proposed in the paper “Y.Li, X.Zhang, and D.Chen,'CSRNet:Dilated Convolutional Neural Networks for Understanding the HighlyCongested Scenes,'in Proc.IEEE Conference on Computer Vision and PatternRecognition, 2018, pp.1091–1100.”; in the paper “L.Liu, J.Chen, H.Wu, T.Chen, G.Li, and L.Lin,'Efficient Crowd Counting via Structured KnowledgeTransfer,'in MM'20:The 28. thACM International Conference on Multimedia, Virtual Event / Seattle, 2020, pp.2645–2654.” The 1 / 4-CSRNet+SKT algorithm proposed in the literature “J.Gao, W.Lin, B.Zhao, D.Wang, C.Gao, and J.Wen, 'C^3Framework: An Open-Source PyTorch Codefor Crowd Counting, 'CoRR, vol.abs / 1907.02724, 2019.” The C3F-VGG and SCAR algorithms proposed in the literature “Q.Wang, J.Gao, W.Lin, and YY uan, 'Learning from Synthetic Data for CrowdCounting in the Wild, 'in Proc.IEEE Conference on Computer Vision and Pattern Recognition, 2019, pp.8198–8207.” The comparison results are shown in Table 1 and Table 2:
[0069] Table 1
[0070]
[0071] Table 2
[0072]
[0073]
[0074] As can be seen from Table 1 and Table 2, in non-pre-trained networks, except for the SHHA dataset, the counting accuracy of the present invention is the best. Compared with the pre-trained network at the bottom of the list, the method of the present invention has better results and is still competitive. In addition, the network parameter volume of the present invention is only 0.15MB, with the lowest floating-point operation volume (only 1.32G) and the shortest reasoning time (only 2.6ms). According to the floating-point operation theory and reasoning time, the model of the present invention is faster than all existing methods and reaches a speed of 381.7FPS on GTX 1080Ti. In comparison, other networks with higher accuracy cannot reach this speed and the performance is not ideal.
[0075] The purpose of this invention is to achieve ultra-real-time crowd counting tasks, so the key goal is to achieve fast reasoning while ensuring model efficiency.
[0076] In order to verify the super real-time performance of the present invention, networks with different input sizes are compared on four hardware devices. Table 3 calculates the test codes of GTX 1080Ti, RTX 3090, NVIDIA TX1 and NVIDIA Xavier on the NWPU-Crowd test set. The resolution of the input image is 576×768 and 1024×1024. The comparison results are shown in Table 3:
[0077] Table 3
[0078]
[0079] As can be seen from Table 3, the present invention has super real-time performance. Through the test of GTX 1080Ti, RTX 3090, NVIDIATX1 and NVIDIA Xavier hardware devices, the network achieves super fast reasoning, and the running time does not exceed 15ms (resolution is 576×768). It can be seen that the network is suitable for ultra-real-time tasks while ensuring the calculation accuracy.
Claims
1. A super real-time crowd counting method based on a stem network-encoder-decoder architecture, characterized in that: The steps include: Step 1: Input the original image into the STEM network for processing; Step 1-1: The size of the original input image is 3×W×H. The resolution of the input image is changed to 1 / 4 of the original image through multiple convolution operations to obtain preliminary features. Step 1-2: Input the preliminary features into the channel disorder information mixing module; The deep feature channel is divided into two branches, one of which performs 1×1, 3×3, and 1×1 convolution operations in sequence, and the other is used as a marker; then, the deep features of the two branches are concatenated along the channel dimension; finally, the channel random information mixing module is input to randomly shuffle the order of deep features to achieve information mixing communication; Step 2: Load the mixed features into the encoder for processing; Step 2-1: Form a multi-scale feature branch by gradually downsampling, with the lowest resolution being Step 2-2: After two more stages, each stage repeats the module group twice, and each module includes two conditional channel weighting modules CCW in series and a multi-branch local fusion module MLF; The MLF module fuses the multi-scale features output by the CCW module through downsampling and summation operations; specifically, in the MLF module, it is divided into two parts: 1) Layers with the same resolution, the number of channels and scale size are the same, and multi-scale features are directly used; 2) Layers with different resolutions have different numbers of channels and scales. Channel conversion and downsampling are performed from high resolution to low resolution through a series of 3×3 convolutions with a step size of 2. The number of branches is 3, corresponding to a resolution of 1 / 2 of the original image. times, the corresponding number of channels is: 36, 64, 96; Step 3: Use the feature pyramid model FPN as the decoder and combine multi-scale features to process feature maps of different sizes; The decoding process is carried out by upsampling and lateral connection. First, the low-resolution feature map is upsampled twice. Then, the sampling result is fused with the feature map of the same size generated by 1×1 convolution to locate the crowd. Finally, 3×3 convolution is used to eliminate the aliasing effect of upsampling. The decoding process is iterated repeatedly until the final density map is created; After the FPN decoding operation, the three feature maps of different scales output by FPN are combined, and then feature regression is performed through two 1×1 convolution operations to make the number of channels 1, and the final prediction map is output to achieve ultra-real-time crowd counting; Step 4: Define the loss function; The input image is converted into a density map and the mean square error loss is used; the formula is as follows: Where Loss is the mean square error loss, N represents the number of input images, represents the predicted density map, represents the ground truth map, j represents the jth input image, represents the Euclidean metric.
2. The ultra-real-time crowd counting method based on the stem network-encoder-decoder architecture according to claim 1 is characterized in that: The convolution operation in step 1-1 uses a serial stack of large-core convolution layers with convolution kernel sizes of 9×9, 7×7, and 5×5 to expand the receptive field during feature extraction and use downsampling to reduce the image resolution to 3. The ultra-real-time crowd counting method based on the stem network-encoder-decoder architecture according to claim 1 is characterized in that: The channel disorder information mixing module Shuffle Block was proposed by Ma et al. in the document "N.Ma, X.Zhang, H.Zheng, and J.Sun, 'ShuffleNetV2: Practical Guidelines for EfficientCnn Architecture Design, '2018." 4. The ultra-real-time crowd counting method based on the stem network-encoder-decoder architecture according to claim 1 is characterized in that: The conditional channel weighting module CCW was proposed by Yu et al. in the document "C.Yu, B.Xiao, C.Gao, L.Yuan, L.Zhang, N.Sang, and J.Wang, 'Lite-HRNet: A Lightweight High-Resolution Network, 'in Proc.IEEE Conference on Computer Vision and Pattern Recognition, 2021, pp.10440–10450."; Modifications are made based on the channel disorder information mixing module: the 1×1 convolution in the channel disorder information mixing module is replaced by element-by-element weighted operation to reduce complexity; specifically, the first and second 1×1 convolutions are replaced by the cross-resolution weight calculation function and the spatial weight calculation function for channel weighting respectively; and there are n branches in the nth stage, and in the nth branch, the element weighted operation is as follows: Y n =W n ⊙X n (2) Where W n is the weight graph, X n is the input image, Y n is the output image and ⊙ is the element-wise multiplication operator.
5. The ultra-real-time crowd counting method based on the stem network-encoder-decoder architecture according to claim 1, characterized in that: The FPN was proposed by Lin et al. in the document "T.Lin, P.Dollár, B.Ross, K.He, B.Hariharan, and J.Serge, 'Feature Pyramid Networks for Object Detection,' in Proc.IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp.936–944."
Citation Information
Patent Citations
Crowd counting method based on coding-decoding structure multi-scale convolutional neural network
CN111242036A
Dense crowd counting method based on multi-scale feature pyramid network
CN113011329A