Multi-scale-based feature extraction crowd counting method and system
Through the VGG-16 backbone network combined with the multi-scale feature extraction method of SAFMN and LSGA modules, the problem of target scale difference in monitoring images is solved, the robustness and accuracy of the crowd counting algorithm are improved, and it is suitable for real-time counting in complex scenarios.
Patent Information
- Application Number
- CN202510729540.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-06-03
AI Technical Summary
The prior art affects the robustness and counting accuracy of the crowd counting algorithm when dealing with the problem of target scale difference in monitoring images.
Image features are extracted using VGG-16 backbone network, and multi-scale feature fusion and optimization are performed through SAFMN and LSGA modules. The spatial correlation is modeled using the Gaussian position matrix to generate an optimized feature map to predict the population density distribution.
It improves the feature extraction ability of multi-scale human head targets, enhances the generalization ability of the model in complex scenarios, is suitable for real-time counting requirements for high-density and multi-scale population scenarios, and significantly improves counting accuracy.
Smart Images

Figure CN120236251A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of crowd counting, and specifically, to a multi-scale feature extraction-based crowd counting method and system. Background Art
[0002] Crowd counting technology can perform millisecond-level density measurement on the personnel distribution in the monitored area. When the detection value breaks through the preset safety threshold, the system immediately triggers a multi-level early warning mechanism, forming a complete closed-loop from risk identification to emergency intervention, and significantly improving the safety prevention and control capabilities in public places. Compared with traditional manual monitoring methods, the crowd counting algorithm shows obvious superiority in terms of accuracy and real-time response. This algorithm greatly improves the efficiency of information processing, reduces the economic and time costs of early warning decisions, and provides a more solid guarantee for the personal safety in public places. In practical applications, this technology promotes the transformation of the security mode from passive monitoring to active defense, and provides multi-dimensional data reports to provide strong decision-making support for the security department. In the security management of large-scale events, this technology realizes accurate modeling of personnel flow by combining multiple data sources, providing important technical support for the security construction of smart cities.
[0003] As the core technology of the intelligent perception system, the crowd counting algorithm is crucial for the modernization of social governance. It can analyze the crowd density, flow, and aggregation situation in real time, providing decision-making basis for multiple fields. It plays an important role in public safety, public health, commercial operation, smart city construction, etc., and has become an important technical support for the refined management and digital transformation of cities. The crowd counting algorithm constructs a data fusion analysis system by collecting and analyzing the distribution data of the crowd in real time, providing accurate data decision-making basis for urban management departments. This algorithm not only significantly improves the emergency response efficiency of large public events, but also continuously expands the application boundaries through its unique knowledge transfer mechanism: supporting quantitative analysis of cell colonies in the biomedical field and realizing intelligent monitoring of the scale of livestock herds in precision agriculture. This technology is continuously empowering the digital transformation of industries and promoting the evolution of social governance towards a data-driven intelligent model.
[0004] Generally speaking, as an important research direction in the field of computer vision, crowd counting technology has attracted much attention due to its practical application value in fields such as public security, urban planning, and commercial decision-making. However, in actual application scenarios, significant target scale differences still have a significant impact on the algorithm robustness and counting accuracy. Therefore, in view of the above current situation, there is an urgent need to provide a multi-scale feature extraction-based crowd counting method and system to overcome the deficiencies in current practical applications. Summary of the Invention
[0005] The purpose of the present invention is to provide a multi-scale feature extraction-based crowd counting method and system, aiming to solve the problems in the above-mentioned background technology.
[0006] The present invention is implemented as follows. A multi-scale feature extraction-based crowd counting method includes the following steps: Step 1: Extract the primary features of the input image through the VGG-16 backbone network; Step 2: Input the primary features into the MSGM, where the MSGM includes SAFMN and LSGA; Step 3: Use the SAFMN to perform multi-scale feature fusion and modulation on the input primary features, and input the features output by the SAFMN into the LSGA. Model the spatial correlation through the Gaussian position matrix to generate an optimized feature map; Step 4: Predict the crowd density distribution based on the optimized feature map and calculate the number of people.
[0007] As a further solution of the present invention: The SAFMN includes FMM, and the FMM includes SAFM and CCM.
[0008] As a further solution of the present invention: The SAFM realizes multi-scale feature extraction through the following steps: Divide the input features into four groups in the channel dimension, and perform adaptive max-pooling with different downsampling rates respectively; Aggregate the multi-scale features through 1×1 convolution to generate an attention map and modulate the input features.
[0009] As a further solution of the present invention: The CCM realizes the expansion and compression in the channel dimension through 3×3 convolution and 1×1 convolution, and combines the GeLU activation function to enhance the non-linear expression ability.
[0010] As a further solution of the present invention: The LSGA realizes spatial correlation modeling through the following steps: Regard the input features as independent pixel labels, and generate a spatial weight matrix through a two-dimensional Gaussian kernel; Introduce the Gaussian position matrix into the attention mechanism to dynamically adjust the feature weights; introduce a spatial weight matrix based on the two-dimensional Gaussian function to quantify the spatial correlation between pixels. The pixels closer to the center have higher weights in feature fusion. The formula is: ; Among them, is the standard deviation, is the spatial position coordinate, , h and w respectively refer to the height and width of the image.
[0011] As a further solution of the present invention: The attention calculation process of the LSGA is as follows: ; Among them, Q represents the query matrix respectively, X is the original input matrix, d is the dimension of each head in the multi-head attention mechanism, B is the relative position deviation parameter, the meaning of T is the transpose operation, Attention() refers to the output result of the attention mechanism, and Softmax() represents the normalized exponential function operation; Using the Gaussian absolute position to replace the relative position deviation, the attention relationship in LSGA can be expressed as: ; Among them, G is the Gaussian position matrix.
[0012] As a further solution of the present invention: In the processing of the output result of the SAFMN, a hierarchical fusion strategy is adopted to integrate cross-scale feature information, and the specific steps are as follows: Element-wise add the original scale feature map and the high-level feature map with a large receptive field to generate a fused feature; Input the fused feature into a cascaded module composed of pointwise convolution and global average pooling in sequence to achieve channel dimension reduction and global context information extraction; Perform secondary fusion on the refined feature and the original feature, and calculate the spatial-channel joint attention weight matrix through the Sigmoid activation function; Dynamically calibrate the input feature using the attention weight matrix and output the optimized feature representation.
[0013] A multi-scale feature extraction crowd counting system uses the multi-scale feature extraction crowd counting method as described above. The system includes: Backbone network module: Adopt the first thirteen layers of VGG-16 to perform primary feature extraction on the input image and output the primary feature map; MSGM: Includes SAFMN and LSGA, which are used to perform multi-scale feature fusion and optimization on the primary feature map and output the optimized feature map; Density estimation module: Used to calculate the crowd density distribution in the monitoring area according to the optimized feature map, and then obtain the number of people.
[0014] As a further solution of the present invention: The SAFM includes an adaptive max pooling layer group, a feature concatenation layer, and a 1×1 convolutional layer, which are used to achieve the aggregation and attention modulation of multi-scale features; the CCM includes a 3×3 convolutional layer, a 1×1 convolutional layer, and a GeLU activation layer, which are used to achieve cross-channel feature interaction.
[0015] As a further solution of the present invention: the LSGA includes a layer normalization layer, a residual connection structure, a spatial relationship modeling layer based on a two-dimensional Gaussian function, and a multi-layer perceptron, which are used for spatially recalibrating the feature map and classifying and outputting.
[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: Through the cross-scale feature aggregation of the SAFMN module and the Gaussian attention mechanism of the LSGA module, the problem of target scale difference in surveillance images is effectively solved, and the feature extraction ability for multi-scale human head targets is improved; By feature fusion and attention mechanism, background noise is suppressed, and the generalization ability of the model in complex scenarios is enhanced, which is applicable to the real-time counting requirements of high-density and multi-scale crowd scenarios; The ablation experiment shows that, compared with only using the backbone network, after introducing the SAFMN module, the MAE on the ShanghaiTech-A dataset drops from 58.8 to 55.7, and the MSE drops from 102.1 to 94.3; after further adding the LSGA module to form MSGM, the MAE drops to 53.9 and the MSE drops to 87.6, and significant accuracy improvement is also shown on the ShanghaiTech-B dataset, which proves the effectiveness of the method of the present invention. Description of the Drawings
[0017] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required to be used in the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0018] Figure 1 It is a schematic diagram of the network structure of VGGNet in the present invention.
[0019] Figure 2 It is a schematic diagram of the structure of the improved spatial adaptive feature modulation module in the present invention.
[0020] Figure 3 It is a schematic diagram of the structures of the SAFM module and the CCM module in the present invention.
[0021] Figure 4 It is a schematic diagram of the multi-scale feature context information fusion structure in the present invention.
[0022] Figure 5 It is a schematic diagram of the LSA and LSGA structures in the present invention.
[0023] Figure 6This is a schematic diagram of the LSGA transformer module structure in the present invention. Specific embodiments
[0024] Next, the technical solutions of the present invention will be clearly and completely described in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0025] The following further explains and illustrates the present invention in conjunction with specific embodiments.
[0026] Please refer to Figures 1 - 6 , a multi-scale feature extraction crowd counting method provided by an embodiment of the present invention, the method includes the following steps: Step 1: Extract the primary features of the input image through the VGG-16 backbone network; Step 2: Input the primary features into a multi-scale Gaussian feature extraction module (Multi-scale Gaussian Feature Extraction Module, MSGM), and the MSGM includes a spatially-adaptive feature modulation module (Spatially-Adaptive Feature Modulation, SAFMN) and a lightweight self-Gaussian attention module (Light Self-Gaussian-Attention, LSGA); wherein the MSGM performs context association on the features extracted by the front-end network to capture multi-scale head target feature information; First, perform multi-scale representation learning on the primary features extracted by the backbone network through the improved SAFMN, and then introduce LSGA to implement feature space recalibration, effectively suppressing background noise interference through dynamic channel weighting, thereby enhancing the discriminative expression of target features. This cascaded structure realizes the progressive optimization from multi-scale feature fusion to adaptive feature screening, significantly improving the feature separability in complex scenarios.
[0027] Step 3: Use SAFMN to perform multi-scale feature fusion and modulation on the input primary features, and input the features output by SAFMN into LSGA, model the spatial correlation through a Gaussian position matrix, and generate an optimized feature map; Step 4: Predict the crowd density distribution based on the optimized feature map and calculate the number of people.
[0028] A multi-scale feature extraction crowd counting system that uses the above-mentioned multi-scale feature extraction crowd counting method, the system includes: Backbone network module: The first thirteen layers of VGG-16 are adopted to perform primary feature extraction on the input image and output the primary feature map; MSGM: It includes SAFMN and LSGA, which are used to perform multi-scale feature fusion and optimization on the primary feature map and output the optimized feature map; Density estimation module: It is used to calculate the crowd density distribution in the monitoring area based on the optimized feature map, and then obtain the number of people.
[0029] In the embodiment of the present invention, through the cross-scale feature aggregation of the SAFMN module and the Gaussian attention mechanism of the LSGA module, the problem of target scale difference in the monitoring image is effectively solved, and the feature extraction ability for multi-scale human head targets is improved; background noise is suppressed through feature fusion and attention mechanism, and the generalization ability of the model in complex scenarios is enhanced, which is suitable for the real-time counting requirements of high-density and multi-scale crowd scenarios.
[0030] In an embodiment of the present invention, please refer to Figures 1 - 6 , VGG-16 is selected as the backbone (Backbone Network, BBN) of the network. This backbone network is responsible for extracting feature information and sending these features into subsequent modules. This is aimed at further refining detailed features that are conducive to capturing multi-scale crowd information and enhancing the network's ability to count multi-scale crowds. VGGNet mainly consists of convolutional layers, max-pooling layers, fully connected layers, and soft-max layers. Its core contribution lies in replacing large-size convolutional kernels with multiple stacked 3×3 small convolutional kernels, reducing the number of parameters while increasing the network depth, and significantly improving the image classification performance using a deeper network structure. The network adopts a structure of continuous convolutional layers and 2×2 max-pooling layers, and is connected to a fully connected layer at the end. Although the fully connected layer has a large number of parameters and a high computational cost, VGGNet is still widely used as a benchmark model in computer vision tasks due to its simple structure and strong feature extraction ability.
[0031] To effectively solve the semantic difference and scale inconsistency problems existing in cross-level feature maps, in this chapter, based on SAFMN, a cross-scale feature aggregation mechanism is introduced to adaptively fuse the semantic features with large receptive fields extracted by the deep network and the high-resolution spatial features of the shallow layer. The overall process is as Figure 2 shown.
[0032] The SAFMN module combines Spatially - Adaptive Feature Modulation (SAFM) and Cross - Channel Mixing (CCM) into a Feature Mixing Module (FMM), and proposes a network architecture with SAFM modules and CCM modules as the basic building blocks. This network architecture is carefully designed. First, a 3×3 convolutional network is used to perform preliminary feature extraction on the input data to obtain shallow features F0, which contain the basic information of the input data and preliminary feature representations. Subsequently, the network uses stacked FMMs to perform in - depth processing and transformation on the shallow features F0, further refining and mining the effective information in the features to generate deep features F1 with richer semantic information. Finally, the network organically fuses the shallow features F0 and the deep features F1, making full use of their complementarity to jointly generate the final result for more accurate prediction or classification.
[0033] Among them: The SAFM module is the core part. It adaptively modulates features using local information, enabling the model to select the most suitable modulation method for each pixel position and enhancing the adaptability to different regions. The CCM module is responsible for interacting with features in different channels to further enhance the ability to restore image details.
[0034] The overall structure diagram of the SAFM network is as shown in Figure 3 the left - hand figure in []. First, the input features normalized in the channel dimension are divided into four groups of components, and then they are respectively fed into the adaptive max - pooling layers, where the down - sampling layers have different down - sampling rates, which are 0, 2, 4, and 8 respectively. Since we hope to select discriminative features to learn non - local feature interactions, an adaptive max - pooling operator is applied to the input features to collect information. Given the input feature X, this process can be expressed by the formula: ; where, corresponds to the channel splitting operation, is a 3×3 depth convolution, represents quickly up - sampling to the original resolution at a specific level through nearest interpolation of the features, represents pooling the input features to the size of, refers to the i - th sub - feature after channel splitting; represents the aggregated representation of X0; represents the aggregated representation of.
[0035] Then, these extracted short-term or long-term features are aggregated by concatenating them along the channel dimension and performing 1×1 convolution, which can be expressed as: ; where, represents the concatenation operation, is the 1×1 convolution.
[0036] After obtaining the aggregated representation , a non-linear mapping is performed through an activation function to estimate the attention map, and the input x is adaptively adjusted according to the estimated attention through element-wise product. This process is shown as follows: ; where, represents the GeLU activation function, is the element-wise product, represents X after adaptive adjustment.
[0037] The architecture of CCM is as shown in the right part of Figure 3 , which consists of a 3X3, 1X1 convolution and a GeLU activation layer. The 3X3 convolution first doubles the input feature channels, and then the 1X1 convolution reduces the feature channels back to the same as the input.
[0038] SAFM and CCM are integrated into a unified feature mixing module (FMM) to select representative features. Then FMM can be expressed as: ; where, is the LayerNorm layer, and X, Y, and Z are intermediate features.
[0039] After obtaining the output result of SAMFN, to effectively integrate cross-scale feature information and alleviate the problem of semantic and scale inconsistency, this paper adopts a hierarchical fusion strategy to co-process the original scale feature map F and the high-level feature map with a large receptive field. As shown in Figure 4 , the process includes the following steps: First, an element-wise addition operation is performed on F and to generate a fused feature; Second, the generated fused feature is sequentially input into a concatenated module composed of pointwise convolution and global average pooling to respectively achieve channel dimension reduction and global context information extraction; Subsequently, the refined feature is fused with for the second time, and then a spatial-channel joint attention weight matrix is calculated through the Sigmoid activation function; Finally, an attention mechanism is used to dynamically calibrate the input feature, and an optimized feature representation is output.
[0040] Finally, this paper introduces the Lightweight Self-Gaussian Attention (LSGA) module. This module combines CNN and Transformer to form a three-stage feature processing mechanism. In the feature extraction stage, a convolutional layer is used for local perception to capture shallow semantic features such as image edges and textures. Secondly, in LGSA, the original input matrix is reused to replace the traditional key-value matrix generation operation. Finally, a spatial relationship modeling layer based on the two-dimensional Gaussian function is introduced to effectively characterize the correlation between the central region and the neighborhood in the feature map. The overall introduction of this model is as follows: The original Light Self-Attention Mechanism (LSA) is shown in Figure 5 (a) of. In the multi-head attention mechanism, relative position bias parameters are introduced into each head, as shown in Figure 5 (b) of. Then LSA can be expressed as: ; Among them, Q represents the query matrix respectively, X is the original input matrix, d is the dimension of each head in the multi-head attention mechanism, B is the relative position bias parameter, the meaning of T is the transpose operation, Attention() refers to the output result of the attention mechanism, and Softmax() represents the normalization exponential function operation.
[0041] The present invention abandons the traditional way of generating image patch tokens and instead treats each pixel as an independent token to maintain the spatial correlation of HSI (hue, saturation, intensity) data. For any spatial neighborhood in the input image, the spectral features of the central pixel dominate, while the contribution of neighboring pixels decays exponentially with the spatial distance, which conforms to the typical spatial autocorrelation law of HSI data. This model effectively quantifies the spatial correlation strength between pixels through the smooth decay characteristics of the Gaussian distribution, making pixels closer to the center have higher feature fusion weights. For this reason, this paper constructs a spatial weight matrix based on the two-dimensional Gaussian kernel: ; Among them, is the standard deviation, is the spatial position coordinate, , h and w respectively refer to the height and width of the image.
[0042] Using the Gaussian absolute position to replace the relative position bias, the attention relationship in LSGA can be expressed as: ; Among them, G is the Gaussian position matrix.
[0043] The overall LSGA transformer module, as shown in Figure 6As shown. In this module, first, layer normalization is performed on the input feature map, and then the output data obtained through LSGA is connected with the original input in a residual connection. Next, at the end of the module, normalization operations, multi-layer perceptron processing, and residual connection are performed in sequence. After being processed by the LSGATransformer module twice, the feature map is sent to a linear layer for flattening, and finally, the classification output of the feature is achieved through this linear layer.
[0044] In an embodiment of the present invention, this solution mainly includes three models, namely: Model 1: The backbone network model, which consists of the first thirteen layers of VGG-16; Model 2: The Backbone+SAFMN model, which adds a spatial adaptive feature modulation module on the basis of Model 1; Model 3: The Backbone+MSGM model, which adds a lightweight self-Gaussian attention module on the basis of Model 3 to form a complete branch structure; The ablation experiment shows that compared with only using the backbone network, after introducing the SAFMN module, the MAE on the ShanghaiTech-A dataset drops from 58.8 to 55.7, and the MSE drops from 102.1 to 94.3; after further adding the LSGA module to form MSGM, the MAE drops to 53.9, and the MSE drops to 87.6, and there is also a significant improvement in accuracy on the ShanghaiTech-B dataset, which proves the effectiveness of the method of the present invention; Table 1 Results of ablation experiment
[0045] To evaluate the counting performance of the crowd counting model, the present invention adopts two evaluation metrics: Mean Absolute Error (MAE) and Mean Squared Error (MSE).
[0046] (1) Mean Absolute Error: This metric objectively reflects the overall generalization ability of the model in the pixel-level density estimation task by calculating the absolute value of the crowd quantity estimation error frame by frame and performing data-level averaging. The calculation process is as shown in the formula: ;
[0047] (2) Mean Squared Error: This metric assigns higher weights to outlier errors through a quadratic penalty mechanism and can sensitively reflect the estimation bias of the model in extremely dense regions. The calculation process is as shown in the formula: ; where n is the number of samples; is the actual number of people in the i-th image; is the number of people predicted by the model for the i-th image.
[0048] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multi-scale based feature extraction method for crowd counting, characterized in that, The method includes the following steps: Step 1: Extract the primary features of the input image through the VGG-16 backbone network; Step 2: Input the primary features into the MSGM, where the MSGM includes SAFMN and LSGA; Step 3: Use SAFMN to perform multi-scale feature fusion and modulation on the input primary features, and input the features output by SAFMN into LSGA to model the spatial correlation through the Gaussian position matrix, generating an optimized feature map; Step 4: Predict the crowd density distribution based on the optimized feature map and calculate the number of people.
2. The multi-scale feature extraction-based crowd counting method according to claim 1, wherein The SAFMN includes FMM, and the FMM includes SAFM and CCM.
3. The multi-scale feature extraction-based crowd counting method according to claim 2, wherein The SAFM realizes multi-scale feature extraction through the following steps: Divide the input features into four groups in the channel dimension and perform adaptive max-pooling with different downsampling rates respectively; Aggregate multi-scale features through 1×1 convolution, generate an attention map and modulate the input features.
4. The multi-scale based feature extraction crowd counting method according to claim 2, wherein The CCM realizes the expansion and compression in the channel dimension through 3×3 convolution and 1×1 convolution, and combines the GeLU activation function to enhance the non-linear expression ability.
5. The multi-scale feature extraction-based crowd counting method according to claim 1, wherein The LSGA realizes spatial correlation modeling through the following steps: Regard the input features as independent pixel labels and generate a spatial weight matrix through a two-dimensional Gaussian kernel; Introduce the Gaussian position matrix into the attention mechanism to dynamically adjust the feature weights; introduce the spatial weight matrix based on the two-dimensional Gaussian function to quantify the spatial correlation between pixels, and the pixels closer to the center have higher weights in feature fusion. The formula is: ; wherein, is the standard deviation, is the spatial position coordinate, , h and w respectively refer to the height and width of the image.
6. The multi-scale feature extraction-based crowd counting method according to claim 1, characterized in that The attention calculation process of the LSGA is: ; Among them, Q represents the query matrix respectively, X is the original input matrix, d is the dimension of each head in the multi-head attention mechanism, B is the relative position bias parameter, the meaning of T is the transpose operation, Attention() refers to the output result of the attention mechanism, and Softmax() represents the normalization exponential function operation; Replace the relative position bias with the Gaussian absolute position, then the attention relationship in LSGA can be expressed as: ; Among them, G is the Gaussian position matrix.
7. The multi-scale feature extraction-based crowd counting method according to claim 1, characterized in that In the processing of the output result of the SAFMN, a hierarchical fusion strategy is adopted to integrate cross-scale feature information. The specific steps are as follows: Add the original-scale feature map and the high-level feature map with a large receptive field element by element to generate a fused feature; Input the fused features into the cascade module composed of pointwise convolution and global average pooling in sequence to realize channel dimension reduction and global context information extraction; Perform secondary fusion on the refined features and the original features, and calculate the spatial-channel joint attention weight matrix through the Sigmoid activation function; Dynamically calibrate the input features using the attention weight matrix and output the optimized feature representation.
8. A multi-scale based feature extraction crowd counting system, characterized in that, Apply the multi-scale based feature extraction crowd counting method as described in claim 2. The system includes: Backbone network module: Adopt the first thirteen layers of VGG-16 to extract the primary features of the input image and output the primary feature map; MSGM: Includes SAFMN and LSGA, which are used to perform multi-scale feature fusion and optimization on the primary feature map and output the optimized feature map; Density Estimation Module: It is used to calculate the crowd density distribution in the monitoring area based on the optimized feature map, and then obtain the number of people.
9. The multi-scale feature extraction based crowd counting system according to claim 8, wherein The SAFM includes an adaptive max pooling layer group, a feature concatenation layer, and a 1×1 convolutional layer, which are used to achieve the aggregation and attention modulation of multi-scale features; The CCM contains a 3×3 convolutional layer, a 1×1 convolutional layer, and a GeLU activation layer, which are used to achieve cross-channel feature interaction.
10. The multi-scale feature extraction-based crowd counting system according to claim 8, wherein The LSGA includes a layer normalization layer, a residual connection structure, a spatial relationship modeling layer based on a two-dimensional Gaussian function, and a multi-layer perceptron, which are used to perform spatial recalibration and classification output on the feature map.
Citation Information
Patent Citations
Dense crowd counting method based on multi-scale feature pyramid network
CN113011329A
Crowd counting method based on multi-scale space guide perception aggregation network
CN114694102A
Crowd density estimation method based on spatial context learning network
CN116403152A
Method for image motion deblurring, apparatus, electronic device and medium therefor
US20240404025A1