Methods, apparatus, equipment and storage media for analyzing pedestrian density
By using a deep learning detection architecture and Gaussian kernel density estimation algorithm, combined with an abnormal activity recognition model, a heat map of crowd density is generated, which solves the problem of real-time perception and risk warning in traditional crowd monitoring methods, and realizes automated, semantic, and real-time perception and risk prevention and control of high-density crowd places.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN QIYANG SPECIAL EQUIP TECH ENG CO LTD
- Filing Date
- 2026-01-21
- Publication Date
- 2026-06-02
AI Technical Summary
Traditional crowd monitoring methods are insufficient for real-time perception and risk warning in high-density crowd areas, and cannot effectively identify abnormal crowd activities.
A deep learning detection architecture is adopted, combined with a Gaussian kernel density estimation algorithm and an abnormal activity recognition model. Through pedestrian detection and abnormal activity recognition, a heat map of pedestrian flow density is generated and displayed on the monitoring platform.
It enables automated, semantic, and real-time perception and risk control of crowd status in high-density locations, and can accurately identify abnormal types such as pushing and crowding.
Smart Images

Figure CN122135279A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of crowd density analysis, and more specifically, to a crowd density analysis method, apparatus, device, and storage medium. Background Technology
[0002] With the acceleration of urbanization and the increasing frequency of large-scale public events, the safety management of high-density crowds in places such as subway stations, squares, and stadiums faces severe challenges. Traditional crowd monitoring mainly relies on manual inspections or simple video playback, which makes it difficult to achieve real-time perception of crowd status and risk warning. Summary of the Invention
[0003] In view of the above problems, this application proposes a method, apparatus, equipment and storage medium for analyzing pedestrian density, which can solve the above problems.
[0004] In a first aspect, embodiments of this application provide a method for analyzing pedestrian density. The method includes: acquiring video images of a scene to be monitored; inputting the video images into a preset pedestrian detection model for inference to determine the pedestrian positions, number of pedestrians, pedestrian density distribution map, and confidence level in each frame; the preset pedestrian detection model is a single-stage or two-stage deep learning detection architecture; if the confidence level is less than a preset threshold, inputting the video images into a preset abnormal activity recognition model for recognition to determine the type of abnormal crowd activity; the preset abnormal activity recognition model includes a 3D residual network, a region proposal network, and a region graph convolutional network; generating a pedestrian density heatmap based on the pedestrian density distribution map using a Gaussian kernel density estimation algorithm; and visually displaying the pedestrian positions, number of pedestrians, type of abnormal crowd activity, and pedestrian density heatmap on a monitoring platform.
[0005] Secondly, embodiments of this application also provide a pedestrian density analysis device, which includes: an acquisition module for acquiring video images of a scene to be monitored; a first determination module for inputting the video images into a preset pedestrian detection model for inference, determining the pedestrian positions, number of pedestrians, pedestrian density distribution map, and confidence level in each frame of the image; the preset pedestrian detection model is a single-stage or two-stage deep learning detection architecture; a second determination module for inputting the video images into a preset abnormal activity recognition model for recognition if the confidence level is less than a preset threshold, determining the type of abnormal crowd activity; the preset abnormal activity recognition model includes a 3D residual network, a region proposal network, and a region graph convolutional network; a generation module for generating a pedestrian density heatmap based on the pedestrian density distribution map and using a Gaussian kernel density estimation algorithm; and a display module for visually displaying the pedestrian positions, number of pedestrians, type of abnormal crowd activity, and pedestrian density heatmap on a monitoring platform.
[0006] Thirdly, embodiments of this application also provide a crowd density analysis device, including a processor, a memory, and one or more application programs; the one or more application programs are stored in the memory and configured to be executed by the processor to implement the above-described crowd density analysis method.
[0007] Fourthly, embodiments of this application also provide a computer-readable storage medium storing program code, wherein the above-mentioned crowd density analysis method is executed when the program code is run by a processor.
[0008] The technical solution provided in this application includes the following method: acquiring video images of the scene to be monitored; inputting the video images into a preset pedestrian detection model for inference to determine the pedestrian positions, number of pedestrians, pedestrian density distribution map, and confidence level in each frame of the image; the preset pedestrian detection model is a single-stage or two-stage deep learning detection architecture; if the confidence level is less than a preset threshold, the video images are input into a preset abnormal activity recognition model for recognition to determine the type of abnormal crowd activity; the preset abnormal activity recognition model includes a 3D residual network, a region proposal network, and a region graph convolutional network; based on the pedestrian density distribution map, a Gaussian kernel density estimation algorithm is used to generate a pedestrian density heatmap; and visually displaying the pedestrian positions, number of pedestrians, type of abnormal crowd activity, and pedestrian density heatmap on a monitoring platform. Therefore, the preset pedestrian detection model first outputs the pedestrian location, number, density distribution and confidence level. Then, the abnormal activity recognition model that integrates 3D residual network, region proposal network and region graph convolutional network is used to accurately identify abnormal types such as pushing and crowding. The pedestrian location, number, abnormal type and heat map are displayed in a unified visual display on the monitoring platform, thereby realizing the automated, semantic and real-time perception and risk prevention of the status of crowds in high-density places. Attached Figure Description
[0009] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments and drawings obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0010] Figure 1 A flowchart illustrating a method for analyzing pedestrian density provided in an embodiment of this application is shown.
[0011] Figure 2 A schematic diagram of a crowd density analysis device provided in an embodiment of this application is shown.
[0012] Figure 3A schematic diagram of the structure of a crowd density analysis device provided in an embodiment of this application is shown.
[0013] Figure 4 This illustration shows a schematic diagram of the structure of a computer-readable storage medium provided in an embodiment of this application. Detailed Implementation
[0014] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the invention as detailed in the appended claims.
[0015] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0016] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0017] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0018] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0019] References to "one embodiment" or "some embodiments" in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized.
[0020] This application provides a method, apparatus, device, and storage medium for analyzing pedestrian density. The method includes: acquiring video images of the scene to be monitored; inputting the video images into a preset pedestrian detection model for inference to determine the pedestrian positions, number of pedestrians, pedestrian density distribution map, and confidence level in each frame of the image; the preset pedestrian detection model is a single-stage or two-stage deep learning detection architecture; if the confidence level is less than a preset threshold, inputting the video images into a preset abnormal activity recognition model for recognition to determine the type of abnormal crowd activity; the preset abnormal activity recognition model includes a 3D residual network, a region proposal network, and a region graph convolutional network; generating a pedestrian density heatmap based on the pedestrian density distribution map using a Gaussian kernel density estimation algorithm; and visually displaying the pedestrian positions, number of pedestrians, type of abnormal crowd activity, and pedestrian density heatmap on a monitoring platform.
[0021] Therefore, the preset pedestrian detection model first outputs the pedestrian location, number, density distribution and confidence level. Then, the abnormal activity recognition model that integrates 3D residual network, region proposal network and region graph convolutional network is used to accurately identify abnormal types such as pushing and crowding. The pedestrian location, number, abnormal type and heat map are displayed in a unified visual display on the monitoring platform, thereby realizing the automated, semantic and real-time perception and risk prevention of the status of crowds in high-density places.
[0022] This invention provides a method for analyzing pedestrian density. The execution entity of this method includes, but is not limited to, at least one of the following pedestrian density analysis devices that can be configured to execute the method provided in this application: a server, a terminal, etc. In other words, this pedestrian density analysis method can be executed by software or hardware installed on a terminal device or a server device. The software can be a blockchain platform. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cluster of cloud servers. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0023] Please see Figure 1 , Figure 1 This application provides a schematic flowchart of a method for analyzing pedestrian density, as illustrated in an embodiment. Figure 1 As shown, the method may include steps 110 to 150.
[0024] In step 110, video images of the scene to be monitored are acquired.
[0025] Deploy image acquisition modules in the monitored scene (e.g., densely populated areas such as subway stations, squares, and shopping mall entrances and exits) to acquire continuous video streams (i.e., video images) of the monitored scene in real time or offline.
[0026] In one specific implementation, the image acquisition module includes, but is not limited to, a fixed camera, a panoramic camera, or an intelligent monitoring terminal, whose field of view covers the area of crowd activity to be analyzed.
[0027] In some implementations, the step "acquiring video images of the scene to be monitored" may include the following steps: (1) Acquire raw video frames at a preset frame rate using an image acquisition module deployed in the scene to be monitored; (2) Perform Gaussian filtering on the original video frames to obtain the denoised video frames; (3) Enhance the contrast of the denoised video frames to obtain the enhanced video frames; (4) Perform foreground target segmentation on the enhanced video frames and extract the regions of interest containing pedestrians; (5) Based on the face detection algorithm, locate the face region within the region of interest, and perform blur processing on the face region to obtain the video image.
[0028] In one specific implementation, the image acquisition module acquires raw video frames of the scene to be monitored at a preset frame rate of 3 frames per second to ensure that dynamic pedestrian flow in the scene to be monitored is captured.
[0029] To improve the robustness of subsequent abnormal crowd activity identification, the system performs Gaussian filtering on the original video frames before extracting keyframes to obtain denoised video frames. Specifically, in actual monitoring scenarios, there are often interference factors such as changes in lighting, sensor noise, or compression artifacts. These noises may be misjudged as significant edges or textures by the Region Proposal Network (RPN) in subsequent steps, leading to the generation of redundant or jittery region proposals. At the same time, even slight jitter at the center point of a region can be amplified in the calculation of displacement distance, thus affecting the accuracy of crowd movement intensity.
[0030] Therefore, in some implementations, a two-dimensional Gaussian kernel is used to perform smoothing filtering on each frame of the original video image. The standard deviation of the Gaussian kernel is... The window size can be preset according to the actual video resolution and noise level (e.g., Window size is This operation effectively suppresses high-frequency noise while preserving key low-frequency information such as crowd outlines and overall motion structure.
[0031] To further improve the robustness of crowd region detection under complex lighting conditions (e.g., backlighting, dim lighting, strong shadows, etc.), after Gaussian filtering for denoising, the system performs contrast enhancement processing on the denoised video frames to obtain enhanced video frames.
[0032] Specifically, in actual monitoring scenarios, due to uneven lighting or limitations in the dynamic range of camera equipment, crowd areas in the original video frames often exhibit localized underexposure or overexposure, resulting in blurred human contours and loss of texture details. This directly affects the ability of the region proposal network in subsequent steps to perceive densely populated areas, leading to missed detections or inaccurate region boundaries. Furthermore, the depth and spatiotemporal features extracted by the 3D residual network in subsequent steps may lack discriminative power due to low contrast, thus affecting the accuracy of identifying abnormal crowd activity types. Therefore, in some implementations, contrast enhancement methods such as adaptive histogram equalization or gamma correction are used to process the denoised video frames.
[0033] To further focus on the area of crowd activity and reduce interference from irrelevant backgrounds, coarse segmentation of foreground targets was performed on the enhanced video frames to initially extract regions of interest containing potential pedestrians.
[0034] In some implementations, foreground object segmentation can be achieved using lightweight methods. For example, inter-frame differencing based on adaptive thresholds, simple background models (e.g., Gaussian mixture models, GMMs), or single-frame saliency detection are not intended to precisely segment each pedestrian silhouette, but rather to coarsely filter out static backgrounds or regions far from the crowd, thereby narrowing the search scope of the Region Proposal Network (RPN) in subsequent steps.
[0035] Based on lightweight face detection algorithms (such as MTCNN, YOLO-Face, or OpenCV's built-in Haar cascade classifier), all face regions are located within the region of interest; subsequently, Gaussian blur, pixelation, or occlusion processing is applied to each detected face region to eliminate identifiable biometric information.
[0036] To ensure the security and compliance of personal privacy data, a face detection algorithm is used to scan video frames within the region of interest. It is important to clarify that this face detection algorithm is only used to locate areas in the image that contain human-like facial features; its goal is to identify potential facial contours, not to perform individual identification or verification.
[0037] Once a face region is located using the face detection algorithm, the system will immediately perform blurring processing on the face region. Preferably, techniques such as Gaussian blur can be used to irreversibly obfuscate the pixel information of the face region, making it impossible to identify specific personal characteristics.
[0038] Therefore, it can be seen that through a series of synergistically optimized preprocessing operations such as image acquisition, Gaussian filtering for noise reduction, contrast enhancement, foreground target segmentation, and directional blurring of the face region, the image quality and pedestrian feature recognition are significantly improved. At the same time, noise interference is effectively suppressed, target details in low-contrast scenes are enhanced, and the region of interest containing pedestrians is accurately focused. Furthermore, the face region is irreversibly blurred without destroying the overall structure and spatial distribution information of the human body, thus achieving an organic unity between privacy protection and detection performance. The resulting video images have high robustness, high fidelity, and high compliance.
[0039] In step 120, the video images are input into a preset pedestrian detection model for inference to determine the pedestrian positions, number of pedestrians, pedestrian density distribution map, and confidence level in each frame of the image.
[0040] In the embodiments of this application, the preset pedestrian detection model is a standard object detection model that has been trained and its parameters fixed before deployment. The preset pedestrian detection model is a single-stage or two-stage deep learning detection architecture. For example, the preset pedestrian detection model can be either YOLOv5 or Faster R-CNN.
[0041] All samples in the dataset used to train the preset pedestrian detection model have undergone the same preprocessing steps as in step 110 (including Gaussian filtering, contrast enhancement, foreground segmentation, and Gaussian blurring of the face region) to ensure that the preset pedestrian detection model has robust perception capabilities for privacy-protected images.
[0042] In the inference stage of the preset pedestrian detection model, video images are input frame by frame into the preset pedestrian detection model, and the bounding boxes, class labels and confidence scores of each pedestrian are output. After non-maximum suppression (NMS), a set of non-overlapping pedestrian detection boxes in each frame is obtained, and the number of pedestrians is directly counted and the position of each pedestrian is accurately located. Based on the center points of all detection boxes, a continuous and smooth pedestrian density distribution map is generated by Gaussian kernel density estimation (KDE).
[0043] In step 130, if the confidence level is less than a preset threshold, the video image is input into a preset abnormal activity recognition model for recognition to determine the type of abnormal activity of the crowd.
[0044] The preset threshold is an empirical criterion established before system deployment. When the detection confidence level is lower than the preset threshold, it indicates significant uncertainty regarding pedestrian targets in the current frame. This uncertainty may be due to bounding box divergence caused by high overlap of crowds or motion blur caused by rapid movement. While these situations are insufficient for accurate counting and density modeling, they are precisely early visual signs of abnormal activities such as stampedes, blockades, and sudden gatherings. Therefore, a confidence-driven two-level analysis mechanism is adopted: high-confidence results are directly used for regular pedestrian density output, while low-confidence results serve as trigger signals to activate a dedicated preset abnormal activity recognition model to identify the type of abnormal crowd activity. This avoids the resource waste caused by running a high-computational-power anomaly model for every frame, while ensuring timely and accurate capture of potential risk scenarios, achieving an optimal balance between computational efficiency and early warning sensitivity.
[0045] The pre-defined abnormal activity recognition model includes a 3D residual network, a region proposal network, and a region graph convolutional network. The 3D residual network is responsible for extracting the spatiotemporal features of video clips, the region proposal network generates high-quality pedestrian region candidates on the feature map, and the region graph convolutional network models the spatial relationships and motion interactions between these regions using a graph structure. In the embodiment of this application, the 3D residual network has a depth of 50 layers and a total of 16 residual modules. The structure of the 3D residual network is shown in Table 1. Table 1 As shown in Table 1, the kernel size of the first convolutional layer (Conv1) of the 3D residual network is 5×7×7. The kernel sizes of the second, third, fourth, and fifth convolutional layers (Conv2, Conv3, Conv4, and Conv5) are detailed in the second column of Table 1. The time stride of the first convolutional layer (Conv1) is 1, the input segment size is 3×224×224, and the output size is 32×112×112. Specifically, the input downsampling operation is performed by the second, third, fourth, and fifth convolutional layers (Conv2, Conv3, Conv4, and Conv5), all with a stride of 2.
[0046] Region graph convolutional networks abstract pedestrian regions located by region proposal networks into graph nodes and encode the physical space and motion semantic relationship into graph edges. Through end-to-end graph learning, they achieve interpretable discrimination of abnormal group behavior.
[0047] Abnormal crowd activity types can be defined as specific categories of dynamic group behavior that deviate from normal human behavior patterns in surveillance video scenarios and may indicate safety risks or emergencies. These behavior types do not refer to arbitrary non-standard actions, but rather to abnormal event patterns with clear semantics and discriminative characteristics defined through statistical analysis of a large amount of real-world scenario data and expert annotation.
[0048] Typical types of abnormal crowd activity include, but are not limited to: pushing, crowding, running away, lingering, reverse flow, and sudden increases in local density.
[0049] The identification of abnormal activity types relies on joint criteria of multi-source features: on the one hand, the spatiotemporal features captured by the 3D residual network reflect the intensity of motion and temporal continuity of the overall scene; on the other hand, the inter-regional interaction relationship modeled by the region graph convolutional network can reveal whether there is high-density aggregation, directional conflict or sudden displacement of local crowds; in addition, the intensity of crowd movement calculated by combining the area change rate of region proposals in adjacent frames with the displacement distance of the center point can further quantify the degree of dynamic anomaly.
[0050] When the classification confidence after the fusion of the above multi-dimensional features is lower than the preset threshold (for example, the model is not certain enough to distinguish between "normal walking" and "slow gathering"), the system will trigger a more refined abnormal activity recognition model to make a secondary judgment and finally output the specific type of abnormal activity of the crowd.
[0051] Since abnormal crowd activities (e.g., pushing, crowding, fleeing) are essentially anomalous changes in the dynamic spatial relationships and collaborative movement patterns among multiple pedestrians, rather than a simple superposition of individual behaviors, constructing a regional graph convolutional network with the aforementioned structure can explicitly model the physical proximity, movement consistency, and interaction tightness between individuals. Furthermore, by aggregating neighborhood information through graph convolution, abnormal topological patterns such as "reverse movement within short distances," "multi-node centripetal aggregation," and "local high connectivity oscillations" can be automatically learned.
[0052] Specifically, in order to achieve end-to-end modeling of the interaction relationships between pedestrian areas, in some implementations, this pedestrian density analysis method further includes the following steps: (1) The initial region proposal is used as a graph node, and each graph node has a corresponding region feature vector; the initial region proposal is generated by the depth spatiotemporal features extracted by the 3D residual network; (2) For any two graph nodes, the first learnable linear transformation and the second learnable linear transformation are used to map the corresponding region feature vectors to obtain the similarity between any two graph nodes. (3) Construct an adjacency matrix based on the similarity of all node pairs, and normalize each row of the adjacency matrix to obtain a normalized similarity adjacency matrix. (4) Based on the normalized similarity adjacency matrix and learnable convolution kernel parameters, configure graph convolutional layers to form a region graph convolutional network.
[0053] In the embodiments of this application, the basic form of the region graph convolutional network can be represented as: in, It is a diagonal matrix formed by Fourier transform. This refers to multiplication operations between matrices of the same order. As input to graph convolution, The parameters in the Chebyshev polynomial convolution kernel are... Information for each graph node, For all nodes in the graph structure Let be the degree matrix of the nodes. Let be the adjacency matrix of the nodes.
[0054] However, directly applying the above-mentioned spectral domain graph convolution form can easily lead to numerical instability in deep networks. Therefore, this application employs a simplified first-order approximation and introduces self-loop connections to normalize the graph convolution operation, resulting in the following form: in, Given an undirected graph with self-loop connections, the adjacency matrix is... ( )for The corresponding degree matrix, the first row element is No. The sum of all elements in the row. Input signals for different channels The set, The parameter weight matrix of the convolution kernel. This is the output matrix of the normalized convolution, which is equivalent to the output features of each convolutional layer.
[0055] The nodes in the constructed region graph convolutional network are region proposals extracted from video frames through a 3D residual network and a region proposal network. That is, the initial region proposals generated by the region graph convolutional network are used as graph nodes, and each node corresponds to a fixed-dimensional region feature vector (e.g., 256-dimensional) extracted from the deep spatiotemporal feature map output by the 3D residual network. This feature vector simultaneously encodes the appearance representation and short-term motion trend of pedestrians in the region, providing a semantic basis for subsequent relationship modeling.
[0056] The graph structure can be represented as follows: .in, For video frames A regional proposal, For the correlation between regional proposals, These are the learnable weight coefficients.
[0057] To measure the potential interaction strength between any two pedestrian regions, two independent learnable linear transformations (i.e., two sets of weight matrices) are applied to the feature vectors of the two regions respectively, and the dot product of the transformed vectors is calculated as the similarity. This design enables the model to adaptively learn which feature combinations (e.g., position offset + velocity direction) best reflect the discriminative cues of "pushing" or "clustering", rather than relying on manually defined rules.
[0058] In some implementations, any two graph nodes include a first graph node and a second graph node; the step "for any two graph nodes, map their corresponding region feature vectors through a first learnable linear transformation and a second learnable linear transformation respectively to obtain the similarity between any two graph nodes" may include the following steps: (1) Perform a first learnable linear transformation on the region feature vector of the first graph node to obtain the first mapping feature; the first learnable linear transformation is implemented by a d×d learnable weight matrix; (2) Perform a second learnable linear transformation on the regional feature vectors of the nodes in the second graph to obtain the second mapping features; the second learnable linear transformation is implemented by a d×d learnable weight matrix; (3) Determine the similarity between the nodes of the first graph and the nodes of the second graph based on the inner product of the first mapping feature and the second mapping feature.
[0059] In one specific implementation, two regional proposals and Similarity between It can be described as: in, , , and The dimensions are respectively The learnable weight matrix is used. The similarity between regions in a video frame can be calculated using the above equation. The value of each edge is normalized so that its sum is 1. These normalized values form the adjacency matrix of the graph.
[0060] The similarity of all node pairs is organized into an adjacency matrix, and Softmax normalization is performed on each row to ensure that the weighted influence of each node on its neighbors is 1. This ensures the numerical stability of graph propagation and makes the normalized value interpretable as "the probability distribution of a pedestrian being influenced by other pedestrians". Using this normalized adjacency matrix as a graph structure constraint, and combining learnable graph convolution kernel parameters (e.g., GCN layer weights), 2–3 layers of graph convolution operations are stacked to achieve neighborhood feature aggregation. The final output of each node feature has been integrated with its own and the spatiotemporal anomaly clues of the interacting objects, which can be directly used by the classifier to identify specific anomaly types such as "pushing", "clustering", and "running away". Furthermore, high-response regions can be located through gradient backtracking to achieve interpretable early warning.
[0061] Proposed by two regions and Similarity between Substituting the adjacency matrix formed by the corresponding equations into the basic form of the region graph convolutional network and normalizing it, we obtain: in, The output feature map of the convolutional layer has a size of [size missing]. Let be the adjacency matrix of the graph nodes, with size . , It is extracted from each video frame. Composed of regional characteristics A dimensional vector, which is also the node input information of a graph convolutional network; The size is The weight matrix can be learned through backpropagation of the graph convolutional network.
[0062] Compared to traditional adjacency matrices built on geometric distance or fixed thresholds (which often suffer from sparsity leading to a large number of zero elements, thus weakening information propagation and even causing training instability), the adjacency matrix based on learnable similarity used in this application has fully connected characteristics and continuous values. This not only avoids the problem of rigidly severing regional relationships but also enhances the model's ability to model complex crowd interaction patterns. By constructing a region graph convolutional network, adaptive modeling of complex interaction relationships between crowd regions in a video is achieved. Specifically, the network first uses initial region proposals generated by a 3D residual network as graph nodes, each carrying a high-dimensional region feature vector. Then, it maps the features of any two nodes through two independent learnable linear transformations and measures their similarity by the inner product of the mapping results. Next, it constructs an adjacency matrix based on the similarity of all node pairs and normalizes its rows to form a normalized similarity adjacency matrix. Finally, it combines this adjacency matrix with learnable convolutional kernel parameters to configure graph convolutional layers, completing the construction of the region graph convolutional network.
[0063] The aforementioned techniques for constructing regional graph convolutional networks abandon the traditional graph construction methods that rely on manual rules or fixed topologies. This enables the model to learn end-to-end which combinations of regions best reflect the semantic associations of abnormal crowd behavior (e.g., pushing, crowding, or running away), thereby significantly improving the perception and discrimination accuracy of crowd interactions in dense and dynamic scenes.
[0064] Based on the region graph convolutional network constructed above, the identification of abnormal crowd activity types is completed. Specifically, in some embodiments, the step "inputting video images into a preset abnormal activity identification model for identification and determining the type of abnormal crowd activity" may include the following steps: (1) Extract a preset number of key frames from the video image and adjust each key frame to a preset pixel size to form an input video segment; (2) Input the input video clip into the 3D residual network and extract the depth spatiotemporal feature matrix; (3) Based on the deep spatiotemporal feature matrix, an initial region proposal is generated by the region proposal network that removes the bounding box regression module; (4) Perform global average pooling on the deep spatiotemporal feature matrix to generate the first feature vector; (5) Construct a region graph structure based on the initial region proposals; the nodes of the region graph structure are the initial region proposals, and the edges are based on the similarity between the initial region proposals; (6) Input the region graph structure into the region graph convolutional network and output the second feature vector; (7) Calculate the crowd movement intensity and determine the third feature vector based on the area of the initial region proposal and the displacement distance of the center point in adjacent frames; (8) Determine the type of abnormal activity of the population based on the first feature vector, the second feature vector and the third feature vector.
[0065] In one specific implementation, a predetermined number of keyframes are extracted from the input video image to cover a sufficient time window for capturing dynamic changes in the crowd. For example, a 5-10 second video stream is selected, keyframes are extracted at a frequency of 6 frames per second, and 32 consecutive keyframes are selected to form an input video clip. Furthermore, each keyframe in the input video clip is uniformly adjusted to a predetermined pixel size (e.g., 224×224) to ensure consistent input scale, thereby forming a standardized input video clip.
[0066] The input video clip is fed into a pre-trained 3D residual network for processing. The input video clip will eventually be transformed into a deep spatiotemporal feature matrix of size (T×H×W×d) (T is the length of the video frame sequence, H×W represents the size of the feature map, and d represents the number of channels).
[0067] The deep spatiotemporal feature matrix is used for identifying abnormal crowd activity types through two branches. Branch 1: Average pooling is performed on the deep spatiotemporal feature matrix extracted by the 3D residual network to generate a d-dimensional vector (i.e., the first feature vector). In the fully connected layer, this d-dimensional vector is used as external knowledge input to the video stream for crowd activity identification.
[0068] Therefore, the deep spatiotemporal feature matrix simultaneously encodes the appearance information of the video in the spatial dimension and its dynamic evolution features in the temporal dimension. To extract the global semantic summary of the video segment, the system performs a global average pooling operation on the deep spatiotemporal feature matrix along its spatial and temporal dimensions, thereby compressing redundant information and obtaining a compact, fixed-dimensional vector, namely the first feature vector.
[0069] In other words, the first feature vector does not focus on local regional details, but rather reflects the macroscopic state of human movement across the entire input video segment. For example, the overall intensity of movement, the trend of crowd density, or the presence of large-scale abnormal movement.
[0070] Branch 2: Based on the deep spatiotemporal feature matrix, initial region proposals are generated by the Region Proposal Network (RPN) that removes the bounding box regression module. The RPN is derived from the proposal generation mechanism of Fast R-CNN, but only its classification-related part is retained. It uses the deep spatiotemporal feature matrix output by the 3D residual network convolutional layer as input to output several initial region proposals that reflect the active areas of potential populations, providing a node foundation for the subsequent construction of the region map structure.
[0071] By removing the bounding box regression module from the original region proposal network, only a portion of the region information is roughly obtained. That is, only the classification branch is retained to roughly obtain the potential active areas of the population. This avoids ineffective optimization for precise localization in high-density, heavily occluded scenarios, reduces computational complexity, and makes region proposal more focused on semantic saliency rather than geometric precision, thus providing more robust and discriminative node inputs for subsequent graph structure construction.
[0072] Specifically, the region proposal network generates a preset number of regions on the deep spatiotemporal feature matrix using a sliding window approach. There are 10 initial region proposals; each initial region proposal corresponds to a spatial location and a corresponding region feature vector, which is obtained by pooling the response of the feature matrix within that region. The initial region proposals collectively form the graph node set of the region graph convolutional network, used for similarity calculation and region graph structure construction. To model the semantic associations and interactions between different crowd regions within a video frame, each initial region proposal is treated as a node in the graph. This node carries a fixed-dimensional region feature vector extracted from the depth-spatiotemporal feature matrix output by the 3D residual network, and any two nodes are connected by an edge. The weight of this edge is determined by the similarity between the corresponding two initial region proposals, thus completing the construction of the region graph structure.
[0073] The region graph structure is input into a pre-configured Region Graph Convolutional Network (ReGCN) for high-order relation reasoning. The ReGCN performs multi-layer aggregation and updating of graph node features based on a normalized similarity adjacency matrix, and finally outputs a second feature vector that incorporates the semantics of interactions between local regions.
[0074] The second feature vector effectively captures the potential association patterns between sub-regions of the crowd, such as whether there are complex behavioral clues such as directional conflict, dense clustering, or local pushing.
[0075] The system also mines dynamic anomaly signals from a temporal dimension: by tracking the rate of change of the area of the same initial region proposed in adjacent video frames and the displacement distance of its center point, the intensity of local crowd movement is quantified, and the crowd movement intensity is calculated accordingly, thereby generating a third feature vector. This vector can sensitively reflect dynamic anomalies such as sudden running, rapid spread, or abnormal lingering.
[0076] Cross-frame matching is performed using initial region proposals from adjacent frames. To ensure matching reliability and avoid background noise interference, this application employs the Intersection over Union (IoU) ratio as the region association criterion: only when frames... Regional proposals in With frames Regional proposals in A match is considered valid only when its IoU value falls within the preset range [0.5, 0.7]. This range can both exclude background locking caused by high overlap (e.g., static objects) and avoid false matches caused by weak correlation, thus ensuring the accuracy of motion estimation.
[0077] Subsequently, based on the successfully matched region proposals, their areas are extracted. and center point displacement distance The overall motion intensity is calculated using a preset crowd motion intensity formula. Specifically, in some implementations, the step "calculating crowd motion intensity based on the area of the initial region proposal and the displacement distance of the center point in adjacent frames" may include: calculating the crowd motion intensity using a preset crowd motion intensity formula based on the area of the initial region proposal and the displacement distance of the center point in adjacent frames.
[0078] In one specific implementation, the expression for the preset crowd exercise intensity formula can be: in, For the intensity of exercise in the population, For the preset quantity, This represents the number of region proposals in a single keyframe. For the first The area of the initial proposed region. This represents the displacement distance between the center points of two keyframes.
[0079] Compared to traditional motion characterization methods that rely on optical flow or frame difference, which are computationally expensive and susceptible to noise interference, this application constructs a lightweight and robust motion intensity metric by proposing geometric properties (area and displacement) of regions, effectively capturing dynamic anomalies of crowds while ensuring real-time performance.
[0080] Therefore, the first feature vector represents the global spatiotemporal context of the entire input video segment, the second feature vector describes the semantic interaction relationship between regions, and the third feature vector describes the local dynamic evolution trend. By fusing the first, second, and third feature vectors, the type of abnormal crowd activity can be determined.
[0081] Specifically, in some implementations, the step of "determining the type of abnormal crowd activity based on the first feature vector, the second feature vector, and the third feature vector" may include the following steps: (1) Perform global average pooling on the second feature vector to obtain the pooled second feature vector; (2) Concatenate the pooled second feature vector, the first feature vector, and the third feature vector to obtain the fused feature; (3) Input the fused features into the fully connected layer and output the identification results of abnormal activity types of the crowd.
[0082] The second feature vector output by the region graph convolutional network (its dimension is...) ,in This represents the initial number of proposed regions. Perform global average pooling on the feature channel count, averaging along the node dimension (i.e., the region dimension) to obtain a... The second feature vector after pooling in dimension 1. This operation effectively aggregates the interaction semantic information of all region nodes, generating a compact global descriptor that represents the overall region relationship pattern.
[0083] Subsequently, the pooled second feature vector is concatenated with the first feature vector (the result of global average pooling from the deep spatiotemporal feature matrix, reflecting the macroscopic state at the video level) and the third feature vector (calculated based on the area change and center point displacement of the initial region proposal in adjacent frames, representing the local dynamic intensity) to form a high-dimensional fused feature. This fused feature simultaneously covers three key dimensions: global context, regional interaction semantics, and spatiotemporal motion dynamics, constituting a multi-granular joint representation of crowd behavior.
[0084] Finally, the fused features are input into at least one fully connected layer, and the probability distribution of each preset abnormal activity category is output through the Softmax function or other classifiers to obtain the final identification result of the abnormal activity type of the crowd, such as "pushing", "gathering", "running away", "lingering" or "normal movement".
[0085] By fusing the first, second, and third feature vectors, not only are the discriminative advantages of each feature branch retained, but end-to-end training also enables the fully connected layer to automatically learn the importance weights of each feature dimension, significantly improving the model's generalization ability and recognition accuracy in complex real-world scenarios.
[0086] Therefore, this application determines the types of abnormal crowd activities to avoid simply assessing the "high or low density of people" when analyzing crowd density in the monitored scene. This application achieves a leap from quantitative statistics to qualitative judgment by identifying specific types of abnormal activities (e.g., "gathering", "pushing", "reverse flow").
[0087] In step 140, a heat map of pedestrian density is generated based on the pedestrian density distribution map and using the Gaussian kernel density estimation algorithm.
[0088] To transform pedestrian density distribution maps into intuitive visual information, we employed the Gaussian Kernel Density Estimation (KDE) algorithm to generate pedestrian density heatmaps. For example, relying on precisely acquired pedestrian density distribution data, we then processed this data using Gaussian kernel density estimation techniques to smooth and refine the pedestrian distribution patterns, thereby generating a continuous and detailed pedestrian density surface.
[0089] The generated heatmap uses color coding to visually represent the population density levels of different areas. Red represents high-density areas, indicating areas with high pedestrian traffic; green represents low-density areas, indicating relatively sparse pedestrian traffic. This color mapping method allows observers to quickly identify crowd hotspots and potential safety hazards.
[0090] In some implementations, to ensure adequate protection of personal privacy, all individual identification information is removed during the heatmap generation process, retaining only statistically significant density information. This results in a final heatmap that contains no personally identifiable data, focusing instead on displaying overall density trends rather than the behavioral patterns of specific individuals.
[0091] In some implementations, to further enhance data security, all heatmap data is stored in a distributed database and encrypted to ensure the security of data transmission and storage, preventing unauthorized access or disclosure.
[0092] In step 150, the pedestrian locations, number of pedestrians, types of abnormal crowd activities, and crowd density heatmaps are visualized on the monitoring platform.
[0093] The location of pedestrians, the number of pedestrians, the types of abnormal crowd activities, and the heat map of crowd density are all integrated into the monitoring platform for visualization, so as to support efficient and intuitive safety management decisions.
[0094] For example, the monitoring platform is equipped with a crowd monitoring system that can render real-time heat maps of crowd density in various areas (using red-green color coding to represent high / low density) and overlay the detected pedestrian locations and the number of pedestrians in the area. When abnormal crowd activities such as "pushing," "gathering," or "running away" are identified, the system will automatically mark alarm icons in the corresponding geographical area and accurately map the location of the event using a GIS map, making it easier for security personnel to quickly locate and respond.
[0095] In some implementations, this visualization interface is not only used for real-time monitoring but also linked to a feedback mechanism: once an abnormal density event (such as sudden congestion) occurs, the relevant data will be recorded and fed back to the model for incremental training, continuously optimizing detection accuracy; simultaneously, all historical heatmaps, abnormal events, and statistical data are stored in a distributed database, supporting quick queries by time and location, assisting in security trend analysis. The entire system ensures the accuracy, real-time performance, and long-term stability of the visualization display through regular maintenance of cameras and edge devices, and updates to deep learning models and database software.
[0096] Please see Figure 2 , Figure 2This illustration shows a schematic diagram of a crowd density analysis device according to an embodiment of this application. The crowd density analysis device 200 includes: an acquisition module 210, a first determination module 220, a second determination module 230, a generation module 240, and a display module 250. Specifically: The acquisition module 210 is used to acquire video images of the scene to be monitored; The first determining module 220 is used to input video images into a preset pedestrian detection model for inference, and determine the pedestrian position, number of pedestrians, pedestrian density distribution map and confidence level in each frame image; the preset pedestrian detection model is a single-stage or two-stage deep learning detection architecture; The second determining module 230 is used to input the video image into a preset abnormal activity recognition model for recognition if the confidence level is less than a preset threshold, and to determine the type of abnormal activity of the crowd; the preset abnormal activity recognition model includes a 3D residual network, a region proposal network and a region graph convolutional network. The generation module 240 is used to generate a heat map of pedestrian density based on the pedestrian density distribution map and using the Gaussian kernel density estimation algorithm. The display module 250 is used to visualize pedestrian locations, pedestrian numbers, types of abnormal crowd activities, and crowd density heat maps on the monitoring platform.
[0097] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device and module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0098] In the several embodiments provided in this application, the coupling or direct coupling or communication connection between the modules shown or discussed may be an indirect coupling or communication connection through some interface, device or module, and may be electrical, mechanical or other forms.
[0099] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0100] Please see Figure 3 , Figure 3The diagram shows a structural schematic of a crowd density analysis device provided in an embodiment of this application. The crowd density analysis device 300 in this application may include one or more of the following components: a processor 310, a memory 320, and one or more application programs. The one or more application programs may be stored in the memory 320 and configured to be executed by one or more processors 310. The one or more programs are configured to perform the crowd density analysis method as described in the foregoing method embodiments.
[0101] The processor 310 may include one or more processing cores. The processor 310 connects to various parts within the crowd density analysis device 300 via various interfaces and lines, and performs various functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 320, and by calling data stored in the memory 320. Optionally, the processor 310 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 310 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the displayed content; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 310 and may be implemented separately using a communication chip.
[0102] The memory 320 may include random access memory (RAM) or read-only memory (ROM). The memory 320 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 320 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function, instructions for implementing the various method embodiments described below, etc. The data storage area may also store data created during the use of the crowd density analysis device 300.
[0103] Please see Figure 4 , Figure 4The diagram shows a computer-readable storage medium 400 provided in an embodiment of this application. The computer-readable storage medium 400 stores program code, which can be called by a processor to execute the crowd density analysis method described in the above method embodiment.
[0104] The computer-readable storage medium 400 may be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Optionally, the computer-readable storage medium 400 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 400 has storage space for program code 410 that performs any of the method steps described above. This program code can be read from or written to one or more computer program devices. The program code 410 may be compressed, for example, in a suitable form.
[0105] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for analyzing pedestrian density, characterized in that, The method includes: Acquire video images of the scene to be monitored; The video images are input into a preset pedestrian detection model for inference to determine the pedestrian positions, number of pedestrians, pedestrian density distribution map, and confidence level in each frame of the image; the preset pedestrian detection model is a single-stage or two-stage deep learning detection architecture. If the confidence level is less than a preset threshold, the video image is input into a preset abnormal activity recognition model for recognition to determine the type of abnormal activity of the crowd; the preset abnormal activity recognition model includes a 3D residual network, a region proposal network, and a region graph convolutional network; Based on the aforementioned pedestrian density distribution map, a pedestrian density heatmap is generated using a Gaussian kernel density estimation algorithm. The pedestrian locations, the number of pedestrians, the types of abnormal crowd activities, and the crowd density heatmap are visualized on the monitoring platform.
2. The method for analyzing pedestrian density according to claim 1, characterized in that, The acquisition of video images of the scene to be monitored includes: The image acquisition module deployed in the scene to be monitored acquires raw video frames at a preset frame rate; The original video frame is subjected to Gaussian filtering to obtain a denoised video frame; The denoised video frames are then contrast-enhanced to obtain enhanced video frames. Foreground target segmentation is performed on the enhanced video frames to extract regions of interest containing pedestrians; Based on a face detection algorithm, the face region within the region of interest is located, and the face region is blurred to obtain the video image.
3. The method for analyzing pedestrian density according to claim 1, characterized in that, The step of inputting the video image into a preset abnormal activity recognition model for recognition and determining the type of abnormal activity in the crowd includes: Extract a preset number of keyframes from the video image, and adjust each keyframe to a preset pixel size to form an input video segment; The input video clip is input into the 3D residual network to extract the depth spatiotemporal feature matrix; Based on the aforementioned deep spatiotemporal feature matrix, an initial region proposal is generated through the region proposal network of the bounding box removal regression module. Global average pooling is performed on the depth spatiotemporal feature matrix to generate the first feature vector; A region graph structure is constructed based on the initial region proposals; the nodes of the region graph structure are the initial region proposals, and the edges are based on the similarity between the initial region proposals. The region graph structure is input into the region graph convolutional network, and the second feature vector is output. Based on the area of the initial region proposal and the displacement distance of the center point in adjacent frames, the crowd movement intensity is calculated, and the third feature vector is determined. The abnormal activity type of the population is determined based on the first feature vector, the second feature vector, and the third feature vector.
4. The method for analyzing pedestrian density according to claim 3, characterized in that, Determining the type of abnormal activity in the population based on the first feature vector, the second feature vector, and the third feature vector includes: The second feature vector is subjected to global average pooling to obtain the pooled second feature vector. The pooled second feature vector, the first feature vector, and the third feature vector are concatenated to obtain the fused feature; The fused features are input into a fully connected layer, and the identification results of the abnormal activity type of the crowd are output.
5. The method for analyzing pedestrian density according to claim 3, characterized in that, The calculation of crowd movement intensity based on the area proposed in the initial region and the displacement distance of the center point in adjacent frames includes: Based on the area of the initial region proposal and the displacement distance of the center point in adjacent frames, the crowd movement intensity is calculated using a preset crowd movement intensity formula. The expression for the preset population exercise intensity formula is: in, The exercise intensity of the aforementioned population. For the preset quantity, This represents the number of region proposals in a single keyframe. For the first The area of the initial proposed region. This represents the displacement distance between the center points of two keyframes.
6. The method for analyzing pedestrian density according to claim 1, characterized in that, The method further includes: The initial region proposal is used as a graph node, and each graph node has a corresponding region feature vector; the initial region proposal is generated by the depth spatiotemporal features extracted by the 3D residual network; For any two graph nodes, the similarity between the two graph nodes is obtained by mapping their corresponding region feature vectors through a first learnable linear transformation and a second learnable linear transformation. An adjacency matrix is constructed based on the similarity of all node pairs, and each row of the adjacency matrix is normalized to obtain a normalized similarity adjacency matrix. Based on the normalized similarity adjacency matrix and learnable convolutional kernel parameters, graph convolutional layers are configured to form the region graph convolutional network.
7. The method for analyzing pedestrian density according to claim 6, characterized in that, The two graph nodes include the first graph node and the second graph node; The step of mapping the corresponding region feature vectors of any two graph nodes using a first learnable linear transformation and a second learnable linear transformation respectively to obtain the similarity between the two graph nodes includes: A first learnable linear transformation is performed on the region feature vector of the first graph node to obtain a first mapping feature; the first learnable linear transformation is... Learnable weight matrix implementation; A second learnable linear transformation is performed on the region feature vectors of the second graph nodes to obtain the second mapping feature; the second learnable linear transformation is derived from... Learnable weight matrix implementation; The similarity between the first graph node and the second graph node is determined based on the inner product of the first mapping feature and the second mapping feature.
8. A crowd density analysis device, characterized in that, The device includes: The acquisition module is used to acquire video images of the scene to be monitored. The first determining module is used to input the video images into a preset pedestrian detection model for inference, and determine the pedestrian positions, number of pedestrians, pedestrian density distribution map and confidence level in each frame image; the preset pedestrian detection model is a single-stage or two-stage deep learning detection architecture; The second determining module is used to input the video image into a preset abnormal activity recognition model for recognition if the confidence level is less than a preset threshold, and determine the type of abnormal activity of the crowd; the preset abnormal activity recognition model includes a 3D residual network, a region proposal network and a region graph convolutional network; The generation module is used to generate a heat map of pedestrian density based on the pedestrian density distribution map and using a Gaussian kernel density estimation algorithm. The display module is used to visualize the pedestrian location, the number of pedestrians, the types of abnormal crowd activities, and the crowd density heat map on the monitoring platform.
9. A crowd density analysis device, characterized in that, include: One or more processors; Memory; One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, the one or more applications being configured to perform the crowd density analysis method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium contains program code, which can be called by a processor to execute the crowd density analysis method as described in any one of claims 1-7.