Passenger flow statistical system and statistical method based on deep learning
By processing visible light and infrared thermal imaging data through a dual-channel convolutional neural network and a spatiotemporal graph convolutional network, and combining dynamic target tracking and feature optimization, the accuracy and real-time performance issues of passenger flow statistics in complex scenarios are solved, achieving high-precision passenger flow statistics and real-time control.
Patent Information
- Application Number
- CN202511068910.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-11-14
AI Technical Summary
Existing passenger flow statistics methods are inadequate in terms of environmental adaptability, density sensitivity, and real-time performance. In particular, their accuracy and real-time performance are insufficient to meet practical needs in low-light and high-density scenarios.
A parallel dual-channel convolutional neural network is used to process visible light video and infrared thermal imaging data. It combines feature-level fusion, dynamic target tracking and spatiotemporal passenger flow analysis, optimizes feature extraction and tracking through optical flow and Kalman filtering, and uses a spatiotemporal graph convolutional network for prediction.
It improves the accuracy and real-time performance of passenger flow statistics, achieving an accuracy rate of 93% in low-light environments, reducing the ID switching rate to 1.5% in high-density environments, and reducing the prediction latency to less than 1 second, thus meeting the high-precision requirements in complex scenarios.
Smart Images

Figure CN120954052A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer vision and data analysis, specifically relating to a passenger flow statistics system and statistical method based on deep learning. Background Technology
[0002] In today's society, accurate passenger flow statistics are crucial for various public places such as power company service halls, shopping malls, stations, and scenic spots. Traditional passenger flow statistics methods, such as those based on infrared sensing and pressure sensing, have many limitations. Infrared sensing is easily affected by environmental interference, and counting errors can easily occur when multiple people pass through the sensing area at the same time; pressure sensing has high requirements for installation location and ground conditions, and cannot distinguish between different individuals, making it difficult to guarantee statistical accuracy.
[0003] With the development of computer vision technology, video image-based methods for passenger flow statistics have gradually emerged. Early video-based passenger flow statistics methods mainly relied on manually designed features, such as HOG (Histogram of Oriented Gradients) and LBP (Local Binary Pattern), and then combined them with traditional classifiers such as SVM (Support Vector Machine) for pedestrian detection. However, these methods perform poorly in complex scenes, such as when there are drastic changes in lighting or severe pedestrian occlusion, the detection accuracy drops significantly. This is because manually designed features are difficult to comprehensively and accurately describe the complex features of pedestrians, and the generalization ability of traditional classifiers is limited.
[0004] The emergence of deep learning technology has brought new breakthroughs to passenger flow statistics. Deep learning can automatically learn more representative features from large amounts of data, thereby improving the accuracy of detection and statistics. However, current deep learning-based passenger flow statistics algorithms still have some problems. On the one hand, model training requires a large amount of labeled data, and manual data labeling is not only time-consuming and laborious, but also prone to labeling errors. On the other hand, existing models still need further improvement in handling multi-target tracking and real-time performance in complex scenarios. For example, in scenarios with high population density, such as shopping mall promotional events, existing algorithms may experience problems such as target loss and duplicate counting, failing to meet the needs of practical applications for high-precision passenger flow statistics.
[0005] The existing passenger flow statistics system has three major technical bottlenecks:
[0006] 1. Poor environmental adaptability: Traditional visible light solutions have a false detection rate of >40% in low-light scenarios;
[0007] 2. Density-sensitive defect: The fixed-parameter algorithm has an ID switching rate as high as 35% in crowded scenarios (ρ>4 people / ㎡);
[0008] 3. Disconnect between forecasting and control: The delay between statistical results and emergency decision-making response is greater than 8 seconds.
[0009] The invention patent with publication number CN113591876A only implements static scene counting and does not solve the problem of algorithm switching under dynamic density; the invention patent with publication number CN110675443A discloses a thermal imaging fusion scheme that does not involve a feature-level weighting mechanism.
[0010] Therefore, developing a deep learning-based passenger flow statistics method that can perform passenger flow statistics more accurately and efficiently and adapt to complex scenarios is of great practical significance. Summary of the Invention
[0011] The purpose of this invention is to provide a passenger flow statistics system and method based on deep learning to address the problems existing in the prior art, so as to solve the shortcomings of existing passenger flow statistics methods in terms of accuracy, adaptability, density sensitivity, and disconnect from control.
[0012] To achieve the above objectives, the deep learning-based passenger flow statistics system of the present invention includes:
[0013] Multimodal feature extraction module: It adopts a parallel dual-channel convolutional neural network (CNN). The first channel processes the visible light video stream, and the second channel processes the infrared thermal imaging data. The enhanced pedestrian feature map is output through the feature-level fusion layer.
[0014] Dynamic target tracking module: Generates initial values of motion trajectory based on optical flow method, integrates Kalman filter prediction results from SORT algorithm, and dynamically adjusts tracking weights through appearance feature matching function;
[0015] Spatiotemporal Passenger Flow Analysis Module: Utilizes a Spatiotemporal Graph Convolutional Network (STGCN) to perform spatiotemporal modeling of pedestrian trajectories, outputting a regional passenger flow density heatmap and future traffic prediction curves;
[0016] Feedback optimization module: Adaptively adjusts the receptive field size of the CNN feature extraction layer and the search window size of the optical flow method based on the tracking loss rate in occluded scenarios.
[0017] Specifically, the calculation method of the feature-level fusion layer is as follows:
[0018] F fusion =α·σ(Wv·F RGB )+(1-α)·ReLU(Wt·F thermal );
[0019] α = 1 - e ﹣β·Ι
[0020] Where F RGB For visible light characteristic maps, F thermalHere, Wv and Wt are the thermal imaging feature maps, σ is the adaptive light intensity coefficient, β is the ambient light intensity, and I is the attenuation factor. The system relies on thermal imaging features, resulting in a 32% reduction in the false detection rate.
[0021] Compared to data-level fusion (direct fusion of raw data), feature-level fusion extracts key features from each data source before fusion. This removes redundant information while retaining core features useful for the task, improving fusion efficiency. Extracting features from each data source before fusion avoids the high computational cost of directly processing massive amounts of raw data. It also provides more concise input for subsequent models (such as classifiers and predictors), reducing overall computational load and complexity.
[0022] Specifically, the appearance feature matching function is defined as:
[0023] S match =λ·loU(B t B t-1 )+(1-λ)·‖φ(f t )-φ(f t-1 )‖2
[0024]
[0025] Where λ is the dynamic adjustment coefficient, ρ is the population density, and ρ max For the maximum population density, φ(f) t ), φ(f t-1 ) are the appearance feature vectors of the current frame and the previous frame extracted by the CNN, respectively, and |||2 is the L2 norm, loU(B t B t-1 The value represents the intersection-union ratio (IoU) of the target bounding boxes in two consecutive frames, dynamically adjusted based on the crowd density ρ. Appearance features dominate matching, resulting in a 41% decrease in ID switching rate.
[0026] Specifically, the feedback optimization module performs the following operations:
[0027] When the tracking loss rate L track When the occlusion rate is >15%, increasing the kernel size of the last two layers of the CNN backbone by 50% increases the recall rate of occluded scenes by 28%.
[0028] When the average displacement between consecutive frames D move When the pixel size is less than 5 pixels, the search window for the optical flow method is reduced from 32×32 to 16×16, which reduces the computation time by 45%.
[0029] Specifically, the spatiotemporal passenger flow analysis module includes:
[0030] Spatiotemporal graph construction unit: Constructs a spatiotemporal graph with pedestrian trajectory points as nodes and motion relationships as edges;
[0031] Gated graph convolutional units: aggregate spatiotemporal neighborhood information and update node states through a gating mechanism;
[0032] Spatiotemporal prediction unit: based on the final layer node state h i (L) The system predicts passenger flow distribution after time T using a spatiotemporal attention mechanism. The predicted MAE is as low as 6.8% (compared to >22% by traditional methods).
[0033] A deep learning-based passenger flow statistics method, applied to the system, includes the following steps:
[0034] S1. Synchronous acquisition and time alignment of visible light and thermal imaging video streams;
[0035] S2. Extract multimodal features using a dual-channel CNN and fuse them according to the calculation formula of the feature-level fusion layer;
[0036] S3. Initialize pedestrian detection boxes based on fused features, and perform cross-frame target association according to the appearance feature matching function;
[0037] S4. Perform dynamic optimization of feature extraction and tracking parameters based on the feedback optimization module;
[0038] S5. Generate a regional passenger flow heat map and future prediction curve as the passenger flow prediction result using the spatiotemporal passenger flow analysis module.
[0039] Specifically, in step S2, a transfer learning strategy is used for fusion:
[0040] The visible light channel is loaded with COCO pre-trained weights, and the thermal imaging channel is loaded with FLIR pre-trained weights.
[0041] The weights of the fusion layer are updated through end-to-end joint fine-tuning.
[0042] Specifically, the passenger flow prediction results in S5 are used to trigger control commands:
[0043] When the area density is greater than the safe density, a diversion alarm signal is generated. The purpose is to disperse the flow of people or vehicles in the area by reminding them to avoid the risks caused by excessive density.
[0044] When the predicted traffic exceeds the maximum capacity, a flow control command is generated. This is a direct intervention measure that ensures that the traffic does not exceed the system or area's capacity limit by restricting the incoming flow.
[0045] This deep learning-based passenger flow statistics method has several significant technical advantages:
[0046] Employing a deep learning-based passenger flow statistics algorithm, this system enables high-precision personnel detection and tracking in real-time video data under complex scenarios. By extracting spatial features from each frame of the image using a convolutional neural network (CNN), the system can automatically detect personnel targets in the video and track and count personnel based on motion information between consecutive frames. Compared with traditional image processing algorithms, the deep learning algorithm can effectively handle complex conditions such as occlusion and changes in lighting in the scene, ensuring the accuracy and robustness of passenger flow statistics. Full-scene robustness: accuracy >93% in low light (I<20lux), and ID switching rate <1.5% in high density (ρ>5 people / ㎡).
[0047] The system relies on convolutional neural networks (CNNs) for person detection and combines optical flow or target tracking algorithms (such as Kalman filtering or SORT) for continuous person tracking. First, the CNN model processes each frame of the video stream to identify and locate pedestrian targets in the image. Then, based on the target tracking algorithm, the system can track the motion trajectory of the same target across frames, preventing double counting or missed counts. Incorporating a time dimension, the system can analyze passenger flow trends over different time periods, providing data support for traffic control and management in the scenario. The latency from prediction to the generation of control instructions is less than 1 second.
[0048] The passenger flow statistics system and method based on deep learning provided by this invention employs dynamic fusion of illumination and thermal imaging (α function) to address environmental adaptability deficiencies; a density-driven matching mechanism (λ dynamic adjustment) overcomes the challenge of high-density tracking; loss rate feedback optimization enables self-adjustment of algorithm parameters; and a spatiotemporal graph prediction engine establishes a closed loop between statistics and decision-making. It achieves good results in various scenarios, effectively improving the accuracy, adaptability, and real-time performance of passenger flow statistics, and has broad application prospects and practical value. Attached Figure Description
[0049] Figure 1 This is a flowchart of the passenger flow statistics system based on deep learning according to the present invention;
[0050] Figure 2 This is a flowchart of the spatiotemporal passenger flow analysis module. Detailed Implementation
[0051] The technical solution of the present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.
[0052] Example 1
[0053] This embodiment provides a passenger flow statistics system based on deep learning, including:
[0054] Multimodal feature extraction module: It adopts a parallel dual-channel convolutional neural network (CNN). The first channel processes the visible light video stream, and the second channel processes the infrared thermal imaging data. The enhanced pedestrian feature map is output through the feature-level fusion layer.
[0055] Dynamic target tracking module: Generates initial values of motion trajectory based on optical flow method, integrates Kalman filter prediction results from SORT algorithm, and dynamically adjusts tracking weights through appearance feature matching function;
[0056] Spatiotemporal Passenger Flow Analysis Module: Utilizes a Spatiotemporal Graph Convolutional Network (STGCN) to perform spatiotemporal modeling of pedestrian trajectories, outputting a regional passenger flow density heatmap and future traffic prediction curves;
[0057] Feedback optimization module: Adaptively adjusts the receptive field size of the CNN feature extraction layer and the search window size of the optical flow method based on the tracking loss rate in occluded scenarios.
[0058] In this embodiment, the feature-level fusion layer is calculated as follows:
[0059] F fusion =α·σ(Wv·F RGB )+(1-α)·ReLU(Wt·F thermal );
[0060] α = 1 - e ﹣β·Ι
[0061] Where F RGB For visible light characteristic maps, F thermal Here, Wv and Wt are the thermal imaging feature maps, σ is the adaptive coefficient for light intensity, β is the ambient light intensity, and I is the attenuation factor. The system depends on the thermal imaging features; in low light conditions (I < 50 lux), α is approximately 0, and the false detection rate decreases by 32%.
[0062] Compared to data-level fusion (direct fusion of raw data), feature-level fusion extracts key features from each data source before fusion. This removes redundant information while retaining core features useful for the task, improving fusion efficiency. Extracting features from each data source before fusion avoids the high computational cost of directly processing massive amounts of raw data. It also provides more concise input for subsequent models (such as classifiers and predictors), reducing overall computational load and complexity.
[0063] In this embodiment, the appearance feature matching function is defined as:
[0064] S match =λ·loU(B t B t-1 )+(1-λ)·‖φ(f t )-φ(ft-1 )‖2
[0065]
[0066] Where λ is the dynamic adjustment coefficient, ρ is the population density, and ρ max For the maximum population density, φ(f) t ), φ(f t-1 ) are the appearance feature vectors of the current frame and the previous frame extracted by the CNN, respectively, and |||2 is the L2 norm, loU(B t B t-1 The cross-union ratio (CUR) represents the intersection-over-union ratio of the target bounding boxes in two consecutive frames, dynamically adjusted according to the crowd density ρ. High density (ρ = 4 people / m²) 2 When λ = 0.4, appearance features dominate matching, and the ID switching rate decreases by 41%.
[0067] Furthermore, the feedback optimization module performs the following operations:
[0068] When the tracking loss rate L track When the rate is >15%, it indicates that the current network may not be able to extract the features of the target (e.g., target occlusion, pose changes, etc. lead to inaccurate feature capture). Increasing the kernel size of the last two layers of the CNN backbone network by 50% can expand the feature receptive field, enhance the ability to capture global features and contextual information of the target, and thus may reduce the tracking loss rate and increase the recall rate in occluded scenes by 28%.
[0069] When the average displacement between consecutive frames D move When the distance is less than 5px, it indicates that the target's movement is small (such as slow movement or near-stationary), and a large search area is not required to capture the displacement information. Reducing the search window of the optical flow method from 32×32 to 16×16 ensures accurate calculation of small displacements while reducing the computational load (the search window area is reduced to 1 / 4 of its original size), thereby improving algorithm efficiency and reducing computation time by 45%. This is an adaptive optimization strategy that balances accuracy and speed, suitable for tasks that rely on optical flow calculations, such as video target tracking and motion analysis.
[0070] In this embodiment, the spatiotemporal passenger flow analysis module includes:
[0071] Spatiotemporal graph construction unit: Constructs a spatiotemporal graph with pedestrian trajectory points as nodes and motion relationships as edges;
[0072] Gated graph convolutional units: aggregate spatiotemporal neighborhood information and update node states through a gating mechanism;
[0073] Spatiotemporal prediction unit: based on the final layer node state h i (L)The system predicts passenger flow distribution after time T using a spatiotemporal attention mechanism. The predicted MAE is as low as 6.8% (compared to >22% by traditional methods). Specific implementation steps are as follows: Figure 2 As shown.
[0074] In this embodiment, the Gated Graph Convolutional Unit (GGCNU) combines a Graph Convolutional Network (GCN) and a gating mechanism to effectively aggregate spatiotemporal neighborhood information to predict passenger flow distribution at time T. The specific formulas and operations involved are as follows:
[0075] 1. First, perform graph convolution operation.
[0076] Suppose we have a graph structure G = (V, E), where V is the set of nodes and E is the set of edges. For each node i, its feature is represented as h_i^{(l)}, where l represents the number of convolutional layers in the graph.
[0077] The graph convolution operation can be represented as:
[0078]
[0079] Where: N(i) is the set of neighboring nodes of node i. ij It is a normalization constant, usually set as the degree d of node i. i Or a more complex normalized form based on the Graph Laplace matrix (such as...) (This is common in GCNs related to symmetric normalized Laplace matrices), W (l) is the weight matrix of the l-th layer, used to perform a linear transformation on the features of neighboring nodes. (l) It is the bias vector of the l-th layer.
[0080] 2. Gating mechanism
[0081] Introducing gating mechanisms to control the flow of information, including updating gates. and reset door The calculation formula is as follows:
[0082] Update Gate:
[0083] Reset Door:
[0084] in:
[0085] σ is the sigmoid activation function, which compresses the input to the (0,1) interval and is used to calculate the weights of the gate value.
[0086] and These are the weight matrices for the update gate and the reset gate, respectively.
[0087] and These are the bias vectors for the update gate and the reset gate, respectively.
[0088] This indicates the sum of features after convolution of the graph. Features of node i at the previous time step Then, the parts are assembled.
[0089] 3. Hidden state update
[0090] Candidate hidden states are calculated using gate resetting and graph convolution results. Then, the hidden state of the node is updated using the update gate.
[0091] Candidate hidden state:
[0092]
[0093] Where: tanh is the hyperbolic tangent activation function, which maps the input to the interval (-1,1).
[0094] W (l)′ It is the weight matrix used to calculate the candidate hidden states, b (l)′ It is the corresponding bias vector.
[0095] ⊙ represents element-wise multiplication, i.e., the reset gate. Convolution results of graphs Perform element-wise multiplication.
[0096] Hidden status update:
[0097]
[0098] This formula indicates that the hidden state of the current node is the same as the hidden state of the previous time step. and candidate hidden state The weighting and weighting are determined by the update gate. Decide.
[0099] 4. Predict passenger flow distribution at time T in the future
[0100] After multiple gated graph convolution operations, the final hidden state representation of the nodes is obtained. (L is the total number of layers). Then, a fully connected layer maps the hidden state to the predicted passenger flow distribution, assuming the prediction function is f:
[0101]
[0102] Among them W out It is the weight matrix of the fully connected layer, bout It is a bias vector. This refers to the predicted passenger flow distribution at node i at time T in the future. The predicted values can be visualized using charts for easy observation.
[0103] Example 2
[0104] This embodiment provides a deep learning-based passenger flow statistics method, applied to the system described in Embodiment 1, including the following steps:
[0105] S1. Synchronous acquisition and time alignment of visible light and thermal imaging video streams;
[0106] S2. Extract multimodal features using a dual-channel CNN and fuse them according to the calculation formula of the feature-level fusion layer;
[0107] S3. Initialize pedestrian detection boxes based on fused features, and perform cross-frame target association according to the appearance feature matching function;
[0108] S4. Perform dynamic optimization of feature extraction and tracking parameters based on the feedback optimization module;
[0109] S5. Generate a regional passenger flow heat map and future prediction curve as the passenger flow prediction result using the spatiotemporal passenger flow analysis module.
[0110] In this embodiment, step S2 employs a transfer learning strategy for fusion, specifically including:
[0111] The visible light channel is loaded with COCO pre-trained weights, and the thermal imaging channel is loaded with FLIR pre-trained weights.
[0112] The weights of the fusion layer are updated through end-to-end joint fine-tuning.
[0113] In this embodiment, the passenger flow prediction result in S5 is used to trigger control commands:
[0114] When the area density is greater than the safe density, a diversion alarm signal is generated. The purpose is to disperse the flow of people or vehicles in the area by reminding them to avoid the risks caused by excessive density.
[0115] When the predicted traffic exceeds the maximum capacity, a flow control command is generated. This is a direct intervention measure that ensures that the traffic does not exceed the system or area's capacity limit by restricting the incoming flow.
[0116] Example 3:
[0117] This embodiment focuses on controlling passenger flow during the morning rush hour on the subway. The specific steps are as follows:
[0118] 1. Data Acquisition:
[0119] Simultaneous acquisition using a visible light camera and a thermal imager (30fps).
[0120] 2. Feature fusion:
[0121] With ambient illuminance I = 30 lux, α = 0.2, and thermal imaging weighting at 80%,
[0122] Load COCO+FLIR pre-trained weights.
[0123] 3. Tracking and Optimization:
[0124] Detection density ρ = 5.2 people / m² → λ = 0.38
[0125] Tracking loss rate L track =18% → Increase the kernel size of the last two convolutional layers of ResNet50 to 9×9.
[0126] 4. Predictive control:
[0127] STGCN predicts area A 5 minutes later.
[0128] Trigger gate flow control: R = 120 × (1 - 6.8 / 8) = 30 people / minute.
[0129] Comparison of effects:
[0130] index Traditional solution This invention Counting accuracy 76.3% 98.1% Predicting MAE 22.7% 6.8% Command response delay 8.5s 0.3s
[0131] The above examples demonstrate robustness across all scenarios:
[0132] Accuracy > 93% in low light conditions (I < 20 lux)
[0133] High-density (ρ>5 people / ㎡) ID switching rate <1.5%.
[0134] The deep learning-based passenger flow statistics method of this invention can achieve good results in different scenarios, effectively improving the accuracy, adaptability and real-time performance of passenger flow statistics, and has broad application prospects and practical value.
[0135] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them; although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that: any modifications to the specific implementation of the present invention or equivalent substitutions to some technical features can still be made within the spirit and principles of the present invention without departing from the spirit of the technical solutions of the present invention, and all such modifications and substitutions should be covered within the scope of the technical solutions claimed in the present invention.
Claims
1. A passenger flow statistics system based on deep learning, characterized in that, include Multimodal feature extraction module: It adopts a parallel dual-channel convolutional neural network (CNN). The first channel processes the visible light video stream, and the second channel processes the infrared thermal imaging data. The enhanced pedestrian feature map is output through the feature-level fusion layer. Dynamic target tracking module: Generates initial values of motion trajectory based on optical flow method, integrates Kalman filter prediction results from SORT algorithm, and dynamically adjusts tracking weights through appearance feature matching function; Spatiotemporal Passenger Flow Analysis Module: Utilizes a Spatiotemporal Graph Convolutional Network (STGCN) to perform spatiotemporal modeling of pedestrian trajectories, outputting a regional passenger flow density heatmap and future traffic prediction curves; Feedback optimization module: Adaptively adjusts the receptive field size of the CNN feature extraction layer and the search window size of the optical flow method based on the tracking loss rate in occluded scenarios.
2. The passenger flow statistics system based on deep learning according to claim 1, characterized in that, The calculation method for the feature-level fusion layer is as follows: F fusion =α·σ(Wv·F RGB )+(1-α)·ReLU(Wt·F thermal ); α=1-e ﹣β·Ι Where F RGB For visible light characteristic maps, F thermal For thermal imaging feature map, Wv and Wt are trainable weight matrices, σ is the light intensity adaptive coefficient, β is the ambient light intensity, and Ι is the attenuation factor.
3. The passenger flow statistics system based on deep learning according to claim 1, characterized in that, In the data preprocessing step, the appearance feature matching function is defined as: S match =λ·loU(B t ,B t-1 )+(1-λ)·‖φ(f t )-φ(f t-1 )‖2 Where λ is the dynamic adjustment coefficient, ρ is the population density, and ρ max For the maximum population density, φ(f) t ), φ(f t-1 ) are the appearance feature vectors of the current frame and the previous frame extracted by the CNN, respectively, and |||2 is the L2 norm, loU(B t B t-1 The value represents the intersection-union ratio of the target bounding boxes in the two consecutive frames, which is dynamically adjusted according to the crowd density ρ.
4. The deep learning-based passenger flow statistics system according to claim 1, characterized in that, In the model building step, the feedback optimization module performs the following operations: When the tracking loss rate L track When the percentage is >15%, increase the kernel size of the last two layers of the CNN backbone by 50%. When the average displacement between consecutive frames D move When the pixel size is less than 5 pixels, the search window for optical flow is reduced from 32×32 to 16×16.
5. The deep learning-based passenger flow statistics system according to claim 1, characterized in that, The spatiotemporal passenger flow analysis module includes: Spatiotemporal graph construction unit: Constructs a spatiotemporal graph with pedestrian trajectory points as nodes and motion relationships as edges; Gated graph convolutional units: aggregate spatiotemporal neighborhood information and update node states through a gating mechanism; Spatiotemporal prediction unit: based on the final layer node state h i (L) The spatial-temporal attention mechanism is used to predict the passenger flow distribution after time T.
6. A deep learning-based passenger flow statistics method, applied to the system described in any one of claims 1-5, characterized in that, Including the following steps: S1. Synchronous acquisition and time alignment of visible light and thermal imaging video streams; S2. Extract multimodal features using a dual-channel CNN and fuse them according to the calculation formula of the feature-level fusion layer; S3. Initialize pedestrian detection boxes based on fused features, and perform cross-frame target association according to the appearance feature matching function; S4. Perform dynamic optimization of feature extraction and tracking parameters based on the feedback optimization module; S5. Generate a regional passenger flow heat map and future prediction curve as the passenger flow prediction result using the spatiotemporal passenger flow analysis module.
7. The deep learning-based passenger flow statistics method according to claim 6, characterized in that, In step S2, a transfer learning strategy is used for fusion: The visible light channel is loaded with COCO pre-trained weights, and the thermal imaging channel is loaded with FLIR pre-trained weights. The weights of the fusion layer are updated through end-to-end joint fine-tuning.
8. The deep learning-based passenger flow statistics method according to claim 6, characterized in that, Following the passenger flow statistics step, the passenger flow prediction results in S5 are used to trigger control commands: When the area density is greater than the safety density, a diversion alarm signal is generated; When the predicted flow exceeds the maximum capacity, a flow limiting control command is generated.
Citation Information
Patent Citations
Coal briquette area detection method for underground coal conveying image
CN110675443A
Three-dimensional full sound wave anomaly detection imaging method and device
CN113591876A