CNN-LSTM-based large-scale crowd evacuation time rapid prediction method

By combining CNN and LSTM methods, a crowd evacuation time prediction model is constructed, which solves the efficiency and accuracy of large-scale crowd evacuation time prediction, and realizes rapid prediction and real-time monitoring in complex scenarios, improving evacuation efficiency and safety.

CN120387548APending Publication Date: 2025-07-29JIANGSU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510602246.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

In the prediction of large-scale crowd evacuation time, it is difficult to achieve high efficiency and accuracy, especially in complex scenarios, it is impossible to quickly evaluate the degree of crowd congestion and risk level, which affects evacuation efficiency and safety.

Method used

Using a combination method based on convolutional neural network (CNN) and long and short-term memory network (LSTM), a population evacuation simulation model is constructed, spatial features are extracted and timing features are mined, and the mapping relationship between pedestrian distribution map and evacuation time is established to make rapid predictions.

Benefits of technology

It realizes rapid crowd evacuation time prediction in complex scenarios, improves the real-time and generalization capabilities of prediction, can monitor crowd dynamics in real time and provide early warnings, and improves evacuation efficiency and safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005397149470000021
    Figure BDA0005397149470000021
  • Figure BDA0005397149470000024
    Figure BDA0005397149470000024
  • Figure BDA0005397149470000026
    Figure BDA0005397149470000026
Patent Text Reader

Abstract

The invention discloses a CNN-LSTM-based large-scale crowd evacuation time rapid prediction method, and the method comprises the steps: building a crowd evacuation simulation model according to an evacuation place, simulating an evacuation process, and storing an evacuation video and evacuation time data; normalizing the data; constructing a time window, and extracting pedestrian distribution map sequence fragments from continuous time steps through a sliding window to form a data set; a large-scale crowd evacuation time rapid prediction model composed of a convolutional neural network, a recurrent neural network and a regression output layer is constructed, the convolutional neural network extracts spatial features, the recurrent neural network mines time sequence features, and the regression output layer establishes a mapping relation between a crowd distribution map and evacuation time and predicts the evacuation time; and utilizing the trained and optimized large-scale crowd evacuation time rapid prediction model to rapidly predict the large-scale crowd evacuation time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent transportation and computer simulation technology, and in particular to a method for quickly predicting large-scale crowd evacuation time based on convolutional neural networks (CNNs) and long short-term memory networks (LSTMs). Background Art

[0002] With the rapid development of urbanization and the improvement of people's living standards, large, high-density crowds often gather in public places such as train stations, subway stations, and stadiums. When uncontrollable emergencies occur, panic can easily ignite among the crowd, significantly reducing pedestrians' ability to perceive their surroundings. This can further exacerbate congestion and blockage during evacuation, and in severe cases, even threaten pedestrians' lives. Rapidly predicting the evacuation time of large crowds would help predict the severity of crowd congestion in advance, allowing for more efficient evacuation and safer pedestrians.

[0003] Currently, there are two main methods for assessing and predicting crowd evacuation times. The first involves traditional methods based on evacuation experiments and surveys. These methods obtain the required data through actual observations or simulated drills. For example, these methods analyze pedestrians' behavioral characteristics, such as movement speed and path selection, based on evacuation videos, or use questionnaires to understand crowd psychological characteristics (such as panic level and environmental familiarity). Evacuation drills can directly measure evacuation time and congestion, providing basic data for architectural design and model verification. The second method uses computer simulation models to obtain relevant parameters. For example, these methods utilize cellular automata, social force models, and multi-agent models to simulate crowd behavioral decisions and evacuation dynamics, generating key parameters such as evacuation time and path congestion.

[0004] However, the existing methods cannot achieve both high efficiency and accuracy in prediction at the same time, especially in scenarios with large populations, which puts higher demands on the realization of high efficiency and accuracy in prediction. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this application proposes a CNN-LSTM-based method for rapid prediction of large-scale crowd evacuation times. By combining a convolutional neural network (CNN) with a long short-term memory network (LSTM), this method enables rapid prediction of large-scale crowd evacuation times in complex scenarios, assesses crowd congestion levels and risk levels, and thus enables real-time monitoring and early warning of crowd dynamics. This method is suitable for areas with complex architectural environments and dense crowds, and has broad application value in scenarios such as emergency evacuations and major events.

[0006] The technical solution adopted by the present invention is as follows:

[0007] A fast prediction method for large-scale crowd evacuation time based on CNN-LSTM, comprising the following steps:

[0008] Step S1, build a crowd evacuation simulation model according to the evacuation site, use the crowd evacuation simulation model to simulate the evacuation process, and save the evacuation video and evacuation time data;

[0009] Step S2, perform normalization processing on the obtained data; construct a time window, and extract sequence segments of pedestrian distribution maps from continuous time steps through a sliding window to form a data set;

[0010] Step S3, build a fast prediction model for large-scale crowd evacuation time, which is composed of a convolutional neural network, a recurrent neural network and a regression output layer; the convolutional neural network extracts spatial features, the recurrent neural network mines temporal features, and the regression output layer establishes a mapping relationship between the crowd distribution map and the evacuation time to predict the evacuation time;

[0011] Step S4, use the data set obtained in Step S2 to train and optimize the fast prediction model for large-scale crowd evacuation time;

[0012] Step S5, use the fast prediction model for large-scale crowd evacuation time to quickly predict the large-scale crowd evacuation time.

[0013] Furthermore, the method for the convolutional neural network to extract the spatial features of the sequence segments of the pedestrian distribution map is as follows:

[0014] Take the sequence segments of the pedestrian distribution map as the input, use convolutional layers and pooling layers at the front end of the convolutional neural network to gradually compress the spatial dimension; use global average pooling at the end to compress the feature map into a vector; perform normalization processing on each batch of data along the channel dimension.

[0015] Furthermore, the recurrent neural network adopts a multi-layer stacked long short-term memory network LSTM and Dropout processing.

[0016] Furthermore, stack two layers of LSTM, denoted as:

[0017]

[0018] Wherein, respectively represent the hidden state outputs of the first layer and the second layer of LSTM, respectively represent the cell states of the first layer and the second layer of LSTM.

[0019] Furthermore, add Dropout between the two layers of LSTM to randomly mask neurons to prevent overfitting, denoted as:

[0020]

[0021] Among them, represents the hidden state output of the second-layer LSTM.

[0022] Furthermore, the regression output layer is denoted as:

[0023]

[0024] where h L represents the hidden state at the last time step of the LSTM; W reg represents the trainable parameter; b reg represents the bias term for the regression task, which is used to adjust the output baseline of the model.

[0025] Furthermore, the mean squared error is selected as the loss function for the large-scale crowd evacuation time rapid prediction model.

[0026] Furthermore, if there is a long tail in the prediction error, the Huber Loss is used as the loss function for the large-scale crowd evacuation time rapid prediction model.

[0027] Furthermore, through hyperparameter setting and training strategy selection, the large-scale crowd evacuation time rapid prediction model is trained and optimized, including the following steps:

[0028] Step S4.1, set hyperparameters, including the optimizer and regularization;

[0029] Step S4.2, set the training strategy, including the early stopping mechanism to prevent the model from overfitting on the training set and dynamically adjusting the learning rate.

[0030] Furthermore, the normalization process in Step S2 is to convert features with different dimensions and large distribution ranges into a unified standard range.

[0031] Advantages of the present invention:

[0032] (1) Enhanced real-time performance

[0033] The prior art has a slow prediction speed for the large-scale crowd evacuation time and cannot ensure safety in reality. The present invention combines a convolutional neural network (CNN) with a long short-term memory network (LSTM). The convolutional neural network (CNN) extracts spatial features of the pedestrian distribution change, and the long short-term memory network (LSTM) mines the temporal features of the pedestrian distribution change. Finally, a mapping relationship between the pedestrian distribution map sequence segments and the crowd evacuation time is established through the regression output layer, enabling the rapid generation of the evacuation time from the pedestrian distribution map and ensuring real-time performance in the evacuation time prediction.

[0034] (2) Improved generalization ability

[0035] The present invention uses computer simulation software to simulate a large number of heterogeneous scenarios and simultaneously considers different population behaviors, thereby obtaining a more comprehensive and detailed dataset. After training with the model in the present invention, the model has stronger recognition ability for video sequences under different scenarios and different pedestrian distributions, improving the generalization ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 It is a flowchart of a method for quickly predicting large-scale crowd evacuation time based on CNN-LSTM of the present invention.

[0037] Figure 2 It is a schematic diagram of the network structure of CNN-LSTM of the present invention.

[0038] Figure 3 It is a structural diagram of the convolutional neural network (CNN) of the present invention.

[0039] Figure 4 It is a structural diagram of the long short-term memory network (LSTM) of the present invention.

[0040] Figure 5 It is a structural diagram of the regression output layer of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0042] As Figure 1 、 Figure 2 、 Figure 3 、 Figure 4 and Figure 5 shown, a method for quickly predicting large-scale crowd evacuation time based on CNN-LSTM includes the following steps:

[0043] Step S1, with the aid of computer simulation software, investigate the environmental information of relevant actual evacuation sites, construct a crowd evacuation simulation model, use the crowd evacuation simulation model to simulate the evacuation process, save data such as evacuation videos and evacuation times, and verify the obtained data to eliminate some unreasonable data, and finally obtain the dataset required for model training.

[0044] Step S1.1, simulate the scenario, define the topological connection relationship by investigating the geometric structure of relevant buildings; at the same time, considering crowd heterogeneity, set the crowd according to categories such as individual attributes and behavior models.

[0045] Step S1.1.1: Construct a virtual evacuation scenario for the simulation model based on the actual investigated evacuation scenario (such as a stadium). The evacuation scenario model mainly includes building walls, internal obstacles such as desks and chairs, and exits at different positions (including doorways with different widths and sizes).

[0046] Step S1.1.2: After the virtual scenario is constructed, divide pedestrians into different categories according to their attributes, including different ages, moving speeds, and position attributes. After division, randomly add pedestrians in the virtual evacuation scenario to ensure the diversity of the simulation. Among them, use the simulation evacuation model based on social force to calculate the movement speed vector of pedestrians, which takes into account the driving force, the interaction force between pedestrians, and the influence of environmental obstacles. The specific formula is as follows:

[0047]

[0048] Where m i is the mass of the i-th pedestrian, v i is the speed of the i-th pedestrian, t is time, represents the driving force of the pedestrian towards the target point, and the formula is is the expected speed of the i-th pedestrian, τ i is the adjustment time parameter; ∑F ij represents the interaction force between pedestrians, including psychological repulsive force and physical contact force, and the formula is g(r ij -d ij )n ij , A is the amplitude constant of several terms, r ij is the actual distance between different objects, d ij is the equilibrium distance, B is the attenuation length, n ij is the unit vector, k is the elastic coefficient, g is the parameter of the linear function; ∑F iW represents the obstacle force, that is, the force of the wall, etc., W represents the boundary (such as the wall); F f represents the force between the pedestrian and other elements in the environment.

[0049] Step S1.1.3: Combine the above Step S1.1.1 and Step S1.1.2 to construct a complete crowd evacuation simulation model, and simulate the evacuation process, saving data such as evacuation videos and evacuation times.

[0050] Step S1.2: Data recording specification, process the data obtained from the simulation to make it standardized.

[0051] Step S1.2.1, Grid the population density map. First, determine the grid resolution, which is selected according to the scene size (e.g., H×W = 40×60) to ensure that a single grid can accommodate 3 to 5 people. Further, calculate the density. The density calculation formula is as follows:

[0052]

[0053] where h and w respectively represent the row index and column index of the grid.

[0054] Step S1.3, Data verification and cleaning. There may be certain errors in the obtained data, and it is necessary to check the rationality of the data. For example, verify whether the evacuation time conforms to the empirical formula:

[0055]

[0056] where T est is the evacuation time, N is the total number of evacuees, W cxit is the effective width of the exit, and q max is the maximum flow per unit width.

[0057] At the same time, conduct a logical detection of the trajectory to detect whether problems such as an evacuation individual penetrating an obstacle or a sudden change in speed occur.

[0058] Step S1.4, Data scale and storage. Since this simulation is to provide a dataset for subsequent model training, a large number of simulations are required to meet the requirements of model training and testing, and this simulation needs to cover more than 90% of the parameter space.

[0059] In terms of data storage, group according to the scene type, store each simulation as an independent dataset, and at the same time, data compression can also be performed to reduce storage overhead.

[0060] Step S2, Preprocess the data obtained in the previous step, mainly through methods such as data normalization and constructing a time window to improve the quality of the data and optimize the calculation efficiency. It includes:

[0061] Step S2.1, Data normalization. It is the process of converting features with different dimensions and large distribution ranges into a unified standard range to achieve goals such as accelerating model convergence, improving model accuracy, and enhancing model generalization ability. It mainly includes the following two aspects:

[0062] Step S2.1.1, Heat map normalization. It is a special treatment for the heat map. For the grid density, select frame-by-frame Min-Max normalization. The formula is as follows:

[0063]

[0064] where is the normalized density, max(D t ) is the maximum density among all grids, ∈ is an adjustment parameter to prevent the denominator from being zero, and its value is taken as 1e-5.

[0065] Step S2.1.2, Feature Fusion. When splicing the heat map (spatial main feature) and auxiliary features (such as exit flow rate, average speed, etc.) into a multi-channel tensor, it is necessary to ensure that the normalization methods of each channel are consistent:

[0066] (1) Heat map channel: Min-Max to [0,1].

[0067] (2) Velocity field channel: Z-Score normalization, and its formula is:

[0068]

[0069] Among them, x represents the observed value of the original data point, μ represents the mean of the population to which the data belongs, and σ represents the standard deviation of the population to which the data belongs.

[0070] (3) Distance matrix channel: After Log transformation, Min-Max to [0,1], where the Log transformation formula is as follows:

[0071] s = c·log b (1 + r)

[0072] Among them, r represents the original input value; s represents the transformed output value; c represents the contrast adjustment constant used to control the range of the output value; represents the base of the logarithm.

[0073] Step S2.2, Complete the construction of the time window through sliding window sampling and dataset partitioning.

[0074] Step S2.2.1, Sliding window sampling. The sliding window is used to extract local sequences from consecutive time steps as model inputs and associate the corresponding output labels. Its core parameters include:

[0075] (1) Window length (L): The number of time steps of the input sequence (historical observation length), and the autocorrelation function is selected as the statistical basis for determining the window length. Its formula is:

[0076]

[0077] Among them, is the expected value, X t is the observed value of the time series, μ is the mean of the time series, and σ is the variance of the time series.

[0078] (2) Step size (S): The interval number of steps for the window to slide (controlling the overlap degree between samples).

[0079] (3) Prediction offset (H): The time offset of the output label relative to the end of the input window (H > 1 for multi-step prediction).

[0080] Step S2.2.1.1, formulating the input-output relationship. For time series data {X1, X2,..., X T} and labels {y1, y2,..., y T}, where T is the total number of time points, the generation rule for each sample is:

[0081] Input: x i = [X i , X i+1 ,..., X i+L-1 (Shape: L × C × H × W)

[0082] Output: y i = y i+L-1+H (H = 0 for single-step prediction)

[0083] where i is the starting time of the window, and it needs to satisfy i + L - 1 + H ≤ T.

[0084] Step S2.2.2, dataset division. Divide the dataset into a training set, a validation set, and a test set in a ratio of 7:2:1, while maintaining time continuity.

[0085] Step S3, design a large-scale crowd evacuation time rapid prediction model based on CNN-LSTM. The convolutional neural network (CNN) is used to extract spatial features, the recurrent neural network is used to mine temporal features, and the regression output layer is used to establish the mapping relationship between the crowd distribution map and the evacuation time to predict the evacuation time. It includes:

[0086] Step S3.1, convolutional neural network (CNN), which is mainly used to extract the spatial features of the pedestrian distribution map sequence segments.

[0087] Step S3.1.1, convolutional layer. For the L-th layer convolutional operation, the input tensor Output feature map The specific calculation formula is:

[0088]

[0089] where K h ×K w represents the convolutional kernel size, which determines the receptive field size; S represents the stride, which controls the output size reduction rate; C in , C out represent the number of input and output channels respectively, that is, the number of different feature detectors; H in , H outrespectively represent the height sizes of the input and output; W in , W out respectively represent the width sizes of the input and output; is the weight parameter of the convolutional kernel; is the value of the channel at different positions in the feature map; is the bias term of the output channel; σ represents the activation function, introducing non-linearity.

[0090] Step S3.1.2, the pooling layer, whose main objective is to reduce the spatial resolution, enhance translational invariance, and reduce the computational amount. The specific calculation formula is as follows:

[0091]

[0092] Among them, is the feature map after pooling processing, is the input of the pooling layer, is the pooling window centered at (h, w).

[0093] In the design selection, use small-step convolution + pooling at the front end of the CNN to gradually compress the spatial dimension; use global average pooling at the end to replace the fully connected layer and compress the feature map into a vector. The relevant formula is as follows:

[0094]

[0095] Among them, v is the output of the global average pooling, is the output vector.

[0096] Step S3.1.3, batch normalization, that is, normalize each batch of data along the channel dimension to alleviate the internal variable shift, accelerate the training convergence, and at the same time allow the use of a larger learning rate to improve the model generalization ability. The specific formula is as follows:

[0097]

[0098] Among them, is the original value of the input feature map, is the value after normalization, is the final output, μ c , σ c respectively represent the mean and standard deviation of the c-th channel within the batch; γ c , β c respectively represent the learnable scaling and translation parameters.

[0099] Step S3.2, the recurrent neural network is mainly used to mine the temporal features of the pedestrian distribution map sequence segments. In this embodiment, the recurrent neural network adopts the long short-term memory network (LSTM).

[0100] Step S3.2.1, Temporal Modeling. For the input feature vector at time step t and the hidden state h at the previous moment t-1 , the relevant calculation formulas of LSTM are as follows:

[0101] Forget Gate: f t = σ(W f ·[h t-1 , v t + b f )

[0102] Input Gate: i t = σ(W i ·[h t-1 , v t + b i )

[0103] Candidate Cell State:

[0104] Cell State Update:

[0105] Output Gate: o t = σ(W o ·[h t-1 , v t + b o )

[0106] Hidden State Output: h t = o t ⊙ tanh(C t )

[0107] Among them, σ represents the Sigmoid function, which compresses the gating signal to control the information flow ratio; tanh represents the hyperbolic tangent function, which adjusts the value range of the candidate state to. W f , W i , W C , W o are the weights of the forget gate, input gate, candidate cell state, and output gate respectively; b f , b i , b C , b o are the bias vectors of the forget gate, input gate, candidate cell state, and output gate respectively.

[0108] Step S3.2.2, Multi-layer Stacking and Dropout Processing. By stacking two layers of LSTM, the higher layer can learn more abstract temporal patterns, and its extended formula is:

[0109]

[0110] Among them, respectively represent the hidden state outputs of the first and second layers of LSTM, respectively represent the cell states of the first and second layers of LSTM.

[0111] Meanwhile, Dropout is added between the LSTM layers to randomly mask neurons and prevent overfitting. The relevant formula is:

[0112]

[0113] Step S3.3, the regression output layer, is mainly used to establish the mapping relationship between the pedestrian distribution map sequence segment and the evacuation time, and predict the crowd evacuation time. Its formula is:

[0114]

[0115] Among them, represents the hidden state at the last time step of LSTM; represents the trainable parameter. b reg is the bias term.

[0116] Meanwhile, the mean squared error (MSE) is selected as the loss function of the large-scale crowd evacuation time rapid prediction model:

[0117]

[0118] Among them, N is the total number of samples, is the evacuation time of each sample.

[0119] If there is a long tail in the prediction error, the loss function is replaced with Huber Loss to enhance robustness. Its formula is:

[0120]

[0121] Among them, δ is the threshold parameter that determines the switching point between the squared loss and the linear loss; e represents the error between the predicted value and the true value.

[0122] Step S4, after establishing the dataset and the prediction model, start training and optimizing the model. By setting hyperparameters and selecting appropriate training strategies, make the prediction model have higher data processing performance and generalization ability. Including:

[0123] Step S4.1, hyperparameter setting. Hyperparameters directly affect the architecture design of the model and determine the learning ability and expression ability of the model. It is mainly processed from two aspects: the optimizer and regularization.

[0124] Step S4.1.1, Optimizer, which is used to adjust model parameters, minimize the loss function, and the Adam optimizer, i.e., Adaptive Moment Estimation, is selected. Its advantage lies in the adaptive adjustment of the learning rate and is suitable for sparse gradient problems (such as dynamic changes in crowd evacuation prediction). The calculation formula is as follows:

[0125] m t = β1m t-1 + (1 - β1)g t

[0126]

[0127] where m t is the first moment estimate at the current moment, is the corrected first moment estimate, g t is the gradient at the current moment, v t is the second moment estimate at the current moment, is the corrected second moment estimate, θ t is the parameter before update, β1 = 0.9, β2 = 0.999 represent the momentum decay rate; η = 10 -3 represents the initial learning rate; ∈ = 10 -8 represents the numerical stability constant.

[0128] Step S4.1.2, Regularization, which is mainly used to prevent the model from overfitting and improve the generalization ability. In the present invention, Dropout regularization is adopted. This method shows significant advantages in preventing overfitting, enhancing robustness, and optimizing training efficiency through randomness and the ensemble effect. Its training formula is:

[0129] m ~ Bernoulli(1 - p)

[0130] where h drop is the regularization result of the input feature, h is the original input feature, m is the Bernoulli mask, and p = 0.5 represents the neuron dropout probability.

[0131] Step S4.2, Training strategy, including:

[0132] Step S4.2.1, Early stopping mechanism, which can prevent the model from overfitting on the training set. The implementation method of this mechanism is as follows:

[0133] (1) Monitoring metric: Validation set loss

[0134] (2) Stopping condition: If the validation loss does not decrease for P = 10 consecutive epochs, stop training.

[0135] (3) Model saving: Retain the model parameters with the lowest validation loss.

[0136] The mathematical basis of this mechanism is as follows: Assume that the validation loss sequence is If it satisfies:

[0137]

[0138] Then early stopping is triggered.

[0139] Step S4.2.2, learning rate scheduling, is used to dynamically adjust the learning rate to balance the convergence speed and stability. The method used in the present invention is cosine annealing. This method significantly improves the model training efficiency and final performance through dynamically balancing exploration and convergence, the periodic restart mechanism, and the smooth transition characteristic. Its advantages are particularly prominent in complex tasks and large-scale data scenarios. The formula is as follows:

[0140]

[0141] Among them, η t is the learning rate at different time steps, T max = 20 represents the cycle length; η max = 10 -3 , η min = 10 -5 represent the upper bound and lower bound of the learning rate respectively.

[0142] Finally, the performance of the model is evaluated. Quantitatively, MAE, RMSE, etc. are selected as evaluation indicators; at the same time, a time comparison curve and an error heat map are plotted for qualitative analysis of the model. Including:

[0143] 1. Quantitative evaluation, that is, measuring the prediction accuracy of the model through mathematical indicators. The following indicators are used in the present invention:

[0144] (1) Root Mean Square Error (RMSE), which is mainly used to reflect the absolute magnitude of the prediction error and is sensitive to outliers (such as the scene of sudden change in evacuation time). The formula is as follows:

[0145]

[0146] (2) Mean Absolute Error (MAE), which is used to directly measure the absolute average of the error and is more robust than RMSE. The formula is as follows:

[0147]

[0148] (3) Coefficient of Determination (R 2 -Squared), which represents the ability of the model to explain the variation of the data. The value range is (-∞, 1]. The closer it is to 1, the better the fitting. The formula is as follows:

[0149]

[0150] (4) Quantile error, which is used to evaluate the reliability of the model in extreme situations (such as a sharp increase in evacuation time caused by a fire), and its formula is as follows:

[0151]

[0152] Among them, q ∈ (0, 1) is the quantile.

[0153] 2. Qualitative analysis, that is, revealing the spatio-temporal patterns of the model behavior through visualization means, including:

[0154] S1. Time comparison curve. By plotting the curves of the true value and the predicted value changing with time, the trend consistency is visually compared. The implementation steps are as follows:

[0155] (1) Data sampling: Select the complete time series of a certain evacuation scenario in the test set and its predicted sequence

[0156] (2) Plotting rules: Use the simulation time as the horizontal axis and the remaining evacuation time as the vertical axis. In terms of curve annotation, distinguish them in the following way: true value (blue), predicted value (red), key event markers (such as the start time of exit congestion).

[0157] In the analysis of this part, two aspects of lag effect and mutation capture should be considered, that is, respectively judge whether the predicted value continuously lags behind the true value (insufficient model response speed); when the evacuation bottleneck is formed, whether the predicted value can timely reflect the time jump.

[0158] S2. Heat map error, visualizing the distribution of prediction errors on the spatial grid, used to locate the model blind spots to guide the improvement of the model. The implementation steps are as follows:

[0159] (1) Error calculation: For each grid (h, w), calculate its mean absolute error:

[0160] (2) Heat map generation:

[0161] Color mapping: Use a gradient color (such as blue → red) to represent the error from low to high.

[0162] Overlay scene layout: Draw the positions of exits and obstacles on the map, associating the error distribution with the scene structure.

[0163] In the analysis of this part, two aspects of high error regions and dynamic patterns should be considered, that is, respectively judge whether they are concentrated in key positions such as near exits and corners (insufficient model spatial perception); whether the error spreads from the center to the exit over time (defect in time series modeling).

[0164] Step S5: Use the large-scale crowd evacuation time rapid prediction model to rapidly predict the large-scale crowd evacuation time.

[0165] The above embodiments are only used to illustrate the design concept and characteristics of the present invention, and the purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. The protection scope of the present invention is not limited to the above embodiments. Therefore, all equivalent changes or modifications made based on the principles and design concepts disclosed by the present invention are within the protection scope of the present invention.

Claims

1. A rapid prediction method for large-scale crowd evacuation time based on CNN-LSTM, characterized in that, It includes the following steps: Step S1: Build a crowd evacuation simulation model according to the evacuation site, use the crowd evacuation simulation model to simulate the evacuation process, and save the evacuation video and evacuation time data; Step S2: Normalize the obtained data; construct a time window, and extract sequence segments of pedestrian distribution maps from continuous time steps through a sliding window to form a data set; Step S3: Build a large-scale crowd evacuation time rapid prediction model, which consists of a convolutional neural network, a recurrent neural network, and a regression output layer; the convolutional neural network extracts spatial features, the recurrent neural network mines temporal features, and the regression output layer establishes a mapping relationship between the crowd distribution map and the evacuation time to predict the evacuation time; Step S4: Use the data set obtained in Step S2 to train and optimize the large-scale crowd evacuation time rapid prediction model; Step S5: Use the large-scale crowd evacuation time rapid prediction model to rapidly predict the large-scale crowd evacuation time.

2. The rapid prediction method for large-scale crowd evacuation time based on CNN-LSTM according to claim 1, wherein The method for the convolutional neural network to extract the spatial features of the sequence segments of the pedestrian distribution map is as follows: Take the sequence segments of the pedestrian distribution map as the input, use convolutional layers and pooling layers at the front end of the convolutional neural network to gradually compress the spatial dimension; use global average pooling at the end to compress the feature map into a vector; perform normalization processing on each batch of data along the channel dimension.

3. A rapid prediction method for large-scale crowd evacuation time based on CNN-LSTM according to claim 1, characterized in that The recurrent neural network adopts multi-layer stacked long short-term memory network LSTM and Dropout processing.

4. A rapid prediction method for large-scale crowd evacuation time based on CNN-LSTM according to claim 3, characterized in that Stack two layers of LSTM, denoted as: Among them, respectively represent the hidden state outputs of the first and second layers of LSTM, respectively represent the cell states of the first and second layers of LSTM.

5. A method for quickly predicting the evacuation time of a large-scale crowd based on CNN-LSTM according to claim 4, characterized in that, Add Dropout between the two layers of LSTM to randomly mask neurons to prevent overfitting, denoted as: Among them, represents the hidden state output of the second-layer LSTM.

6. A rapid prediction method for large-scale crowd evacuation time based on CNN-LSTM according to claim 1, characterized in that The regression output layer is denoted as: Among them, h L represents the hidden state of the last time step of the LSTM; W reg represents the trainable parameter; b reg represents the bias term for the regression task, which is used to adjust the output baseline of the model.

7. A rapid prediction method for large-scale crowd evacuation time based on CNN-LSTM according to claim 1, characterized in that Select mean squared error as the loss function of the large-scale crowd evacuation time rapid prediction model.

8. A rapid prediction method for large-scale crowd evacuation time based on CNN-LSTM according to claim 1, characterized in that If there is a long tail in the prediction error, use Huber Loss as the loss function of the large-scale crowd evacuation time rapid prediction model.

9. A rapid prediction method for large-scale crowd evacuation time based on CNN-LSTM according to claim 1, characterized in that, Through hyperparameter setting and training strategy selection, train and optimize the large-scale crowd evacuation time rapid prediction model, including the following steps: Step S4.1: Set hyperparameters, including optimizers and regularization; Step S4.2: Set training strategies, including an early stopping mechanism to prevent the model from overfitting on the training set and dynamically adjusting the learning rate.

10. A method for rapidly predicting the evacuation time of a large-scale crowd based on CNN-LSTM according to claim 1, characterized in that, The normalization processing in Step S2 is to convert features with different dimensions and large distribution ranges into a unified standard range.