Passenger flow prediction method based on mobile phone signaling and subway turnstile data association modeling
By dynamically preprocessing and multi-dimensional feature encoding the subway gate and mobile phone signaling data, combined with deep learning models and historical baselines, the problem of insufficient perception in the prediction of passenger flow for major events in existing technologies has been solved, and accurate passenger flow prediction and proactive operation intervention have been achieved.
Patent Information
- Application Number
- CN202511960234.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-24
- Publication Date
- 2026-01-23
AI Technical Summary
Existing subway passenger flow forecasting methods lack the ability to perceive the macro-level flow of people in major event areas, fail to effectively integrate mobile phone signaling data, struggle to capture the complete passenger flow evolution chain from surrounding gatherings to station influx, and fail to construct a business closed loop of forecasting-identification-control, resulting in operational delays.
By dynamically preprocessing real-time signaling data, gate data, and activity dynamic information, a real-time standard input snapshot is constructed. Multidimensional spatiotemporal feature encoding is used to drive a deep learning model. Anomaly identification and risk assessment are performed in combination with historical baselines and site capacity to generate a closed-loop management and control execution plan.
It enables accurate prediction and proactive intervention of passenger flow during major events, breaking through the limitations of perception from a single data perspective, and improving the accuracy of prediction of explosive passenger flow and the initiative of operation.
Smart Images

Figure CN121390477A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of big data and artificial intelligence technology, and more specifically, to a passenger flow prediction method based on the correlation modeling of mobile phone signaling and subway gate data. Background Technology
[0002] Against the backdrop of rapid development in urban rail transit, subways have become a core mode of transportation for carrying passenger flows to major events such as sporting events, expos, and performances. These passenger flows are typically characterized by high bursts of activity, extremely concentrated spatiotemporal distribution, and high uncertainty. Without accurate advance prediction and effective capacity management, severe platform congestion and even stampedes can easily occur. Currently, industry research on subway passenger flow forecasting largely focuses on using historical automatic fare collection system data combined with machine learning algorithms for modeling. Some existing technologies utilize dynamic graph clustering and graph trend network models to handle passenger flow forecasting during holidays. However, when facing the specific high-pressure scenario of major events, existing forecasting schemes based on single data sources or general models show significant shortcomings.
[0003] Specifically, traditional methods rely heavily on data from subway turnstiles, lacking the ability to perceive the macro-level flow of people around event venues. They fail to effectively integrate external data, such as mobile phone signaling, which reflects pre-gathering trends, resulting in an inability to capture the complete passenger flow evolution chain from surrounding gatherings to station influx. Furthermore, existing models are primarily designed for regular daily commuter traffic, failing to fully incorporate the attributes of major events. This makes it difficult to effectively distinguish the essential differences between daily passenger flow and sudden event-related surges, leading to low sensitivity to the specific temporal pattern of "pre-gathering period – outbreak period – dispersal period," failing to meet the practical need for accurate predictions several hours in advance. In addition, existing passenger flow anomaly identification methods often treat stations as isolated analytical units, calculating deviations only by longitudinally comparing them to their own historical baselines, completely ignoring the spatial topological relationships inherent in the subway network system. This results in a failure to integrate network structure information into the core anomaly detection algorithm. Finally, existing solutions often only output predicted values, failing to construct a "prediction-identification-control" business loop, causing operational responses to often be reactive and delayed.
[0004] Therefore, there is an urgent need for a passenger flow prediction method based on the correlation modeling of mobile phone signaling and subway gate data. Summary of the Invention
[0005] This application is made in order to solve the above-mentioned technical problems.
[0006] According to one aspect of this application, a passenger flow prediction method based on mobile phone signaling and subway gate data association modeling is provided, which includes: Dynamic preprocessing is performed on real-time signaling data streams, real-time gate data streams, and real-time activity dynamic information to obtain a real-time standard input snapshot; Feature encoding is performed on real-time standard input snapshots to obtain real-time inferred feature vectors; The real-time inferred feature vectors are input into the trained passenger flow prediction model to obtain the future passenger flow prediction sequence. Based on historical baseline data and station capacity parameters, passenger flow anomalies and risks are identified and assessed in future passenger flow forecast sequences to obtain the spatiotemporal risk level distribution. Based on the knowledge base of control strategies, hierarchical control decisions are made on the spatiotemporal risk level distribution to obtain a closed-loop control execution plan.
[0007] Compared with existing technologies, this application provides a passenger flow prediction method based on the correlation modeling of mobile phone signaling and subway gate data. First, it collects mobile phone signaling data covering macro-level passenger flow, micro-level gate data, and information on major events in real time, then dynamically cleans and aligns the data spatiotemporally to construct a panoramic standard input. Subsequently, it uses multi-dimensional spatiotemporal feature encoding to drive a deep learning model, analyzing the temporal evolution patterns during the event and predicting future passenger flow sequences. Next, it combines historical baselines and station capacity to calculate anomaly deviations and determine risks, identifying spatiotemporal risk levels. Finally, it matches decisions based on a management knowledge base to generate a closed-loop execution plan including capacity scheduling and flow control measures. This effectively overcomes the limitations of perception from a single data perspective and the temporal lag of conventional models, thereby achieving accurate prediction and proactive intervention of the entire passenger flow chain during major events. Attached Figure Description
[0008] The above and other objects, features, and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the embodiments of this application to explain this application and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.
[0009] Figure 1 This is a flowchart of a passenger flow prediction method based on the association modeling of mobile phone signaling and subway gate data according to an embodiment of this application.
[0010] Figure 2 This is a data flow diagram of a passenger flow prediction method based on mobile phone signaling and subway gate data association modeling according to an embodiment of this application.
[0011] Figure 3 This is a flowchart of sub-step S1 of the passenger flow prediction method based on the association modeling of mobile phone signaling and subway gate data according to an embodiment of this application.
[0012] Figure 4This is a flowchart of sub-step S2 of the passenger flow prediction method based on the association modeling of mobile phone signaling and subway gate data according to an embodiment of this application.
[0013] Figure 5 This is a flowchart of sub-step S4 of the passenger flow prediction method based on the association modeling of mobile phone signaling and subway gate data according to an embodiment of this application.
[0014] Figure 6 This is a flowchart of sub-step S41 of the passenger flow prediction method based on the association modeling of mobile phone signaling and subway gate data according to an embodiment of this application. Detailed Implementation
[0015] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0016] To address the problems mentioned above, this application proposes a passenger flow prediction method based on the correlation modeling of mobile phone signaling and subway gate data. Figure 1 This is a flowchart of a passenger flow prediction method based on the association modeling of mobile phone signaling and subway gate data according to an embodiment of this application. Figure 2 This is a data flow diagram for a passenger flow prediction method based on mobile phone signaling and subway gate data association modeling, according to an embodiment of this application. Figure 1 and Figure 2 As shown, the passenger flow prediction method based on the correlation modeling of mobile phone signaling and subway gate data includes the following steps: S1, dynamically preprocessing the real-time signaling data stream, real-time gate data stream, and real-time activity dynamic information to obtain a real-time standard input snapshot; S2, performing feature encoding on the real-time standard input snapshot to obtain a real-time inferred feature vector; S3, inputting the real-time inferred feature vector into the trained passenger flow prediction model to obtain a future passenger flow prediction sequence; S4, based on the historical baseline database and station capacity parameters, performing passenger flow anomaly identification and risk assessment on the future passenger flow prediction sequence to obtain a spatiotemporal risk level distribution; S5, based on the control strategy knowledge base, performing hierarchical control decisions on the spatiotemporal risk level distribution to obtain a closed-loop control execution plan.
[0017] In the aforementioned passenger flow prediction method based on the correlation modeling of mobile signaling and subway gate data, step S1 involves dynamically preprocessing the real-time signaling data stream, real-time gate data stream, and real-time activity dynamic information to obtain a real-time standard input snapshot. Because the real-time signaling data stream and real-time gate data stream originate from different sources, the data exhibits inherent heterogeneity in terms of time granularity, spatial coordinate system, and data format. Furthermore, real-time activity dynamic information (such as delays and actual number of entrants) is dynamically changing. Directly inputting these multi-source, heterogeneous, and unaligned data into the model would make it difficult for the model to capture the correlation features between different data, thus affecting the accuracy of the prediction. Therefore, in the technical solution of this application, the real-time signaling data stream, real-time gate data stream, and real-time activity dynamic information are dynamically preprocessed to obtain a real-time standard input snapshot, thereby eliminating the heterogeneity of multi-source data and ensuring high consistency of data in the spatiotemporal dimensions. In this way, a structured data slice containing the current passenger flow at each station's gates, the distribution of associated signaling, and the latest activity status can be constructed, providing a high-quality and standardized input foundation for subsequent feature encoding and model inference. This effectively avoids prediction bias caused by data quality issues and improves the model's ability to perceive sudden changes in passenger flow during major events.
[0018] Specifically, the real-time signaling data stream, real-time gate data stream, and real-time activity dynamic information collection process are as follows: First, a high-concurrency, multi-source heterogeneous data access layer is constructed at the bottom layer. Three independent data receiving channels are opened in parallel to connect to the signaling data interface of the telecom operator, the AFC system interface of the subway clearing center, and the dynamic information release interface of the activity venue management, respectively. For the real-time signaling data stream, the signaling logs generated by the interaction between the user's mobile terminal and the base station are continuously captured and parsed to extract the encrypted user unique identifier IMSI, interaction timestamp, and base station location code LAC and CI information; for ... signaling logs generated by the interaction between the user's mobile terminal and the base station management center. The real-time gate data stream monitors the entry and exit transaction records of all stations across the entire subway network, capturing the card number hash value, the card swiping time accurate to the second, the transaction type, and the station ID of each transaction. For real-time event dynamic information, the system obtains the current actual number of attendees, performance delay status notifications, and crowd heat data of core areas uploaded by the venue through API polling or event push mechanisms. Subsequently, the above three high-frequency data streams are merged into a unified message queue middleware for caching, and the continuous data streams are sliced according to a preset 15-minute time granularity and packaged into raw data blocks to be processed.
[0019] In particular, in one specific embodiment, Figure 3 This is a flowchart of sub-step S1 of the passenger flow prediction method based on the association modeling of mobile phone signaling and subway gate data according to an embodiment of this application. Figure 3As shown, step S1 includes: S11, performing anomaly cleaning on the real-time signaling data stream and the real-time gate data stream to obtain cleaned signaling data blocks and cleaned gate data blocks; S12, dynamically updating and marking the real-time activity dynamic information based on a preset static activity information database to obtain a dynamically updated activity state and stage marker set; S13, aligning the cleaned signaling data blocks, the cleaned gate data blocks, and the dynamically updated activity state and stage marker set in a spatiotemporal dimension based on a base station-site spatial mapping table to obtain a real-time standard input snapshot.
[0020] Specifically, step S11 involves anomaly cleaning of the real-time signaling data stream and the real-time gate data stream to obtain cleaned signaling data blocks and cleaned gate data blocks. It should be understood that real-time collected signaling data may contain issues such as missing base station location information and format errors, while the gate data stream may exhibit abnormal transaction times and invalid site IDs. These invalid data can severely interfere with the accuracy of subsequent modeling. Therefore, this application performs anomaly cleaning on the real-time signaling data stream and the real-time gate data stream, eliminating invalid records through multi-dimensional verification to ensure the accuracy and integrity of the data. This yields clean signaling data blocks and gate data blocks, providing a reliable data foundation for subsequent spatiotemporal correlation feature calculations and avoiding prediction biases caused by outliers.
[0021] Specifically, in one possible embodiment, step S11 is implemented as follows: For real-time signaling data streams, key fields such as user identifier, base station location, and timestamp are verified for each record, and records with empty fields or non-compliant formats are removed. For real-time gate data streams, the reasonableness of transaction time, the validity of site ID, and the logic of transaction type are checked, and transaction records that exceed normal operating hours, have invalid site codes, or are duplicates are deleted. After cleaning, the valid data is divided into blocks according to time windows and stored to form well-structured cleaned data blocks.
[0022] Specifically, step S12 involves dynamically updating and marking real-time event dynamic information based on a pre-set static event information database to obtain a dynamically updated event state and stage marker set. It should be understood that due to the high degree of uncertainty in the actual execution of major events, situations such as concerts being delayed, sports events having extended end times due to overtime, or actual attendance not matching expectations frequently occur. Relying solely on a pre-fixed static schedule for feature construction can lead to severe timing biases in the model's judgment of peak arrival times. For example, in the event delay, the model might still predict the departure peak based on the original planned time, resulting in false alarms or missed alarms. Therefore, in the technical solution of this application, the real-time event dynamic information is further dynamically updated and marked based on a pre-set static event information database to obtain a dynamically updated event state and stage marker set. This establishes a dynamic benchmark mechanism capable of real-time perception and correction of the actual event progress, ensuring that the time characteristics input to the model are strictly synchronized with the actual event progress in the physical world. This enables subsequent feature encoding steps to accurately capture the true start and end times of the three key passenger flow evolution stages: the pre-gathering period, the burst period, and the evacuation period. This ensures that the neural network model activates the prediction weights for burst passenger flow within the correct time window, thereby significantly improving the robustness and timeliness of passenger flow prediction in scenarios where there are temporary changes in the event time.
[0023] Specifically, in one possible embodiment, step S12 is implemented as follows: First, a preset static activity information database is initialized and loaded. This database stores basic data such as the planned start time, planned end time, venue location coordinates, and expected number of participants for large-scale events. Then, the system continuously monitors the incoming real-time activity dynamic information stream, which includes real-time notifications of activity status changes pushed by the event organizer or venue, such as delayed start times, encore performances leading to delayed end times, or dynamic statistics of actual attendance. When a status change notification is received, the corresponding entry in the static information database is immediately overwritten with the latest timestamp and attendance data from the real-time activity dynamic information, generating the optimal activity status estimate for the current moment. Then, based on the updated key time nodes, the time difference between the current system time and the activity start and end times is calculated, and the current time period is mapped to a specific activity stage marker according to a preset time window threshold logic. For example, when the current time is calculated to be within the interval from 3 hours before the updated start time of the activity to the official start time of the activity, this period is marked as the pre-gathering period; when the current time is within the interval from the updated end time of the activity to 1 hour after the end of the activity, this period is marked as the outbreak period; when the current time is within the interval from 1 hour to 3 hours after the end of the updated activity, this period is marked as the evacuation period. The final output is a dynamic dataset containing the accurately corrected activity attributes and the current stage label.
[0024] Specifically, in step S13, based on the base station-site spatial mapping table, the cleaned signaling data blocks, cleaned gate data blocks, and dynamically updated activity status and stage marker sets are spatiotemporally aligned to obtain a real-time standard input snapshot. It should be understood that because the spatial coordinates of signaling data are based on base stations, and gate data is based on subway stations, the spatial references of the two types of data are inconsistent, and the temporal granularity of the multi-source data differs, making direct data association and fusion impossible. Therefore, this application further performs spatiotemporal alignment operations on the cleaned signaling data blocks, cleaned gate data blocks, and dynamically updated activity status and stage marker sets based on the base station-site spatial mapping table to unify the spatiotemporal reference of the data, thereby achieving effective fusion of multi-source data. This allows for the formation of a spatiotemporally consistent real-time standard input snapshot, fully preserving the correlation between passenger flow data and activity information, and providing structured data support for subsequent multi-dimensional feature extraction and model inference.
[0025] Specifically, in one possible embodiment, step S13 is implemented as follows: First, based on a pre-built base station-site spatial mapping table, the base station locations in the cleaned signaling data are converted into corresponding subway station identifiers to achieve spatial dimension unification. Second, the time granularity of the signaling data and the gate data is uniformly resampled to 15 minutes to ensure time dimension consistency. Finally, using the time window and the station identifier as a joint primary key, the aligned passenger flow data is associated and spliced with the dynamically updated activity status and stage marker set to form a real-time standard input snapshot containing passenger flow, spatial, and activity information.
[0026] In the aforementioned passenger flow prediction method based on the correlation modeling of mobile phone signaling and subway gate data, step S2 involves feature encoding of the real-time standard input snapshot to obtain a real-time inference feature vector. It should be understood that because the real-time standard input snapshot contains multi-source heterogeneous data, its form is scattered and lacks structured feature expression, it cannot be directly input into a deep learning model for inference. Furthermore, the spatiotemporal evolution of passenger flow during major events is closely related to event attributes and station characteristics, requiring systematic characterization to support accurate prediction. Therefore, this application further performs multi-dimensional feature encoding on the real-time standard input snapshot, integrating dynamic spatiotemporal correlation features, periodic and event attribute features, and static station features to construct a feature system that comprehensively reflects the influencing factors of passenger flow. This transforms the raw data into a standardized vector recognizable by the model, providing high-quality feature input for subsequent passenger flow prediction and significantly improving the model's ability to capture the temporal abrupt changes and spatial concentration characteristics of passenger flow during major events.
[0027] In particular, in one specific embodiment, Figure 4 This is a flowchart of sub-step S2 of the passenger flow prediction method based on the association modeling of mobile phone signaling and subway gate data according to an embodiment of this application. Figure 4As shown, step S2 includes: S21, performing dynamic spatiotemporal correlation feature calculation on the number of signaling users and gate passenger flow of each station in the real-time standard input snapshot, including the current time window and the previous time window, to obtain a dynamic feature set; S22, performing periodic and activity attribute feature encoding on the timestamps, activity stage markers, and activity attribute data in the real-time standard input snapshot to obtain an encoded feature set; S23, extracting static features of the station from the static station feature library; S24, performing multi-dimensional feature concatenation on the encoded feature set, dynamic feature set, and static features to obtain a real-time inference feature vector.
[0028] Specifically, step S21 involves calculating dynamic spatiotemporal correlation features for the number of signaling users and gate passenger flow at each station within the current and previous time windows in the real-time standard input snapshot to obtain a dynamic feature set. It should be understood that since the number of signaling users and gate passenger flow in the real-time standard input snapshot are only discrete data within a single time window, they cannot reflect dynamic evolutionary characteristics such as crowd gathering speed and flow concentration. These characteristics are crucial for distinguishing passenger flow during major events from daily passenger flow and directly affect the sensitivity of the prediction model to sudden changes in passenger flow. Therefore, this application further calculates dynamic spatiotemporal correlation features based on relevant data from the current and previous time windows to quantify the spatiotemporal change trend of passenger flow. In a specific example of this application, step S21 includes: calculating dynamic spatiotemporal correlation features using the following formula:
[0029]
[0030] in, This represents the number of signaling users at each site during the current time window. This represents the number of signaling users at each site in the previous time window. This represents the total number of passengers entering the target station within the current time window. This represents the total passenger flow entering all stations across the entire subway network within the current time window. For signaling aggregation rate, This refers to the concentration of passenger flow. Specifically, it calculates the ratio of the difference between the current number of signaled users and the previous time window to the number of users in the previous time window. Essentially, it quantifies the rate of population growth per unit time. Mathematically, it transforms discrete user data into a relative value reflecting the intensity of temporal change. Positive values indicate that the population is converging; the larger the value, the faster the convergence. Zero or negative values correspond to no convergence or dispersal. This method accurately captures the temporal evolution intensity of passenger flow from the pre-gathering period to the peak of major events. Then, by calculating the ratio of passenger flow entering the target station to the total passenger flow entering the entire network, the degree of concentration of passenger flow spatial distribution is quantified. Mathematically, this normalizes the passenger flow of a single station within the context of the entire network's passenger flow, with a value ranging from 0 to 1. The closer the value is to 1, the more concentrated the passenger flow is towards the target station. This effectively separates the difference between the dispersed passenger flow of daily life and the spatially concentrated passenger flow caused by major events, highlighting the core position of the target station's passenger flow during specific periods. The two formulas complement each other from the dimensions of temporal evolution intensity and spatial distribution concentration, jointly providing quantitative mathematical support for the model to characterize the dynamic spatiotemporal correlation of passenger flow. In this way, the dynamic patterns of passenger flow during major events can be accurately captured, providing the model with core features reflecting real-time changes in passenger flow and effectively improving the model's prediction accuracy for short-term explosive passenger flow.
[0031] Specifically, step S22 involves periodically and activity-attribute-specific encoding of the timestamps, activity stage markers, and activity attribute data in the real-time standard input snapshot to obtain an encoded feature set. It should be understood that since timestamps, activity stage markers, and activity attribute data are mostly non-numerical information, they cannot be directly parsed by the model. Furthermore, this information contains periodic patterns in passenger flow and unique influencing factors specific to major events, which are crucial for the model's accurate adaptation to major event scenarios. Therefore, this application further encodes these types of data with periodic and activity-attribute-specific features, transforming non-numerical information into standardized numerical features to enhance the model's ability to perceive time-periodic patterns and the impact of events. This allows the model to effectively distinguish the differences in the impact of different time periods and types of activities on passenger flow, accurately identify the special temporal characteristics of passenger flow during major events, and provide strong support for predicting changes in passenger flow in advance.
[0032] Specifically, in one possible embodiment, step S22 is implemented as follows: First, temporal feature encoding is performed to... Based on the basic periodic coding (where T=24h), a new stage marker for major events is added, clearly defining the pre-gathering period as 3 hours before the event, the peak period as 1 hour after the event, and the evacuation period as 1-3 hours after the event. Simultaneously, temporal characteristics such as the start / end time difference of the event and the peak time of passenger flow are incorporated to enhance the characterization of sudden changes in passenger flow patterns. Secondly, spatial and passenger flow correlation feature coding is performed, optimizing the spatial grid precision to 500m×500m. The focus is on statistically analyzing the signaling density, straight-line distance between stations and venues, and traffic accessibility within a 3km radius of major event venues, generating a station-event correlation mapping table. Simultaneously, based on mobile signaling data, the crowd gathering rate (growth rate of inflow per unit time), dwell density, and flow concentration in the area surrounding the event are calculated. Finally, the event attributes are quantitatively coded, and event-specific features are constructed according to event type (sports event, concert, exhibition, etc.), expected number of participants (classified as less than 10,000, 10,000-30,000, and more than 30,000), and event time (weekday / weekend, daytime / nighttime). The impact intensity of different types of events on passenger flow is quantified, and all coding results are finally integrated to form a coded feature set.
[0033] Specifically, step S23 involves extracting static features of stations from a static station feature library. It should be understood that the inherent attributes of stations (such as their spatial relationship with event venues and accessibility) have a long-term impact on passenger flow distribution patterns. These attributes are relatively stable and do not require real-time calculation; solidifying them as static features improves feature construction efficiency. Furthermore, static features complement dynamic and event attribute features, enhancing the completeness of the feature system. Therefore, this application further extracts static features of corresponding stations from a pre-built static station feature library to supplement the feature dimensions reflecting the inherent attributes of the stations. This allows the model to comprehensively consider both spatiotemporal static constraints and dynamic changes, accurately characterizing the carrying and attracting capacity of different stations for major events, and further improving the spatial accuracy of passenger flow prediction.
[0034] Specifically, in one possible embodiment, step S23 is implemented as follows: First, the static site feature library pre-stores the core static attributes of each site, including the site's latitude and longitude, straight-line distance to the main event venue, walking time and number of transfers to the venue, site size level, and surrounding transportation connections. Second, based on the site ID in the real-time standard input snapshot, a precise matching query is performed in the static site feature library. Finally, all static attribute information of the matching sites is extracted and organized into a structured feature vector according to a preset format to ensure that the static feature dimensions of each site are consistent, laying the foundation for subsequent multi-dimensional feature stitching.
[0035] Specifically, in step S24, the encoded feature set, dynamic feature set, and static features are multi-dimensionally concatenated to obtain a real-time inference feature vector. It should be understood that since the encoded feature set, dynamic feature set, and static features characterize passenger flow influencing factors from different dimensions, but are scattered and dimensionally independent, they cannot form a unified model input, and the numerical ranges of each feature differ, which may affect the model training and inference effects. Therefore, this application further performs multi-dimensional concatenation and standardization processing on the encoded feature set, dynamic feature set, and static features according to preset rules to integrate multi-dimensional features and form a feature vector with a unified structure and standardized values. In this way, all key factors such as dynamic spatiotemporal changes, periodic patterns, the influence of activity attributes, and the static characteristics of the site can be integrated into a single feature vector, meeting the input requirements of deep learning models, providing full-dimensional, high-quality feature support for efficient model inference, and ensuring the accuracy and stability of passenger flow prediction.
[0036] Specifically, in one possible embodiment, step S24 is implemented as follows: First, the feature concatenation order is determined to be dynamic feature set, encoded feature set, and static features, ensuring that the positions of features in each dimension are fixed. Second, the numerical features in the encoded feature set, dynamic feature set, and static features are normalized to map all feature values to the same numerical range, eliminating the influence of differences in units. Then, the standardized dynamic feature vector, encoded feature vector, and static feature vector are horizontally concatenated in sequence to form a one-dimensional long vector. Finally, the concatenated vector undergoes dimension verification and format conversion to ensure that its dimension completely matches the dimension of the input layer of the trained passenger flow prediction model, ultimately generating a real-time inference feature vector.
[0037] In the aforementioned passenger flow prediction method based on the correlation modeling of mobile phone signaling and subway gate data, step S3 involves inputting the real-time inferred feature vector into the trained passenger flow prediction model to obtain a future passenger flow prediction sequence. In a specific example of this application, the trained passenger flow prediction model includes a two-layer LSTM module, a multi-head attention module, a Transformer encoder module, and a decoding prediction module. It should be understood that, due to the highly nonlinear, sudden, and complex temporal dependence characteristics of subway passenger flow during major events, and the obvious pre-aggregation, burst, and dispersal phases in the passenger flow evolution process, a single linear statistical model or a simple recurrent neural network is difficult to simultaneously capture long-term trends and accurately characterize key abrupt changes, easily leading to lag in the prediction of peak passenger flow times or biased intensity estimation. Therefore, in the technical solution of this application, the real-time inferred feature vector is further input into the trained passenger flow prediction model to obtain a future passenger flow prediction sequence, thereby utilizing a hybrid model architecture composed of a two-layer LSTM, a multi-head attention mechanism, and a Transformer encoder to deeply mine multi-level spatiotemporal correlation information in the input features. In this way, LSTM can capture sequence dependencies, attention mechanism can be used to enhance sensitivity to key moments such as the end of the event, and combined with the global modeling capability of Transformer, it can accurately predict changes in passenger flow within a long time window of 3 hours in the future, ensuring that the prediction results are accurate in terms of trends and can also keenly capture short-term peak passenger flow.
[0038] Specifically, in one possible embodiment, step S3 is implemented as follows: First, the constructed real-time inference feature vector sequence is pre-loaded into the hybrid neural network model of the inference engine. The data flow first passes through a two-layer LSTM module, which uses its internal input gate and forget gate mechanisms to extract the slow aggregation trend in the early stage of the activity and the explosive growth features after the end of the activity, respectively. Then, the output hidden state sequence is transmitted to the multi-head attention module. This module calculates the attention weights at different time steps to enhance the features of key time nodes such as the start and end times of the activity. The weighted feature sequence is further input into the Transformer encoder module, which uses the self-attention mechanism to capture long-term correlation dependencies throughout the entire cycle. Finally, the decoding and prediction module maps the high-dimensional features into specific numerical sequences and outputs the predicted passenger flow of each station at a 15-minute granularity for the next 3 hours.
[0039] Specifically, the training process of the passenger flow prediction model is as follows: First, an initial dataset containing 20,000 positive samples and 80,000 negative samples is constructed based on the gate and signaling data during historical major events. Considering the sparsity of samples from major events, the SMOTE synthetic minority class oversampling algorithm is used to enhance high-risk passenger flow samples. At the same time, random undersampling is implemented for daily off-peak passenger flow samples, dynamically adjusting the ratio of positive to negative samples to 1:3. The data is then divided into training, validation, and test sets in an 8:1:1 ratio. Subsequently, the system initializes and constructs a hybrid network architecture containing a two-layer LSTM module, a multi-head attention module, and a Transformer encoder module. The processed temporal feature matrix is input into this network, and the data sequentially passes through the first layer of the LSTM to capture the data before the event. The model employs a multi-head attention mechanism to weight key time points, and finally, a Transformer encoder integrates long-term correlations across the entire timeframe. During backpropagation, a combined loss function is calculated between predicted and true values. This loss function is a weighted average of mean squared error loss and risk level classification loss. Based on this loss value, the Adam optimizer updates the model weights with an initial learning rate of 0.001 and a cosine annealing strategy. Model performance is evaluated on a validation set after each iteration. Convergence is determined when the validation set loss value fluctuates less than 0.003 for six consecutive iterations and the test set prediction accuracy is not less than 88%. The final deployable model parameters are then output. To quantify the optimization objective during training, the following combined loss function is used for calculation:
[0040] in, This represents the total loss value during model training, used to guide the backpropagation of gradients. This represents the total number of samples in the current training batch. Indicates the first The actual passenger flow data for each sample Indicates the first The model predicts passenger flow values for a sample. This is the mean squared error term, used to measure the numerical accuracy of regression predictions. This is the weighting coefficient for the regression loss, and its value is set to 0.7. The number of risk levels (e.g., low, medium, and high). Indicates the first Each sample belongs to category The true probability (0 or 1). The model predicts the first... Each sample belongs to category The probability is given by the logarithmic summation term within parentheses, which is the cross-entropy classification loss, used to measure the accuracy of risk level classification. The weighting coefficient for the classification loss is set to 0.3.
[0041] In the aforementioned passenger flow prediction method based on the correlation modeling of mobile phone signaling and subway gate data, step S4 involves identifying passenger flow anomalies and determining risks based on a historical baseline database and station capacity parameters to obtain a spatiotemporal risk level distribution for the future passenger flow prediction sequence. It should be understood that since the future passenger flow prediction sequence only reflects changes in passenger flow values and does not incorporate historical normal levels and station carrying capacity limits, it is impossible to directly determine whether an abnormal cluster is caused by a major event, and a single-dimensional judgment can easily lead to misjudgment or omission of risks. Therefore, this application further integrates a historical baseline database and station capacity parameters, and conducts dual judgment through anomaly deviation calculation and risk index quantification to accurately distinguish between normal and abnormal passenger flow and quantify the risk level of different stations. This allows for a risk level distribution covering the entire time and space, clearly presenting high, medium, and low-risk stations and time periods, providing accurate decision-making basis for subsequent hierarchical management and ensuring the targeted implementation of management measures.
[0042] In particular, in one specific embodiment, Figure 5 This is a flowchart of sub-step S4 of the passenger flow prediction method based on the association modeling of mobile phone signaling and subway gate data according to an embodiment of this application. Figure 5 As shown, step S4 includes: S41, based on the historical baseline database, calculating and identifying passenger flow anomaly deviations in the future passenger flow prediction sequence to obtain a prediction sequence with anomaly markers; S42, based on station capacity parameters, quantitatively calculating the passenger flow risk index of the prediction sequence with anomaly markers to obtain a passenger flow risk index set; S43, based on risk level classification thresholds, determining and mapping the risk level of the passenger flow risk index set to obtain a spatiotemporal risk level distribution.
[0043] Specifically, in step S41, based on a historical baseline database, abnormal deviations in future passenger flow prediction sequences are calculated and identified based on historical baselines to obtain prediction sequences marked with anomalies. It should be understood that since the predicted future passenger flow values themselves cannot reflect the degree of deviation from normal levels, and the essential difference between passenger flow during major events and daily commuting passenger flow needs to be highlighted through historical comparison, relying solely on predicted values can easily misjudge normal peak passenger flow as abnormal. Therefore, this application further uses a historical baseline database as a reference to calculate the deviation rate of predicted passenger flow relative to the historical baseline, and combines this with spatial correlation correction to accurately identify abnormally concentrated passenger flow caused by major events. This effectively distinguishes between daily passenger flow fluctuations and event-specific abnormal passenger flow, allowing for the selection of core concerns for subsequent risk assessment and ensuring that risk judgment focuses on the truly controllable spatiotemporal points.
[0044] It is understandable that the passenger flow anomaly identification method described in the above embodiments has an inherent technical flaw in calculating the deviation rate of predicted passenger flow relative to the historical baseline. This method treats each subway station as an isolated, homogeneous analytical unit, calculating its deviation rate merely as a longitudinal comparison between the station's predicted passenger flow and its own historical baseline, completely ignoring the spatial topological relationships inherent in the subway network system. The root cause of this approach is the failure to integrate the network's topological structure information into the core algorithm for anomaly detection. It is also understandable that passenger flow anomalies at a station, especially during major events, are usually not isolated occurrences but rather propagate and spread along the subway network. A surge in passenger flow at a core station will inevitably have a significant and predictable chain reaction with its upstream and downstream neighboring stations. Because the original mechanism fails to consider this synergistic effect of passenger flow between topologically adjacent stations, it directly leads to distortion in the deviation rate calculation, resulting in misjudgments or missed detections of risks. For example, a transfer station adjacent to a core station might not reach a set absolute threshold for passenger flow growth, but if this growth is entirely due to passenger overflow from the core station, it becomes a secondary risk that should not be ignored. Similarly, for historically low-passenger-flow stations at the network's periphery, even a tiny absolute increase in passenger flow can easily trigger false alarms, interfering with the priority of operational decisions. To address these technical deficiencies, this application preferably proposes a spatial correction deviation rate calculation method based on topological proximity weighting. This technique introduces the topological structure of the metro network to spatially smooth and correct the original deviation rate of individual stations, thereby more accurately identifying group-based passenger flow anomalies spreading from points to areas.
[0045] In particular, in one specific embodiment, Figure 6 This is a flowchart of sub-step S41 of the passenger flow prediction method based on the association modeling of mobile phone signaling and subway gate data according to an embodiment of this application. Figure 6 As shown, step S41 includes: S411, constructing a proximity weight matrix based on subway line topology data; S412, calculating the individual original deviation rate vector for the future passenger flow prediction sequence based on the historical baseline database; S413, performing spatial correction deviation rate fusion calculation on the individual original deviation rate vector based on the proximity weight matrix to obtain a spatially corrected complete deviation rate sequence; S414, generating a prediction sequence with anomaly markers based on the comparison between the spatially corrected complete deviation rate sequence and the abnormal deviation rate threshold.
[0046] More specifically, step S411 involves constructing a proximity weight matrix based on subway line topology data. It should be understood that a mathematical tool is needed to quantify the strength of the association between any two stations in the network, providing a weighting basis for subsequent spatial correction. Specifically, for any two stations... and The proximity weights between them are calculated using a formula that integrates physical distance and transfer relationships. , means as follows:
[0047] in, Representing the proximity weights, it is a core element of the proximity weight matrix W. Represents the physical distance between stations, i.e., the number of stations between two stations; It is a transfer indicator function, which takes the value of 1 when the two stations are transfer stations, and 0 otherwise; This is a distance attenuation factor used to balance the importance of physical distance and transfer relationships. In other words, two core spatial relationships in the subway network are mathematically modeled: the attenuation effect of physical distance reflects the directness of passenger flow transmission along the line, while the transfer association effect characterizes the intensity of passenger flow interaction between different lines, so as to generate a static weight matrix W that can accurately reflect the topological association of the entire network, providing a key and stable data foundation for the calculation of subsequent steps.
[0048] More specifically, in a particular example of this application, a specific numerical calculation embodiment is given to aid in the illustration: setting a distance attenuation factor. =0.6, site With the site Physical station spacing =2 stations, and the two stations are transfer stations (transfer indication function) =1). First, substitute into the formula. : Calculate the physical distance weight term 0.6×1+21=0.6×0.333≈0.1998, and the transfer-related weight term (1 0.6) × 1 = 0.4, summing the two gives... ≈0.1998 + 0.4 = 0.5998. Therefore, the station... and Because of the transfer connection and the close physical distance, the proximity weight is about 0.6, which reflects the strong topological connection between the two stations in the subway network and provides a reasonable weight basis for subsequent spatial correction.
[0049] More specifically, step S412 involves calculating the individual raw deviation rate of the future passenger flow prediction sequence based on the historical baseline database to obtain an individual raw deviation rate vector. That is, an initial indicator is needed to measure the independent, uncorrected passenger flow growth at each station. Specifically, for each spatiotemporal point... Use it to predict passenger flow Compared with historical baseline passenger flow Perform the calculation.
[0050]
[0051] in, This represents the individual's original deviation rate; This is a smoothing term introduced to prevent the denominator from being zero. In other words, it generates a basic, quantified measure of anomaly for each station. Although it doesn't yet consider spatial correlations, it's a necessary input for subsequent fusion calculations, generating an original deviation rate vector for all stations across the network at a specific time. This vector serves as an intermediate result, providing a data source for the final correction calculation.
[0052] More specifically, in a particular example of this application, a specific numerical calculation embodiment is given to aid in the illustration: at a certain spatiotemporal point Predicted passenger flow =1500 people, historical baseline passenger flow =1000 people, smoothing term =1×10 6. First, substitute the values into the formula. Calculate the numerator 1500 1000 = 500, denominator 1000 + 1 × 10 6≈1000.000001, dividing the two gives... ≈1000.000001 / 500≈0.4999998. This shows that the predicted passenger flow at this spatiotemporal point is about 50% higher than the historical baseline, initially indicating an abnormal passenger flow trend. However, this value does not consider the spatial correlation between stations and only reflects the longitudinal deviation of the stations themselves.
[0053] More specifically, in step S413, based on the proximity weight matrix, the original individual deviation rate vectors are fused to obtain a spatially corrected complete deviation rate sequence. It should be understood that the original deviation rate cannot reflect the network's collaborative effect and needs a mechanism to correlate it with the states of neighboring stations. Specifically, drawing on the core idea of graph convolution, a fusion formula is used to weight and fuse the individual deviation rate of a station with the deviation rates of its neighboring stations to obtain the final spatially corrected deviation rate. , means as follows:
[0054] in, It is the spatial correction deviation rate of the final output; The individual raw bias rate representing other stations in the network; It is a fusion coefficient used to adjust the inter-influence between self-influence and neighbor-influence; while This means for all sites in the network. A summation process is performed, which constructs a dynamic equilibrium model of individual performance and neighborhood influence. The formula... The item retains the site's own abnormal signals, while The term "deviation rate" is a key spatial smoothing term that simulates the propagation and diffusion of passenger flow anomalies in the network topology. If multiple neighbors of a station, especially those with high weights, generally exhibit high deviations, even if the station itself has a low deviation, its corrected deviation rate will be significantly increased. This approach moves away from viewing each station in isolation and considers the entire network's passenger flow situation as a whole, generating a deviation rate indicator that better reflects the true level of network risk. This provides a more accurate and robust basis for subsequent anomaly detection.
[0055] More specifically, in a particular example of this application, a specific numerical calculation embodiment is given to aid in the illustration: Setting the spatial fusion coefficient β = 0.3, a certain site... Individual original deviation rate =0.5, its neighboring stations Proximity weight =0.4, original deviation rate =0.6, neighboring stations Proximity weight =0.3, original deviation rate =0.4 (other non-neighboring sites have a weight of 0, which can be ignored when summing). First, substitute into the formula. : Calculate the contribution of self-bias (1) 0.3) × 0.5 = 0.35, the weighted sum of neighboring stations is 0.4 × 0.6 + 0.3 × 0.4 = 0.24 + 0.12 = 0.36, the neighboring contribution is 0.3 × 0.36 = 0.108, and the sum of the two is... =0.35 + 0.108 = 0.458. Therefore, it can be seen that although the site... Its original deviation rate is 0.5, but due to the collaborative influence of neighboring high-deviation sites, the corrected deviation rate is 0.458. This not only preserves its own abnormal signals but also integrates network topology association information, thus avoiding misjudgments caused by isolated judgments.
[0056] The ultimate goal and achievable technical effect of this approach is to significantly improve the accuracy and sensitivity of passenger flow anomaly identification during major subway events, overcoming the risk misjudgment and underjudgment problems caused by existing methods neglecting the topological relationships between stations. By creatively introducing a spatial correction mechanism, the calculation of the deviation rate is no longer an isolated numerical comparison, but a dynamic evaluation process reflecting network synergy. This allows the system to identify secondary risk areas formed by the spread of core anomalies earlier, even if the absolute increase in passenger flow in these areas has not yet reached the traditional threshold. Simultaneously, through spatial smoothing, false alarms caused by minor fluctuations at historically low-passenger-flow stations are effectively suppressed, making the early warning system more stable and reliable, ensuring that operational decision-making resources can truly focus on risk points with actual network impact. Ultimately, this method outputs a more accurate deviation rate index after spatial correction, providing high-quality data input for subsequent risk classification, early warning issuance, and even the generation of automated control strategies. This comprehensively improves the intelligence level and practical value of the entire passenger flow prediction and control system, ensuring the safety and efficiency of subway operations.
[0057] More specifically, step S414 generates a predicted sequence with anomaly markers based on a comparison between the spatially corrected complete deviation rate sequence and the abnormal deviation rate threshold. It should be understood that while the spatially corrected complete deviation rate sequence eliminates the influence of misjudgments of isolated sites, it lacks clear criteria to define the boundary between normal fluctuations and abnormal clusters. Relying solely on subjective judgment can easily lead to insufficient accuracy in anomaly identification, failing to effectively distinguish between daily passenger flow fluctuations and genuine anomalies caused by major events. Therefore, this application further compares the spatially corrected complete deviation rate sequence with a preset abnormal deviation rate threshold, combining dual identification criteria to screen abnormal spatiotemporal points, thereby achieving accurate marking of abnormal passenger flow associated with major events. This effectively eliminates false anomaly signals, ensuring that all marked abnormal passenger flow is directly related to major events, providing accurate and focused processing targets for subsequent risk index quantification calculations, and guaranteeing the scientific nature of risk assessment.
[0058] Specifically, in one possible embodiment, step S414 is implemented as follows: First, a preset abnormal deviation rate threshold is retrieved, which is determined based on historical passenger flow data from major events and statistical analysis of daily passenger flow fluctuations. Second, the complete spatially corrected deviation rate sequence is traversed, and the deviation rate of each spatiotemporal point is compared with the threshold. If the deviation rate exceeds the threshold, the signaling aggregation rate corresponding to that spatiotemporal point is further verified to ensure it meets the standard and whether the growth in passenger flow at the station conforms to the characteristics associated with major events. Finally, spatiotemporal points that simultaneously meet the deviation rate threshold condition and the correlation characteristic verification are marked as abnormal. The abnormal markers are then integrated with the original future passenger flow prediction sequence to generate a prediction sequence with abnormal markers for each spatiotemporal point.
[0059] Specifically, in step S42, based on station capacity parameters, the passenger flow risk index of the predicted sequence marked with anomalies is quantitatively calculated to obtain a passenger flow risk index set. It should be understood that since passenger flow marked with anomalies only indicates a deviation from historical normal levels, but different stations have different carrying capacities, the risk level caused by abnormal passenger flow of the same scale varies at different stations, and the rate of passenger flow growth exacerbates the risk level. Therefore, this application further combines station capacity parameters and passenger flow growth characteristics to quantitatively calculate the risk index of each abnormal spatiotemporal point, thereby accurately measuring the threat level of abnormal passenger flow to the operational safety of the station. This overcomes the limitation of merely identifying anomalies without assessing risks, providing a quantitative basis for subsequent risk classification and ensuring that high-risk stations receive priority in the allocation of control resources.
[0060] Specifically, step S43 involves determining and mapping the risk level of the passenger flow risk index set based on risk level classification thresholds to obtain a spatiotemporal risk level distribution. It should be understood that since the passenger flow risk index set is only a quantitative value and lacks an intuitive risk level classification, it cannot directly guide the operations department in carrying out tiered management. Furthermore, different risk levels correspond to different handling strategies, requiring clear boundary definition. Therefore, this application further sets risk level classification thresholds based on operational practices and safety standards to map the risk indices to levels, thereby transforming abstract risk indices into concrete risk levels. This allows for the formation of a spatiotemporal risk level distribution with the station as the spatial dimension and the time window as the temporal dimension, clearly presenting the risk status of each station at different times, providing an intuitive and operable basis for matching management strategies.
[0061] Specifically, in one possible embodiment, steps S42 and S43 are implemented as follows: First, a risk assessment index formula is constructed based on site capacity parameters and passenger flow growth characteristics: Subsequently, the risk index for each anomalous spatiotemporal point was calculated based on this formula. And perform level mapping based on preset thresholds: when When a situation is deemed high-risk, it is necessary to activate flow control and emergency transportation capacity; When the risk level is determined to be medium, additional backup trains need to be added; If a site is deemed low-risk, it can continue normal operations. Finally, the risk levels of all sites are integrated according to time windows to form an intuitive spatiotemporal risk level distribution.
[0062] In the aforementioned passenger flow prediction method based on the correlation modeling of mobile phone signaling and subway gate data, step S5 involves classifying and controlling the spatiotemporal risk level distribution based on a control strategy knowledge base to obtain a closed-loop control execution plan. It should be understood that because the spatiotemporal risk level distribution presents differentiated risk states across multiple stations and time periods, the demand for capacity and control varies fundamentally across different risk levels. Furthermore, the lack of systematic strategy support can easily lead to insufficient targeted measures, resource waste, or execution conflicts, preventing the formation of a complete control loop. Therefore, this application further utilizes a control strategy knowledge base to perform precise strategy matching and integration based on risk levels, thereby generating a systematic control plan adapted to different risk scenarios. This enables precise correspondence between risks and control measures, ensuring priority handling of high-risk stations and reasonable adjustment of medium- and low-risk stations, comprehensively improving the safety and efficiency of subway operations during major events, and providing operational departments with directly implementable practical guidance.
[0063] Specifically, in one possible embodiment, step S5 is implemented as follows: First, comprehensively analyze the spatiotemporal risk level distribution data and extract the risk level of each station. Second, retrieve the control strategy knowledge base using the highest risk level of each station as an index to generate a refined control plan: for high-risk stations, output suggestions for increasing capacity (such as shortening the departure interval by 20-30% one hour before the event), emergency additional trains (adding direct trains to transfer stations after the event), and cross-line support plans (dispatching trains from surrounding lines to supplement capacity); for platform control, output practical suggestions such as one-way traffic guidance, security checkpoint expansion, and queue limiting outside the station. Finally, systematically optimize all matched and generated control instructions, verify the consistency of departure interval adjustments on the same line, and form a closed-loop control execution plan with a clear structure that includes specific capacity adjustment ratios and execution actions.
[0064] In summary, the passenger flow prediction method based on mobile phone signaling and subway gate data association modeling, as described in this application, is elucidated. First, it collects mobile phone signaling data covering macro-level passenger flow, micro-level gate data, and information on major events in real time. This data is then dynamically cleaned and spatiotemporally aligned to construct a panoramic standard input. Subsequently, a deep learning model is driven by multi-dimensional spatiotemporal feature encoding to analyze the temporal evolution patterns during the event and predict future passenger flow sequences. Furthermore, it combines historical baselines and station capacity to calculate anomaly deviations and determine risks, identifying spatiotemporal risk levels. Finally, based on a management knowledge base, a closed-loop execution plan including capacity scheduling and flow control measures is generated. This effectively overcomes the limitations of perception from a single data perspective and the temporal lag of conventional models, thereby achieving accurate prediction and proactive intervention of the entire passenger flow chain during major events.
[0065] As described above, the passenger flow prediction system 100 based on the association modeling of mobile phone signaling and subway gate data according to the embodiments of this application can be implemented in various wireless terminals, such as servers with passenger flow prediction algorithms based on the association modeling of mobile phone signaling and subway gate data. In one possible implementation, the passenger flow prediction system 100 based on the association modeling of mobile phone signaling and subway gate data according to the embodiments of this application can be integrated into the wireless terminal as a software module and / or hardware module. For example, the passenger flow prediction system 100 based on the association modeling of mobile phone signaling and subway gate data can be a software module in the operating system of the wireless terminal, or it can be an application developed for the wireless terminal; of course, the passenger flow prediction system 100 based on the association modeling of mobile phone signaling and subway gate data can also be one of many hardware modules of the wireless terminal.
[0066] Alternatively, in another example, the passenger flow prediction system 100 based on the association modeling of mobile phone signaling and subway gate data can also be a separate device from the wireless terminal, and the passenger flow prediction system 100 based on the association modeling of mobile phone signaling and subway gate data can be connected to the wireless terminal via wired and / or wireless networks, and transmit interactive information in accordance with the agreed data format.
Claims
1. A passenger flow prediction method based on the correlation modeling of mobile phone signaling and subway gate data, characterized in that, include: Dynamic preprocessing is performed on real-time signaling data streams, real-time gate data streams, and real-time activity dynamic information to obtain a real-time standard input snapshot; Feature encoding is performed on real-time standard input snapshots to obtain real-time inferred feature vectors; The real-time inferred feature vectors are input into the trained passenger flow prediction model to obtain the future passenger flow prediction sequence. Based on historical baseline data and station capacity parameters, passenger flow anomalies and risks are identified and assessed in future passenger flow forecast sequences to obtain the spatiotemporal risk level distribution. Based on the knowledge base of control strategies, hierarchical control decisions are made on the spatiotemporal risk level distribution to obtain a closed-loop control execution plan.
2. The passenger flow prediction method based on the correlation modeling of mobile phone signaling and subway gate data as described in claim 1, characterized in that, Dynamic preprocessing is performed on real-time signaling data streams, real-time gate data streams, and real-time activity dynamic information to obtain a real-time standard input snapshot, including: Anomaly cleaning is performed on the real-time signaling data stream and the real-time gate data stream to obtain cleaned signaling data blocks and cleaned gate data blocks; Based on a pre-defined static activity information database, real-time activity dynamic information is dynamically updated and events are marked to obtain a dynamically updated set of activity status and stage markers; Based on the base station-site spatial mapping table, the cleaned signaling data blocks, cleaned gate data blocks, and dynamically updated activity status and stage marker sets are spatiotemporally aligned to obtain a real-time standard input snapshot.
3. The passenger flow prediction method based on the correlation modeling of mobile phone signaling and subway gate data as described in claim 1, characterized in that, Feature encoding is performed on a real-time standard input snapshot to obtain a real-time inferred feature vector, including: Dynamic spatiotemporal correlation features are calculated for the number of signaling users and gate passenger flow at each site in the real-time standard input snapshot, which includes the current time window and the previous time window, to obtain a dynamic feature set; Periodically encode the timestamps, activity phase markers, and activity attribute data in the real-time standard input snapshot with activity attribute features to obtain the encoded feature set; Extract static features of a site from a static site feature library; The encoded feature set, dynamic feature set, and static features are concatenated into a multi-dimensional feature set to obtain the real-time inference feature vector.
4. The passenger flow prediction method based on the correlation modeling of mobile phone signaling and subway gate data as described in claim 3, characterized in that, Dynamic spatiotemporal correlation features are calculated for the number of signaling users and gate passenger flow at each site, including the current time window and the previous time window, in the real-time standard input snapshot to obtain a dynamic feature set. This includes calculating the dynamic spatiotemporal correlation features using the following formula: in, This represents the number of signaling users at each site during the current time window. This represents the number of signaling users at each site in the previous time window. This represents the total number of passengers entering the target station within the current time window. This represents the total passenger flow entering all stations across the entire subway network within the current time window. For signaling aggregation rate, This refers to the concentration of the flow.
5. The passenger flow prediction method based on the correlation modeling of mobile phone signaling and subway gate data as described in claim 1, characterized in that, The trained passenger flow prediction model includes a two-layer LSTM module, a multi-head attention module, a Transformer encoder module, and a decoding prediction module.
6. The passenger flow prediction method based on the correlation modeling of mobile phone signaling and subway gate data as described in claim 1, characterized in that, Based on historical baseline data and station capacity parameters, passenger flow anomalies are identified and risks are assessed in future passenger flow forecast sequences to obtain a spatiotemporal risk level distribution, including: Based on the historical baseline database, the passenger flow forecast sequence is calculated and identified based on the historical baseline to obtain a forecast sequence with anomaly markers. Based on the station capacity parameter, the passenger flow risk index is quantitatively calculated for the predicted sequence with anomaly markers to obtain the passenger flow risk index set. Based on the risk level classification threshold, the risk level of the passenger flow risk index set is determined and mapped to obtain the spatiotemporal risk level distribution.
7. The passenger flow prediction method based on the correlation modeling of mobile phone signaling and subway gate data as described in claim 6, characterized in that, Based on a historical baseline database, the project calculates and identifies passenger flow anomaly deviations in the future passenger flow forecast sequence to obtain a forecast sequence with anomaly markers, including: Based on the subway line topology data, a proximity weight matrix is constructed. Based on the historical baseline database, the individual original deviation rate of the future passenger flow prediction sequence is calculated to obtain the individual original deviation rate vector. Based on the proximity weight matrix, the spatially corrected deviation rate is fused and calculated from the original deviation rate vector of individuals to obtain the complete spatially corrected deviation rate sequence. A predicted sequence with anomaly markers is generated by comparing the spatially corrected complete deviation rate sequence with the anomaly deviation rate threshold.
Citation Information
Patent Citations
Subway multi-task passenger flow prediction method fusing space-time multi-dimensional features
CN120688667A
Tourism rail transit passenger flow characteristic analysis method based on mobile phone signaling data
CN121029900A
Subway station large passenger flow early warning method and system based on real-time data
CN121148123A