A real-time prediction method for rail transit passenger flow based on a streaming computing engine

Through the distributed flow processing engine and multi-mode feature construction method, the problems of high-frequency sliding prediction and late data impact in rail transit systems are solved, and efficient and accurate network-level short-term passenger flow prediction is achieved.

CN119558460BActive Publication Date: 2025-07-22YUNNAN NORMAL UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411604589.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-12
Publication Date
2025-07-22
Estimated Expiration
2044-11-12

AI Technical Summary

Technical Problem

The short-term passenger flow prediction scheme of the existing rail transit system cannot meet the demand for high-frequency sliding prediction, the network-level passenger flow calculation efficiency is insufficient, and it fails to effectively deal with the impact of late data on real-time prediction accuracy.

Method used

Using a distributed stream processing engine-based method, a parallel computing algorithm is designed for network-level short-time passenger flow calculation, combining multi-rolling window sample construction and multi-modal feature construction, online feature engineering and prediction are controlled through triggers to alleviate the impact of late data on prediction accuracy.

Benefits of technology

It realizes efficient and real-time network-level short-time passenger flow calculation, improves the generalization ability and accuracy of the prediction model, and maintains high prediction accuracy in high-frequency sliding prediction scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119558460B_ABST
    Figure CN119558460B_ABST
Patent Text Reader

Abstract

The present invention relates to a real-time prediction method for rail transit passenger flow based on a streaming computing engine, belonging to the field of real-time passenger flow prediction. First, obtain the historical ticket card records of the rail transit system, divide them according to time windows, and calculate the short-term inbound and outbound passenger flows of each station. Secondly, in a batch processing environment, use multiple rolling windows to construct time series prediction samples based on the aggregated passenger flow data by sliding, train the prediction model and save it. After that, use Kafka and Spark-streaming to build a pipeline-style real-time computing environment, deploy the network-level short-term passenger flow calculation algorithm, feature management module and prediction control module. Finally, continuously process the real-time ticket card data stream and slide to predict the future short-term passenger flow. The present invention is used to solve the problem of real-time prediction of passenger flow in large-scale rail transit scenarios, meet the high-frequency sliding prediction requirements in streaming scenarios, effectively cope with the impact of late ticket card data, and maintain the latency within 4 seconds when processing 70,000 ticket card records per minute.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a real-time passenger flow prediction method for rail transit networks based on a distributed stream processing engine, belonging to the field of real-time passenger flow prediction. Background Art

[0002] Accurate and reliable short-term passenger flow prediction is the key to the intelligent application of rail transit systems. Rail transit systems operate at high speeds in a closed environment, serving millions of passengers daily. Especially during peak commuting hours, some subway stations face severe congestion, and certain lines or sections bear extremely high passenger flow pressure, posing great challenges to operation management. In this context, accurate short-term passenger flow prediction can provide important decision-making support for operation management, such as passenger flow guidance, train scheduling optimization, and emergency management, thereby improving operation efficiency and resource allocation effectiveness. Currently, many researchers have proposed short-term passenger flow prediction solutions based on big data technology. These methods can be roughly divided into two categories: one is offline prediction based on batch processing, and the other is real-time prediction based on stream processing. Compared with batch processing methods, stream processing can better meet the real-time requirements of passenger flow prediction and intelligent decision-making. Although existing solutions have to some extent solved the technical problems of short-term passenger flow prediction, there are still the following deficiencies:

[0003] (1) It cannot meet the requirements of high-frequency sliding prediction. In real application scenarios, we expect to obtain passenger flow dynamics at a higher frequency. Compared with rolling prediction, sliding prediction is more suitable for real-time passenger flow prediction due to its high-frequency characteristics. However, existing solutions all use the task scenario of rolling prediction, construct samples for training and evaluating the model based on passenger flow aggregated by rolling windows, resulting in insufficient generalization ability of the prediction model and low sliding prediction accuracy, and cannot meet the requirements of high-frequency sliding prediction. Therefore, there is still a large room for optimization in improving the model prediction accuracy in the sliding prediction scenario for existing solutions;

[0004] (2) The real-time passenger flow calculation efficiency is insufficient. Existing solutions often only calculate the passenger flow of a few stations, or only target a single entry and exit mode. They cannot efficiently calculate the entry and exit passenger flows of all stations in the entire rail transit network at the same time. Therefore, improving the efficiency of network-level entry and exit passenger flow calculation remains a problem to be solved;

[0005] (3) The impact of late data on real-time prediction is not fully considered. Due to the instability of data acquisition systems and networks, the card swiping records generated by subway turnstiles usually have delays when transmitted to the data center. These delayed data will cause errors between the real-time passenger flow calculated by the stream processing system and the actual passenger flow, and affect subsequent online feature engineering and prediction links, thereby reducing the accuracy of real-time prediction. Therefore, studying how to deal with the impact of late data on prediction accuracy can fill the current research gap.

[0006] For example, the patent "CN114118555A, A Distributed Subway Passenger Flow Prediction Method Based on Improved KNN" designs a method that combines the KNN and LightGBM algorithms for short-term passenger flow prediction in the subway scenario, and predicts the subway passenger flow in the next hour based on historical card-swipe data. This method uses big data components HDFS and Spark to expand the computing power of the solution, but it is only limited to the offline sliding prediction scenario and cannot meet the requirements of high-frequency real-time prediction. In addition, due to its relatively large time granularity, it cannot provide more fine-grained passenger flow information.

[0007] The patent "CN110276474A, A Method for Short-Term Passenger Flow Prediction at Urban Rail Transit Stations" uses the parallel computing engine Spark to train an Elman neural network for short-term passenger flow prediction in the subway scenario. The time granularity used in this solution is 15 minutes, which can provide more fine-grained passenger flow information compared to one hour. However, since this solution is a batch processing solution, it cannot meet the timeliness requirements of real-time prediction.

[0008] The patent "CN110782060A, A Method and System for Short-Term Prediction of Rail Transit Section Passenger Flow Based on Big Data Technology" designs a systematic solution for the short-term prediction scenario of rail transit section passenger flow. This solution uses HDFS to store historical ticket card data, historical passenger flow data, and real-time passenger flow data; uses the message middleware Kafka to access real-time card-swipe data; uses Spark to calculate the real-time section passenger flow, and predicts the section passenger flow in the next 5 minutes based on the real-time section passenger flow, historical section passenger flow, and historical inbound passenger flow. Although this solution uses stream processing technology to calculate and predict the real-time section passenger flow, it does not support the high-frequency sliding prediction requirement, cannot predict the inbound and outbound passenger flow of each station on a large spatio-temporal scale (rail transit network), and does not propose effective measures to deal with the impact of late data on the prediction accuracy. Summary of the Invention

[0009] The technical problem to be solved by the present invention is to perform real-time network short-term passenger flow prediction in the scenario where a large amount of passenger card-swipe ticket card data is continuously generated in the rail transit system, which can support some data-driven intelligent applications in the rail transit scenario, such as transfer recommendation, train scheduling, and emergency management.

[0010] The present invention analyzes the existing short-term passenger flow prediction solutions and their deficiencies. Specifically, there are still three deficiencies in the existing solutions:

[0011] (1) It does not support efficient large-scale network-level real-time passenger flow calculation. Most of the existing short-term passenger flow prediction schemes assume that the real-time short-term passenger flow is aggregated or provided by a third party; a few real-time short-term passenger flow calculation schemes propose methods for real-time calculation of short-term passenger flow, but the calculated short-term passenger flow is the station-level short-term passenger flow or short-term cross-section passenger flow, rather than the network-level in-out passenger flow.

[0012] (2) Only the rolling prediction scenario is considered. The existing short-term passenger flow prediction schemes only construct samples and train and evaluate models for the rolling prediction scenario. Although the models trained by this method can be used in the high-frequency sliding prediction scenario, there will be a problem of decreased prediction accuracy, and there is still a large room for improvement in the generalization ability of the models.

[0013] (3) The impact of late data on real-time prediction is ignored. When the existing schemes perform real-time prediction of short-term passenger flow, they do not consider the impact of late data on the real-time prediction effect. Specifically, late data will cause an error between the real-time calculated short-term passenger flow and the true value, and this error will further propagate along the data processing pipeline to the input features of the prediction model, thereby affecting the accuracy of real-time prediction and resulting in the online prediction accuracy being lower than the offline prediction.

[0014] The technical solution of the present invention is as follows: Based on the above analysis, a real-time passenger flow prediction method for rail transit networks based on a distributed stream processing engine is proposed. This method uses Spark as the distributed stream processing engine, and designs and implements a parallel calculation algorithm for efficiently calculating the network-level short-term in-out passenger flow, improving the efficiency and timeliness of real-time network-level short-term passenger flow calculation. A multi-rolling window sample construction method is designed. This method constructs samples based on the aggregated passenger flow by sliding, generating more samples and containing more detailed and rich passenger flow laws, which helps to improve the generalization ability of the model. A real-time multi-mode feature construction method is designed. Through the batch-stream fusion method, multi-mode passenger flow features including real-time mode, daily mode, and weekly mode are constructed by combining real-time and historical short-term passenger flow data. With the help of accurate daily mode and weekly mode features, the error propagated by late data to the multi-mode passenger flow features is diluted, thereby alleviating the impact of late data on the online prediction accuracy and improving the online prediction accuracy. A trigger-controlled online feature engineering and prediction method is designed. This method specifies the sliding interval , trigger interval and watermark size as the parameters for process control, and let , . In this way, the stream processing application will trigger calculations within each , and the previous The second calculation processes the late-arriving data belonging to the previous sliding time window, updates its calculation status, and improves the accuracy of real-time short-term passenger flow calculation and prediction. This method only performs online feature engineering and prediction for the second calculation, and can obtain the best timeliness in the first prediction within each , achieve the same accuracy as the offline prediction in the th prediction, and find the balance between timeliness and accuracy in the intermediate predictions.

[0015] The specific steps of a real-time passenger flow prediction method for rail transit based on a streaming computing engine are as follows:

[0016] Step1: Collect all historical smart card tapping records generated by the rail transit ticket card collection system to obtain a static offline card-swiping data set;

[0017] Step2: Perform data preprocessing on the offline card-swiping data set, and use a sliding window to calculate the offline network-level short-term inbound and outbound passenger flows according to the set time granularity and sliding interval;

[0018] Step3: Use the multi-mode feature sample construction method with multiple rolling windows to process the offline network-level short-term inbound and outbound passenger flows, and construct offline multi-mode feature samples;

[0019] Step4: Use the offline multi-mode feature samples to train and evaluate the pre-trained model;

[0020] Step5: Save the pre-trained model in a serialized manner;

[0021] Step6: Deploy a message middleware, a persistent key-value storage service, a streaming computing engine, a feature engineering controller, a prediction control module, and a real-time data simulation module to provide an available stream processing environment;

[0022] Step7: Use the real-time data simulation module to simulate a large amount of real-time card-swiping data streams generated by the rail transit system according to the parameter settings at the time of deployment, and connect them to the message middleware;

[0023] Step8: Use the streaming computing engine to obtain the real-time card-swiping data stream from the message middleware, perform data preprocessing, parallelly calculate the real-time network-level short-term inbound and outbound passenger flows, and write them into the persistent key-value storage service;

[0024] Step9: Input the real-time network-level short-term inbound and outbound passenger flows, and use the trigger control mechanism composed of the feature engineering controller and the prediction control module to control the process of online feature engineering and prediction, and output the prediction results;

[0025] Step10: Repeat the loop of Step7 - Step9 to form a continuous flow processing pipeline that includes real - time ticket card data stream access, data pre - processing, short - term passenger flow calculation, multi - mode feature construction, and passenger flow prediction.

[0026] Specifically, in Step1, the offline card - swiping dataset consists of historical card - swiping records of the rail transit system over a past period of time and is static. Each card - swiping record includes fields such as card number, transaction time, station name, in - out station flag, and transaction amount. Among them, only the transaction time, station name, and in - out station flag are the core fields required for network - level short - term passenger flow calculation and prediction.

[0027] Specifically, in Step2, the data pre - processing performed on the offline card - swiping dataset includes extracting the core fields of each card - swiping record, standardizing the station names, and mapping the standard stations to station numbers. To implement these pre - processing logics, we use Spark SQL for development and encapsulate it into a function. To standardize the processing flow and improve the performance of the program; when calculating the offline network - level short - term in - out station passenger flow using a sliding window with a specified time granularity and sliding interval, first group the pre - processed offline card - swiping dataset according to the sliding window, and use the parallel algorithm designed by the present invention for each group to calculate the offline network - level short - term in - out station passenger flow. The parallel algorithm is implemented with the help of Spark UDAF (User - Defined Aggregation Function), and its input is the card - swiping dataset generated by subway stations in the rail transit system within a time window , which has partitions distributed on different nodes. Let represent the partition number, then the specific implementation process of the parallel algorithm is as follows:

[0028] First, initialize a long integer array with a length of and an initial value of 0 on each partition, which is used to store intermediate calculation results. The first elements store the inbound passenger flow, and the last elements store the outbound passenger flow;

[0029] Then, traverse all the card - swiping data records in each partition in parallel, and calculate and update the long integer array according to the in - out station flag and station number of the card - swiping records;

[0030] After that, wait for all partitions to finish traversing the local partition datasets, and then collect the into a list storing long integer arrays ;

[0031] Finally, pairwise merge the long integer arrays in each partition in the list of long integer arrays by element-wise addition, and finally obtain the network-level short-term inbound and outbound passenger flows within the time window.

[0032] The logic of the above offline data preprocessing and calculating the offline network-level short-term inbound and outbound passenger flows is similar to the logic of the online data preprocessing and calculating the real-time network-level short-term inbound and outbound passenger flows in Step8. Therefore, functions and parallel algorithms can be reused in Step8. The difference is that the processing object of the offline preprocessing and short-term passenger flow calculation in Step2 is the static offline data set, while the processing object of the real-time preprocessing and short-term passenger flow calculation in Step8 is the dynamic unbounded data stream. Due to the characteristics of the unbounded data stream, such as being dynamic, unbounded, and the data stream may arrive out of order due to network latency, it is more difficult to process compared to the static offline data set. Therefore, a watermark mechanism needs to be introduced, and the output mode and trigger mechanism of the stream processing need to be considered.

[0033] Specifically, the method for constructing multi-mode feature samples with multiple rolling windows in Step3 is expressed as follows:

[0034] Let represent any moment within the operation time period of the rail transit system, represent the time granularity of the network-level short-term inbound and outbound passenger flows, represent the time window of the network-level short-term inbound and outbound passenger flows of the rail transit system, represent the number of historical short-term passenger flow records used when constructing the features of a single mode, represent the time interval between adjacent historical short-term passenger flow records when constructing the input features, represent 24 hours, represent the mode of passenger flow, and represent the real-time mode, daily mode, and weekly mode respectively. Then , , represent the network-level short-term inbound and outbound passenger flows of the real-time mode, daily mode, and weekly mode respectively, , , represent the passenger flow characteristics of the real-time mode, daily mode, and weekly mode respectively. The relationship between the network-level short-term inbound and outbound passenger flows of different modes and is as follows:

[0035]

[0036] Specified mode Corresponding passenger flow characteristics Are expressed as follows:

[0037]

[0038] Let Represent the multi-mode feature sample of the network-level short-term in-out passenger flow at time t:

[0039]

[0040] Let the number of stations in the rail transit system be , Represent the th station, Represent the in-out identification, Represent in, Represent out, then the network-level short-term in-out passenger flow Is a long integer array with a length of 2 , consisting of the short-term in passenger flow and short-term out passenger flow of each station. Specifically, let Represent the in-out identification of station within the time window be The station-level short-term passenger flow, and its value is a long integer variable, then the network-level short-term in passenger flow and the network-level short-term out passenger flow Are defined as follows:

[0041]

[0042] Then it is the network-level in-out passenger flow containing and (Unless otherwise specified, the short-term passenger flow mentioned below is the network-level in-out passenger flow , so it will be simply referred to as short-term passenger flow below), and can be expressed as follows:

[0043]

[0044] Let the total operating time of the rail transit system on a certain day be , the sliding time interval for calculating the short-term passenger flow be , then there is , , . That is Can be divided into by a rolling window of size non-overlapping time slices; It can also be divided into time slices with overlapping parts by a sliding time window of size and interval ; is times of . At this time, let , denote the multi-pattern feature sample of the th time slice. Then the sample set obtained based on the traditional feature engineering method is represented as follows:

[0045]

[0046] The sample set obtained by using the multi-pattern feature sample construction method based on multiple rolling windows is represented as follows:

[0047]

[0048] The multi-pattern feature sample construction method of multiple rolling windows is equivalent to dividing the short-term passenger flow of sliding aggregation into groups, and the time interval of the passenger flow records within each group is . Then, samples are constructed for each group according to the traditional feature engineering method to obtain the sample set of each group, and finally they are merged to obtain the complete sample set. The number of samples constructed by this method is about times that of the traditional method, which can provide more and finer-grained useful information for model training, improve the generalization ability of the pre-trained model, and improve its performance in the sliding prediction scenario.

[0049] Let the start operation time of a certain day of the rail transit system be , and the end operation time be , and the sliding time interval of sliding prediction be . Then the sliding prediction of the network-level short-term passenger flow can be defined as:

[0050]

[0051] To intuitively compare the differences between sliding prediction and rolling prediction, the definition of the rolling prediction of the network-level short-term passenger flow is as follows:

[0052]

[0053] Specifically, the difference between sliding prediction and rolling prediction is that the of rolling prediction, while the of sliding prediction can be specified more flexibly. When the When available, sliding prediction can be performed with the aid of a rolling prediction model. Compared with rolling prediction, sliding prediction can specify a sliding time interval and provide more frequent prediction information. At the same time granularity, the number of sliding predictions is times that of rolling prediction. If is fixed and the time granularity is larger, the sliding prediction provides richer prediction information compared with the rolling prediction.

[0054] Specifically, since the rail transit system usually stops operating at night, within the time range before every morning , it is impossible to construct complete multi-modal features. Therefore, we use the historical passenger flow records during a period of time before the rail transit system stopped operating the previous day ( ) to supplement the multi-modal features. Assume is the total length of the daily shutdown time of the rail transit system, and represents the time the rail transit system has been in operation at time . Then the construction formula of the multi-modal feature

[0055]

[0056] Specifically, in Step 4, when training and evaluating the model, use the offline multi-modal feature sample set constructed in Step 3, and take the specified date and time as the boundary to divide the sample set into a training set, a validation set, and a test set. Then train multiple prediction models on the training set, test the generalization ability of the models on the validation set, and evaluate the prediction performance of the models on the test set.

[0057] Specifically, in Step 5, select the prediction model with the best performance and save it as a binary file in a serialized manner and save it to the storage system.

[0058] Specifically, the steps for initializing the stream processing module in Step 6 are as follows:

[0059] Step 6.1: Cluster planning and construction. Allocate different numbers of nodes and hardware resources for the message middleware Kafka and the streaming computing engine Spark; install the corresponding software dependencies on each node, and set the IP addresses and ports used for component communication and other configurations to build a Kafka cluster and a Spark cluster.

[0060] Step 6.2: Deploy the message middleware. Create a Topic that accesses the real-time card swiping data stream and is consumed by Spark-Streaming according to the specified Topic (topic) name and the number of partitions.

[0061] Step6.3: Deploy the persistent key-value storage service. Enable the persistent key-value storage service, and pull and persist the short-term passenger flow data of the past week from the distributed storage system HDFS.

[0062] Step6.4: Deploy the streaming computing engine. Specify the parameters for running the real-time short-term passenger flow prediction task, and submit the Spark Structured Streaming application to the Spark cluster. At this time, the stream processing application will create the daemon process Driver of the application on the Driver node of the Spark cluster, load the code of the application in the Driver, and distribute it to the worker nodes for running.

[0063] Step6.5: Deploy the feature engineering controller. When the daemon process Driver loads and runs the code of the application, it will create an instance of the feature engineering controller of the feature engineering controller for subsequent feature engineering.

[0064] Step6.6: Deploy the prediction control module. When the daemon process Driver loads and runs the code of the application, it will deserialize and load the pre-trained model serialized and saved in Step5 into the memory of the Driver for caching, for subsequent real-time prediction.

[0065] Step6.7: Deploy the real-time data simulation module. Specify various parameters such as the date and start and end times for simulating historical card-swipe data, the load expansion multiple, the communication URL of Kafka, and the latency simulation function to start the real-time data simulation module.

[0066] After completing the above steps, the stream processing module is initialized, and each component will provide an available stream processing environment for the stream processing pipeline of continuous real-time short-term passenger flow prediction.

[0067] Specifically, the role of the feature engineering controller in Step 6 is to cache the historical and real-time network-level short-term inbound and outbound passenger flow data in the daemon process Driver of the application as the data source for online feature engineering, avoiding a large amount of cross-node network transmission (shuffle) caused by the second time window-based streaming aggregation. In addition, the feature engineering controller abstracts the logic of online feature engineering into 4 functions to standardize the process of online feature engineering and facilitate improvement and expansion. We designed a custom type (Class) to implement the feature engineering controller. When deploying the feature engineering controller, the daemon process Driver of the application will create an instance of the feature engineering controller for subsequent control processes. Specifically:

[0068] The feature engineering controller The encapsulated member variables are shown in Table 1. Among them, the real-time passenger flow cache and the historical passenger flow cache are respectively used to cache historical and real-time network-level short-term inbound and outbound passenger flow data. TreeMap[Long, Array[Long]] represents a sorted key-value pair set with a key of type Long and a value of an array of type Long. The remaining member variables are used to assist in the process control of online feature engineering, such as judging whether it is the current date for initialization , the current event time that stores the latest short-term passenger flow event time and so on.

[0069] Table 1 Important Member Variables of FEController

[0070]

[0071] Table 2 Core Functions of FEController

[0072]

[0073] The four encapsulated core member functions are shown in Table 2. The functions of each are as follows:

[0074] Initialization function : Based on the incoming timestamp (the event time of the latest network-level short-term inbound and outbound passenger flow), update auxiliary variables such as the current date , calculate the time range of the network-level short-term inbound and outbound passenger flow required to initialize the real-time passenger flow cache and the historical passenger flow cache , and query and load the corresponding data records from the persistent key-value storage service to initialize the two passenger flow caches.

[0075] Real-time Passenger Flow Data Cache Update Function : Pass in the real-time network-level short-term inbound and outbound passenger flow to update the real-time passenger flow cache . If a new short-term passenger flow record is appended, remove the oldest record in the real-time passenger flow cache .

[0076] Real-time Feature Time Series Calculation Function : Based on the incoming current event time calculate the time index sequence of the network-level short-term inbound and outbound passenger flow records required for real-time mode features (let be , then correspond to )

[0077] Multi - mode feature constructor : The role is to pass in (that is ), calculate and . And use to index data records from the real - time passenger flow cache to construct real - time mode passenger flow features ; from using and to index data records from the historical passenger flow cache to construct daily - mode and weekly - mode passenger flow features and ;

[0078] Specifically, the role of the predictive control module in Step 6 is to deserialize the pre - trained model serialized and saved in Step 5 in the daemon process Driver of the application, load it from the specified storage path into the memory of Driver for caching. At the same time, initialize the Onnx Runtime runtime environment for model inference in Driver to ensure that the pre - trained model trained in the Python environment can perform model inference in the Java programming environment.

[0079] Specifically, the role of the real - time data simulation module in Step 6 is to simulate the rail transit system to generate ticket - card data streams and write them into the message middleware. It is implemented by a multi - process program and does not depend on the stream processing application program. When starting, parameters such as the date and start - end time for simulating historical card - swiping data, the load expansion multiple, the communication URL of Kafka, and the delay simulation function need to be specified. The delay simulation function is used to simulate the total transmission delay of card - swiping data from the turnstile to the station proxy server and then forwarded to the data center. . To simulate , we designed the following delay simulation function:

[0080]

[0081] In the above formula, U(a, b) represents the uniform distribution function with the lower bound a and the upper bound b; represents generating a random number uniformly distributed between with a probability of .

[0082] After starting the real - time data simulation module with the specified parameters in Step 7, its working process is as follows:

[0083] Step 7.1: Data Loading. Read the card - swiping data within the specified date - time range according to the submitted parameters.

[0084] Step 7.2: Data Pre - processing. Expand the total amount of data in the dataset according to the load expansion multiple. Traverse each record in the "card - swiping time" column and add the simulated delay function randomly generated to generate a new column "writing time".

[0085] Step 7.3: Data Writing. Based on the "writing time" column, index the sub - dataset of the card - swiping dataset loaded in Step 7.2 second by second and write it into the message middleware Kafka.

[0086] Step 7.3.1: Initialization of Loop Variables. Let the start time of the simulated passenger flow writing specified by the parameter be and the end time of the simulated passenger flow writing be ;

[0087] Step 7.3.2: System Time Acquisition. Obtain the system time when entering a new round of loop ;

[0088] Step 7.3.3: Loop Condition Judgment. Judge whether is less than . If so, jump to Step 7.3.4; otherwise, jump to Step 7.6;

[0089] Step 7.3.4: Passenger Flow Interception. Filter the sub - dataset whose values in the "writing time" column are within ;

[0090] Step 7.3.5: Passenger Flow Writing. Enable a new auxiliary process to write the intercepted passenger flow data into the specified topic (Topic) in Kafka. When writing, the data will be partitioned first according to the partition strategy of the specified parameter "partition by site" or "random partition" to obtain the partition number, and then written into the corresponding partition of the Topic in Kafka;

[0091] Step 7.3.6: Processing Time Calculation. Obtain the current system time , calculate the processing time used in this round of loop ;

[0092] Step 7.3.7: Main Process Sleep. Make the main process sleep for time;

[0093] Step 7.3.8: Loop Variable Update. Let increase by one second, that is , and jump to Step 7.3.2;

[0094] Step 7.4: Program termination. Close the site simulator and terminate the program operation.

[0095] Specifically, the logic of real-time preprocessing and short-term passenger flow calculation in Step 8 is similar to that in Step 2. Therefore, the functions implemented using Spark SQL in Step 2 and the parallel passenger flow algorithm implemented using Spark UDAF in Step 2 are reused. The difference between Step 8 and Step 2 is as follows: Step 8 processes dynamic and unbounded data streams, with different data access methods, as well as different methods for processing and writing calculation results. The specific construction process is as follows: Step 8.1: The streaming computing engine obtains real-time card-swipe data streams from the message middleware at regular time intervals;

[0096] Step 8.2: Use

[0097] functions to preprocess the real-time card-swipe data streams; and group them based on sliding time windows; Step 8.3: Call the parallel algorithm

[0098] to process the card-swipe data within each group (sliding time window) and calculate the real-time network-level short-term inbound and outbound passenger flows; Step 8.4: Write the calculated real-time network-level short-term inbound and outbound passenger flows into the persistent key-value storage service.

[0099] Specifically, the persistent key-value storage service is deployed to store short-term passenger flow data for the past week and the real-time calculated network-level short-term inbound and outbound passenger flows. Its role is to serve as the data source for initializing the

[0100] operator or the checkpoint (Checkpoint) for process failure recovery. Execute operator or process failure recovery checkpoint (Checkpoint).

[0101] Specifically, the role of the trigger-controlled online feature engineering and prediction in Step 9 is to determine whether to perform online feature engineering and prediction based on the trigger control mechanism and perform process control. To implement the trigger control mechanism, we specify the watermark size and the stream processing trigger interval , and through setting , , to assist in process control, represents the number of times to trigger stream processing within each sliding time interval , represents the number of times to trigger stream processing within each sliding time interval The number of times of triggering online feature engineering and prediction. Each sliding time interval corresponds to a window size of and a sliding time interval of for the sliding time window.

[0102] Specifically, within each sliding interval , the stream processing application will trigger times of calculations, corresponding to micro-batches. However, only the first times of calculations will process the late data in the previous sliding window and update the calculation status. Therefore, this mechanism only performs online feature engineering and prediction on the first times of calculations within each . This method takes into account the impact of late data on the real-time prediction accuracy. By triggering multiple calculations within each to process the late data and update the calculation status, the first prediction within has the best timeliness but is most affected by late data, and the real-time prediction accuracy drops the most. Each subsequent prediction will introduce delay while improving the accuracy of real-time prediction until the th prediction within

[0103] Step9.1: Obtain the triggering time of the current micro-batch ;

[0104] Step9.2: Combine the instance of the feature engineering controller and the real-time network-level short-term inbound and outbound passenger flow , and sequentially determine whether it is necessary to initialize , and whether it is necessary to update the real-time passenger flow data cache in ;

[0105] Step9.2.1: Determine whether the instance of the feature engineering controller is initialized and is empty;

[0106] Step9.2.1.a: If is not initialized and is empty, then jump to Step7;

[0107] Step9.2.1.b: If is not initialized and is non-empty, then obtain the event time of the latest short-term passenger flow record in ​ , and execute the initialization function initialization , and then execute the real-time passenger flow cache update function Update real-time passenger flow cache ;

[0108] Step 9.2.1.c: If Initialize and If not empty, get The event time of the latest short-term passenger flow record in And with the current date Compare and determine whether the date is updated. If the date is updated, the initialization function needs to be executed initialization ; After that, regardless of whether the date is updated, the real-time passenger flow cache update function is executed Update real-time passenger flow cache ;

[0109] Step 9.2.1.d: If Initialize and If it is empty, skip to the next step directly;

[0110] Step 9.3: Example based on feature engineering controller The trigger time of the current micro-batch Calculate the number of calculations within the current sliding window , if we calculate the number of , then with the help of The encapsulated method performs online feature engineering, constructs real-time multi-modal features, and passes them to the pre-trained model cached in the predictive control module. Make real-time predictions and output prediction results;

[0111] Step 9.3.1: Example based on feature engineering controller The trigger time of the current micro-batch Calculate the number of calculations within the current sliding window ;

[0112] Step 9.3.2: Determine the number of calculations in the current window Is it less than the number of predictions within the window? ;

[0113] Step 9.3.2.a: If If not, jump to Step 7.

[0114] Step 9.3.2.b: If If it holds, online feature engineering and prediction are performed. The specific processing flow is as follows:

[0115] First, based on the instance of the feature engineering controller and the number of calculations within the current window judge whether the current sliding window is updated. If it is updated, update the current event time for subsequent process control.

[0116] Secondly, calculate the time index sequence of the short-term passenger flow records required for multi-mode features. Based on the current event time calculate the time index sequence required to construct the real-time mode . Assume the current event time corresponds to the time stamp , then the time index sequence is .

[0117] Then, construct multi-mode features in real time. Use the time index sequence to obtain the short-term passenger flow records required to construct the real-time mode passenger flow features from the real-time passenger flow cache and construct ; based on calculate and , and use and to obtain the short-term passenger flow records required for the daily mode and weekly mode passenger flow features from the historical passenger flow cache and construct and ; finally, splice to obtain the multi-mode feature .

[0118] Finally, call the pre-trained model cached by the Driver daemon process , and pass in the multi-mode features constructed in real time( , ) for online passenger flow prediction and output the real-time prediction results.

[0119] Step9.4: Perform loop control, jump to Step7, and start a new round of real-time short-term passenger flow prediction.

[0120] In any case, this processing flow will eventually return to Step7 to start a new round of processing. This is because the processing object of stream processing is the unbounded data stream that continuously arrives at the stream processing module. Therefore, the stream processing module needs to form a continuous stream processing pipeline including real-time ticket data stream access, data preprocessing, short-term passenger flow calculation, multi-mode feature construction, and passenger flow prediction.

[0121] The beneficial effects of the present invention are:

[0122] (1) The parallel computing algorithm for efficiently calculating the short-term inbound and outbound passenger flows at the network level proposed by the present invention can efficiently and scalably aggregate massive swiping data streams in real time, and simultaneously complete the calculation of short-term inbound passenger flows and short-term outbound passenger flows;

[0123] (2) The method for constructing multi-rolling window samples proposed by the present invention constructs samples based on short-term passenger flows with sliding aggregation. The sample data set can provide more detailed and rich passenger flow rules for model training, and can improve the generalization ability and prediction effect of the prediction model in high-frequency sliding prediction scenarios;

[0124] (3) The method for constructing real-time multi-mode features proposed by the present invention can further dilute the error of late data propagated into multi-mode passenger flow features compared with the single-mode feature construction method, thereby alleviating the impact of late data on the online prediction accuracy and improving the online prediction accuracy.

[0125] (4) The online feature engineering and prediction method controlled by a trigger proposed by the present invention fully considers the influence of late data, performs multiple predictions for the same sliding window, and facilitates users to balance the real-time performance and accuracy of predictions. Description of the Drawings

[0126] Figure 1 is the flowchart of the steps of the present invention;

[0127] Figure 2 is the schematic diagram of the real-time short-term passenger flow prediction pipeline constructed by the present invention;

[0128] Figure 3 is the schematic diagram of the present invention for constructing multi-mode passenger flow features in real time and making predictions;

[0129] Figure 4 is the graph of the first online prediction accuracy loss within the sliding window corresponding to different time granularities;

[0130] Figure 5 is the graph of the passenger flow accuracy and the online prediction accuracy loss corresponding to multiple prediction cycles within the sliding window. Detailed Embodiments

[0131] The present invention will be further described below in combination with the drawings and detailed embodiments on the basis of the description of the specification.

[0132] As Figure 1 shown, a real-time passenger flow prediction method for rail transit networks based on a distributed stream processing engine includes the following steps:

[0133] Step1: Collect all historical smart card swiping records generated by the rail transit ticket collection system to obtain a static offline swiping data set;

[0134] Step 2: Preprocess the offline card-swipe dataset, and calculate the short-term inbound and outbound passenger flows at the offline network level using a sliding window according to the set time granularity and sliding interval;

[0135] Step 3: Process the short-term inbound and outbound passenger flows at the offline network level using a multi-mode feature sample construction method with multiple rolling windows to construct offline multi-mode feature samples;

[0136] Step 4: Train and evaluate the pre-trained model using the offline multi-mode feature samples;

[0137] Step 5: Save the pre-trained model in a serialized manner;

[0138] Step 6: Deploy a message middleware, a persistent key-value storage service, a streaming computing engine, a feature engineering controller, a prediction control module, and a real-time data simulation module to provide an available stream processing environment;

[0139] Step 7: Use the real-time data simulation module to simulate a large amount of real-time card-swipe data streams generated by the rail transit system according to the parameter settings at the time of deployment, and connect them to the message middleware;

[0140] Step 8: Use the streaming computing engine to obtain the real-time card-swipe data streams from the message middleware, perform data preprocessing, parallelly calculate the short-term inbound and outbound passenger flows at the real-time network level, and write them into the persistent key-value storage service;

[0141] Step 9: Input the short-term inbound and outbound passenger flows at the real-time network level, and use a trigger control mechanism composed of a feature engineering controller and a prediction control module to control the process of online feature engineering and prediction, and output the prediction results;

[0142] Step 10: Repeat Steps 7 - 9 in a loop to form a continuous stream processing pipeline including real-time ticket card data stream access, data preprocessing, short-term passenger flow calculation, multi-mode feature construction, and passenger flow prediction.

[0143] Figure 1 The processing flow at the bottom corresponds to Steps 1 to 5 in the above steps, which corresponds to the offline batch processing part and will only be executed once. Its function is to train the serialized pre-trained model required when deploying the prediction control module in Step 6. Figure 1The processing flow at the top corresponds to Step 6 to Step 10 in the above steps, which corresponds to the part of real-time stream processing. Among them, Step 6 is used to initialize the stream processing module, providing an available stream processing environment for the subsequent real-time short-term passenger flow prediction pipeline, and it will only be executed once; Step 7 to Step 10 will be executed repeatedly to form a continuous stream processing pipeline for real-time short-term passenger flow prediction. This pipeline is as Figure 2 shown.

[0144] As Figure 2 shown, Figure 2 from left to right and from bottom to top in it, the continuous stream processing pipeline for real-time short-term passenger flow prediction includes the following components and modules: real-time data simulation module, message middleware Kafka, distributed stream processing engine Spark-Streaming, and persistent key-value storage service. The interaction method between them is as follows:

[0145] First of all, the real-time data simulation module simulates the card swiping data within a specified historical time range, and writes the real-time card swiping data stream into the message middleware Kafka every second according to the specified data load multiple and simulation delay;

[0146] Secondly, after the message middleware Kafka receives the new real-time card swiping data stream, it notifies the stream processing engine Spark-Streaming to consume it; after Spark-Streaming sends a consumption request to Kafka, Kafka forwards the real-time card swiping data stream corresponding to the request information and converts it into a streaming structured data set (DataFrame) corresponding to a micro-batch;

[0147] After that, Spark-Streaming uses function to preprocess the streaming DataFrame composed of the real-time card swiping data stream of the current micro-batch, and groups the preprocessed streaming DataFrame with a sliding time window, and calls the parallel algorithm to calculate the real-time network-level short-term inbound and outbound passenger flow , and writes it into the persistent key-value storage service.

[0148] When initializing the stream processing module in Step 6, the stream processing application is submitted to the Spark cluster. The Spark cluster will start the daemon process Driver on the Driver node and create instantiated object . When the continuous stream processing pipeline for real-time short-term passenger flow prediction starts to work, the instance of the feature engineering controller is not initialized. At this time, if the real-time network-level short-term inbound and outbound passenger flow is not empty, it will obtain the event time of the latest network-level short-term inbound and outbound passenger flow records in , and execute the initialization function ( ) Calculate the real-time passenger flow cache and the historical passenger flow cache the time range of the historical short-term passenger flow records required, and initiate a data query request to the persistent key-value service, and pull the query results into the Driver process for initialization . In addition, when the Driver node of the stream processing application fails, restart the stream processing program, and it will still re-perform initialization, and at this time, the persistent key-value storage service is equivalent to an instance of the feature engineering controller for the checkpoint (Checkpoint) of fault recovery.

[0149] Finally, once the instance of the feature engineering controller is initialized, after Spark-Streaming calculates the real-time network-level short-term inbound and outbound passenger flow , and writes it into the persistent key-value storage service, it will collect across nodes to the Driver node, and execute the real-time passenger flow cache update function to update the real-time passenger flow cache , and construct multi-mode features based on the real-time passenger flow cache and the historical passenger flow cache , and finally call the pre-trained model cached in the prediction control module to predict the real-time network-level short-term inbound and outbound passenger flow within the next time interval . This process is as shown. Figure 3

[0150] The present invention will be described in detail below with specific experimental data.

[0151] The implementation process of the present invention is divided into 7 parts, namely: development environment setup, offline preprocessing and short-term passenger flow calculation, offline multi-mode feature sample construction, model training evaluation and saving, stream processing module initialization, continuous stream processing, real-time prediction performance evaluation.

[0152] In the development environment setup stage, it is mainly used to install the development environment required for offline batch processing and real-time stream processing, and the process is as follows:

[0153] ​Plan an eight-node cluster for installing Kafka cluster, Spark cluster, persistent key-value storage service, and real-time data simulation module. Among them, the Spark cluster consists of 1 master node and 3 slave nodes. In the experiment, we plan to start the daemon process Driver of the Spark application on the master node, which serves as the core controller of the Spark application, manages task scheduling, cluster communication, and result collection. The Spark cluster is used to complete offline data preprocessing and sample construction, real-time data preprocessing, real-time network-level short-term inbound and outbound passenger flow calculation, online feature engineering, and real-time passenger flow prediction. Spark natively uses HDFS as the distributed file system by default. In addition, in order to train the pre-trained model in the Python environment and use the pre-trained model for online inference in the Java environment, we additionally installed the AI engine Pytorch and the inference engine Onnx Runtime environment on the Driver node of the Spark cluster; to ensure the fault tolerance of the feature engineering control module in the Driver process, we also deployed the lightweight database SQLite on the Driver node as the persistent key-value storage service. The Kafka cluster consists of 3 nodes, which is used to receive and cache the real-time card-swipe data stream and notify the SparkStructured Streaming application deployed on the Spark cluster to consume the real-time card-swipe data stream. The remaining 1 node is used to deploy the real-time data simulation module, which simulates the real-time card-swipe data streams generated by multiple subway stations in a multi-process manner and writes them into Kafka.

[0154] In the offline preprocessing and short-term passenger flow calculation stage, mainly preprocess the offline card-swipe dataset and calculate the offline network-level short-term inbound and outbound passenger flow. The process is as follows:

[0155] Collect the smart card swipe records of a city's rail transit system for four months to obtain the offline card-swipe dataset and upload it to the distributed file system HDFS cluster;

[0156] Use the Spark offline batch program to read all the swipe records in HDFS and use the preprocessing function implemented by Spark SQL , extract the four fields of card number, transaction time, station name, and inbound / outbound identifier from each swipe record, standardize the station name, and map the station name to the station number to complete the offline preprocessing; then, according to the time granularity of , and the sliding frequency of , group the preprocessed card-swipe dataset by the sliding time window, and call the parallel algorithm to aggregate the card-swipe data of each group to obtain the network-level short-term inbound and outbound passenger flow and write it into HDFS.

[0157] Parallel algorithm The input is the card - swiping data set generated by stations in the rail transit system within a time window , which has sub - regions distributed at different nodes. Let represent the sub - region number, then the specific implementation steps of this parallel algorithm are as follows:

[0158] First, initialize a long - integer array with a length of and an initial value of 0 on each sub - region , which is used to store intermediate calculation results. The first elements store the inbound passenger flow, and the last elements store the outbound passenger flow;

[0159] Then, traverse all the card - swiping data records in each sub - region in parallel, and calculate and update the long - integer array according to the inbound / outbound flag and station number of the card - swiping record;

[0160] After that, wait until all sub - regions have traversed the data sets of their local sub - regions, and then collect the from each sub - region across nodes through the network to the list where the long - integer array is stored;

[0161] Finally, pairwise merge the long - integer arrays of each sub - region in the list of long - integer arrays in an element - by - element addition manner , and finally obtain the network - level short - term inbound / outbound passenger flow within the time window.

[0162] In this embodiment, the number of stations in the rail transit system is 166, the number of sub - regions is 192, and the sliding time interval is 1 minute. In order to compare the impact of late - arriving data on the real - time prediction accuracy at different time granularities during the continuous - flow processing and performance testing phases, multiple groups of time granularities (i.e., 1, 3, 5, 10, 15 minutes) are used to calculate the network - level short - term inbound / outbound passenger flow.

[0163] In the offline multi - mode feature sample construction phase, mainly perform multi - mode feature sample construction based on multiple rolling windows, and the process is as follows:

[0164] Pull the network - level short - term inbound / outbound passenger flow calculated in the offline pre - processing and short - term passenger flow calculation and construction phase from the HDFS cluster of the distributed file system to the Python environment of the Driver node;

[0165] In the Python environment of the Driver node, perform multi - mode feature sample construction based on multiple rolling windows to obtain the sample data set 。

[0166] Divide the sample data set into a training set , a validation set and a test set . When making the division, the division method cannot simply and roughly use the proportional division method, because the sample order is the splicing of time slices divided by multiple rolling window orders, so it is necessary to specify the date and time for splitting to filter the samples.

[0167] In this embodiment, we use the samples of the last week as , the samples of the penultimate week as , and the remaining data as . That is, the date and time for splitting the sample set is from Monday to 00:00 of the last week and the penultimate week.

[0168] In the training, evaluation, and saving stage of the model, the pre-trained model is mainly trained, evaluated, and saved. The process is as follows:

[0169] Select a prediction model. Select multiple open-source and mature short-term passenger flow prediction models, and initialize the network structure and model parameters in the AI engine Pytorch. In this embodiment, 5 classic short-term passenger flow prediction models are selected for experiments. These 5 models are: GCN, GRU, TGCN, ResLSTM, and Conv-GCN;

[0170] The AI engine Pytorch traverses the sample training set constructed in the offline multi-mode feature sample construction stage in a batch-loading manner , and trains the prediction model to update the model parameters. The training is divided into multiple epochs, and each epoch will use to evaluate the generalization ability of the model and update the learning direction of the model.

[0171] After the training is completed, multiple short-term passenger flow prediction models can be obtained. At this time, evaluate and compare multiple prediction models. Here, 4 metrics are used to evaluate the difference between the real network-level short-term passenger flow and the predicted network-level short-term passenger flow . Assuming is the number of samples, the 4 selectable metrics can be formulated as:

[0172]

[0173] Among them, Accuracy is the accuracy rate;

[0174]

[0175] Among them, RMSE is the root mean square error;

[0176]

[0177] Among them, MAE is the mean absolute error;

[0178]

[0179] Among them, WMAPE is the weighted mean absolute percentage error;

[0180] We used the above four evaluation indicators on the above five classic short-term passenger flow prediction models to compare the prediction performance of the models trained by the traditional feature engineering method and the multi-mode feature sample construction method based on multiple rolling windows. The experimental data obtained are as follows:

[0181] The prediction performance of all models trained by the traditional feature engineering method decreased in the sliding scenario compared with the rolling scenario. The accuracy decreased by at least 0.4%, the RMSE increased by at least 1.55, and the MAE increased by at least 0.77. Among them, the ResLSTM model had the most significant decrease, with its accuracy decreasing by 1.07%, the RMSE increasing by 3.74, and the MAE increasing by 1.79. This shows that the models trained by the traditional feature engineering method lack finer-grained learning, have weak model generalization ability, and there is performance loss when facing higher-frequency sliding prediction tasks;

[0182] The prediction accuracy of all models trained by the multi-mode feature sample construction method based on multiple rolling windows has improved compared with the traditional feature engineering method in both the rolling prediction and sliding prediction scenarios. The accuracy of rolling prediction has increased by at least 0.82%, and the sliding prediction has increased by at least 1.16%. The TGCN model has the most obvious improvement effect, with its rolling prediction accuracy increasing by 1.74% and the sliding prediction accuracy increasing by 2.27%. This shows that the multi-mode feature sample construction method based on multiple rolling windows can effectively improve the generalization ability of the model and improve the prediction effect of the model in high-frequency prediction scenarios;

[0183] Among all the prediction models, ConvGCN achieved the best prediction performance. Specifically, the accuracy of ConvGCN for rolling prediction on the sample test set in the last week was 88.95%, the MAE was 21.58, the RMSE was 38.42, and the WMAPE was 9.90%; the accuracy for sliding prediction was 89.00%, the MAE was 21.55, the RMSE was 38.29, and the WMAPE was 9.88%;

[0184] After evaluating multiple short-term passenger flow prediction models, the ConvGCN model with the best offline sliding prediction performance was selected as the pre-trained model and convert it into a binary file in a serialized manner and save it to the specified file path in the distributed file system HDFS.

[0185] In the initialization stage of the stream processing module, the stream processing module is mainly initialized to provide an available stream processing environment for the real-time short-term passenger flow prediction pipeline in the subsequent stages.

[0186] Deploy the message middleware. Create a Topic for real-time access to the card swiping data stream according to the specified Topic (theme) and the number of partitions for Spark-Streaming to consume.

[0187] Deploy the persistent key-value storage service. Enable the persistent key-value storage service to pull and persist the short-term passenger flow data of the past week from the distributed storage system HDFS.

[0188] Deploy the streaming computing engine. Specify the parameters for running the real-time short-term passenger flow prediction task and submit the Spark Structured Streaming application to the Spark cluster. At this time, the stream processing application will create a daemon process Driver of the application on the Driver node of the Spark cluster, load the code of the application in the Driver, and distribute it to the worker nodes Woker for running.

[0189] Deploy the feature engineering control module. When the daemon process Driver loads and runs the code of the application, it will create an instance of the feature engineering controller for subsequent feature engineering.

[0190] Deploy the prediction control module. When the daemon process Driver loads and runs the code of the application, it will deserialize and load the pre-trained model saved in serialization into the memory of the Driver for caching for subsequent real-time prediction.

[0191] Deploy the real-time data simulation module. Start the real-time data simulation module by specifying parameters such as the date and start and end times for simulating historical card swiping data, the load expansion multiple, the communication URL of Kafka, and the delay simulation function.

[0192] After the above deployments are completed, the stream processing module is initialized and can provide an available stream processing environment for the real-time short-term passenger flow prediction pipeline in the subsequent stages.

[0193] In the continuous stream processing stage, it is mainly used to construct a real-time short-term passenger flow prediction pipeline, continuously process the continuously arriving real-time card swiping data stream, and perform short-term passenger flow prediction. The process is as follows:

[0194] Construct a real-time short-term passenger flow prediction pipeline:

[0195] Real-time card swiping data access: The real-time data simulation module generates a large amount of real-time card swiping data streams according to the parameters specified during deployment to simulate the rail transit system and accesses them to the message middleware;

[0196] When specifying parameters for the real-time data simulation module, this embodiment specifies two data multiples: 1x data multiple and 5x data multiple. The former simulates the normal card swiping data stream load, and the latter simulates the high-voltage card swiping data stream load; In addition, when simulating the parameters of the delay function in this embodiment, let , , , , to simulate the delay distribution with a maximum late arrival time within 30 seconds 。This distribution, there is a 50% probability that the network transmission is normal, and the delay is evenly distributed between 0.5 seconds and 3 seconds; There is a 50% probability that the network is abnormal, and a random delay evenly distributed between 3 seconds and 27 seconds is added on the basis of the normal transmission delay. This is more lagging and unstable compared to the transmission delay in the real application scenario. If our invention can work properly in this scenario, then in a scenario with a more stable network environment, our invention can achieve better performance;

[0197] Real-time preprocessing and short-term passenger flow calculation: The streaming computing engine obtains the real-time card swiping data stream from the message middleware, performs data preprocessing, and then calculates the real-time network-level short-term inbound and outbound passenger flows in parallel and writes them into the persistent key-value storage service;

[0198] This process uses the function for data preprocessing and the parallel algorithm for calculating the network-level short-term inbound and outbound passenger flows in the offline preprocessing and short-term passenger flow calculation 。The difference from the offline preprocessing and short-term passenger flow calculation is that: the processing object of the real-time preprocessing and short-term passenger flow calculation is the unbounded data stream, and it is necessary to set watermarks to control the tolerance of the late arrival time of the real-time card swiping data stream, set the output mode of the real-time network-level short-term inbound and outbound passenger flows, and the trigger time interval for triggering the calculation 。

[0199] In this embodiment, the watermark refers to the distribution of the total transmission delay and is designed to be 30 seconds to ensure that all late data can be processed by the stream processing engine. Assuming that the event time of the latest network-level short-term inbound and outbound passenger flow at the last trigger of the calculation is ,then each time the stream processing engine triggers the calculation, it will calculate the water level line according to If the event time of the late data is less than the watermark , then late data will not be processed; if the event time of late data is higher than the watermark , then the late data will be processed and the corresponding state (i.e. calculation result) will be updated at the same time. , stream processing can be triggered more frequently, caching smaller processing delays at the cost of higher resource overhead. To balance resource overhead and processing delay, this embodiment sets 7.5 seconds.

[0200] Trigger-controlled online feature engineering and prediction: input the real-time network-level short-term inbound and outbound passenger flow corresponding to the current micro-batch , and based on the example of feature engineering controller The trigger control mechanism composed of the prediction control module controls the process of online feature engineering and prediction and outputs the prediction results;

[0201] In this embodiment, because the sliding time interval The trigger interval is 1 minute. is 7.5 seconds, so 8 calculations are triggered per minute, each calculation corresponds to a micro-batch. After being initialized, each micro-batch calculates a non-empty real-time network-level short-term inbound and outbound passenger flow When , but only the first 5 micro-batches per minute will be used for online feature engineering and prediction;

[0202] The process of controlling online feature engineering and prediction for a single micro-batch is as follows:

[0203] Get the trigger time of the current micro-batch ;

[0204] Example of combining feature engineering controller and real-time network-level short-term inbound and outbound passenger flow , determine whether initialization is needed , and whether an update is required Real-time passenger flow data cache in ;

[0205] Example of feature engineering controller The trigger time of the current micro-batch Calculate the number of calculations within the current sliding window , if we calculate the number of , then with the help of The encapsulated method performs online feature engineering, constructs real-time multi-mode features, and inputs them into the pre-trained model cached in the prediction control module for real-time prediction, and outputs the prediction results;

[0206] Jump to the steps of real-time passenger flow access, and repeat the above processing flow in a loop to form a stream processing pipeline including real-time ticket card data stream access, data preprocessing, short-term passenger flow calculation, multi-mode feature construction, and passenger flow prediction.

[0207] The real-time prediction performance evaluation stage is mainly used to test the real-time prediction performance of the present invention in the presence of late data, as well as the data processing ability of the real-time short-term passenger flow prediction pipeline constructed by the present invention. The process is as follows:

[0208] Test the real-time prediction performance of this solution in the scenario of late data:

[0209] In order to compare the differences between the single-mode feature model and the multi-mode notification model in the real-time prediction scenario, this embodiment selects the ConvGCN, the best-performing short-term passenger flow prediction model in the offline sliding prediction scenario, as the pre-trained model, and implements the single-mode feature model (SPF-Model) and the multi-mode feature model (MPF-Model) of ConvGCN respectively. Let the difference between the real-time passenger flow prediction accuracy and the offline passenger flow prediction accuracy be the real-time prediction accuracy loss. Then, when the first real-time short-term passenger flow prediction is performed per minute, the influence of late data on the real-time short-term passenger flow accuracy prediction is the greatest. Therefore, we selected the real-time network-level short-term in-out passenger flow accuracy, the accuracy of real-time single-mode features, the accuracy of real-time multi-mode features, the real-time prediction accuracy loss of the SPF-Model, and the real-time prediction accuracy loss of the MPF-Model for the first real-time short-term passenger flow prediction per minute under different time granularities for data visualization, specifically as Figure 4 shown.

[0210] As Figure 4 shown, there are 5 broken lines in the figure. They respectively correspond to the accuracy of real-time network-level short-term in-out passenger flow, the accuracy of real-time single-mode features, the accuracy of real-time multi-mode features, the real-time prediction accuracy loss corresponding to the single-mode feature model, and the real-time prediction accuracy loss corresponding to the multi-mode feature model. It can be found that under the same late data distribution, the smaller the time granularity, the lower the accuracy of the real-time network-level short-term passenger flow, and the lower the accuracy of the real-time input features of the prediction model, resulting in the lower the accuracy of real-time prediction. When the time granularity is 1 minute, the accuracy of the real-time network-level short-term passenger flow is the lowest at 85.612%, and the real-time prediction accuracy losses of the SPF-Model and the MPF-Model are also the largest, at 1.977% and 1.504% respectively. In addition, it is not difficult to find that the multi-mode features can further reduce the error of late data propagated to the prediction model compared with the single-mode features, thereby alleviating the real-time prediction accuracy loss.

[0211] To evaluate the effectiveness of the trigger control mechanism in balancing real-time prediction accuracy and timeliness, this embodiment visualizes the 5 real-time prediction performances per minute of the MPF-Model of ConvGCN at different time granularities, specifically as Figure 5 shown. Figure 5 There are 10 broken lines in it. The 5 broken lines with solid markers represent the real-time prediction accuracy losses of the first 5 online predictions per minute at different time granularities. The 5 broken lines with hollow markers represent the changes in real-time passenger flow accuracy of the first 5 online predictions per minute at different time granularities. It can be found that at different time granularities, the real-time prediction accuracy loss of the first prediction per minute is the largest, and the real-time prediction accuracy loss of the 5th prediction is 0, that is, the real-time prediction accuracy is equal to the offline prediction accuracy. Suppose we believe that when the real-time prediction accuracy loss is less than 0.1%, the accuracy and timeliness of short-term passenger flow prediction can be balanced. Then when the time granularity is 1 minute, the balance point of the accuracy and timeliness of short-term passenger flow prediction can be obtained at the 3rd prediction per minute; when the time granularity is 3 minutes, the balance point of the accuracy and timeliness of short-term passenger flow prediction can be obtained at the 2nd prediction per minute.

[0212] Evaluating the data processing ability of the real-time short-term passenger flow prediction pipeline:

[0213] We use two metrics, latency and throughput, to evaluate the data processing ability of the real-time short-term passenger flow prediction pipeline. Latency represents the time interval from receiving an event to observing the event processing effect in the output. Since latency is generally small and sensitive to system jitter, latency is generally described in percentile. Suppose the processing latency of the stream processing application is , then includes the latency of the stream processing engine reading data from Kafka , the latency of real-time passenger flow calculation , the latency of writing real-time passenger flow into the database , the latency of online feature engineering and the latency of online prediction . The corresponding calculation formula is as follows:

[0214]

[0215] It should be noted that it is necessary to distinguish from the latency of real-time passenger flow prediction . Specifically, suppose the prediction of is completed at moment, then the business latency of this prediction .

[0216] Throughput is an indicator to measure the data processing capacity per unit time, representing the number of records that can be processed per unit time in each link of continuous stream processing. Generally speaking, the lower the latency, the better, and the greater the throughput, the better. These two indicators are not independent of each other. In extreme cases, when the load is very small, the latency of stream processing is extremely small, but the throughput is very low; when the load is very large, the throughput of stream processing is extremely high, but the newly arrived events will be backlogged in the message queue, resulting in a sharp increase in the latency of stream processing. To better test the data processing capacity of the real-time short-term passenger flow prediction pipeline, in this embodiment, the historical short-term passenger flow from 08:04 to 08:54 on a certain day in the test set is selected for playback, and the highest generation frequency of the card-swipe data within this time period can reach 14,000 records per minute. In the test scenario of pressure load, the highest generation frequency of the card-swipe data within this time period can reach 70,000 records per minute.

[0217] When the program is running stably, under the load of 14,000 records per minute, 95% of the is within 3.426 seconds; under the load of 70,000 records per minute, 95% of the is within 3.487 seconds; it can be found that even if the load is increased by 5 times, the stream processing latency does not increase suddenly, which indicates that our invention has good scalability and high efficiency, and also has the potential to handle larger real-time card-swipe data loads.

[0218] Through the above data, it can be shown that a real-time passenger flow prediction method for rail transit based on a streaming computing engine in the field of real-time passenger flow prediction is far higher than the existing methods. It can efficiently and scalably calculate the short-term inbound and outbound passenger flows of the network drama by aggregating massive card-swipe data streams in real time; it can provide the prediction performance of the prediction model in the sliding prediction scenario; it can efficiently implement online feature engineering, construct multi-mode features in real time and avoid a large amount of cross-node network transmission of the second streaming aggregation; it can alleviate the loss of real-time prediction accuracy caused by late data by means of real-time multi-mode features; and it can balance the timeliness and accuracy of real-time prediction through the trigger control mechanism. It can efficiently support real-time network-level short-term passenger flow prediction in a real rail transit scenario with complex network scenarios and late data.

[0219] The specific implementation manners of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above implementation manners, and various changes can be made without departing from the spirit of the present invention within the knowledge scope of those of ordinary skill in the art.

Claims

1. A real-time rail transit passenger flow prediction method based on a streaming computing engine, characterized in that: Step1: Collect all historical smart card tapping records generated by the rail transit ticket card collection system to obtain a static offline card-swiping data set; Step2: Perform data preprocessing on the offline card-swiping data set, and calculate the offline network-level short-term inbound and outbound passenger flows using a sliding window according to the set time granularity and sliding interval; Step3: Use a multi-mode feature sample construction method with multiple rolling windows to process the offline network-level short-term inbound and outbound passenger flows to construct offline multi-mode feature samples; Step4: Use the offline multi-mode feature samples to train and evaluate the pre-trained model; Step5: Save the pre-trained model in a serialized manner; Step6: Deploy a message middleware, a persistent key-value storage service, a streaming computing engine, a feature engineering controller, a prediction control module, and a real-time data simulation module to provide an available stream processing environment; Step7: Use the real-time data simulation module to simulate a large amount of real-time card-swiping data streams generated by the rail transit system according to the parameter settings during deployment, and connect them to the message middleware; Step8: Use the streaming computing engine to obtain the real-time card-swiping data stream from the message middleware, perform data preprocessing, parallelly calculate the real-time network-level short-term inbound and outbound passenger flows, and write them into the persistent key-value storage service; Step9: Input the real-time network-level short-term inbound and outbound passenger flows, and use a trigger control mechanism composed of a feature engineering controller and a prediction control module to control the online feature engineering and prediction process, and output the prediction results; Step10: Repeat Steps 7 - 9 in a loop to form a continuous stream processing pipeline including real-time ticket card data stream access, data preprocessing, short-term passenger flow calculation, multi-mode feature construction, and passenger flow prediction; The specific content of Step9 is as follows: The trigger control mechanism specifies the watermark size for distribution based on late data and the stream processing trigger interval , by setting , , , to achieve auxiliary process control. Within each sliding interval , the stream processing application triggers computations, corresponding to micro-batches, but only the first computations will update the computational state of the previous window. Therefore, only the first computations are used for online feature engineering and prediction. The specific implementation steps are as follows: Step9.1: Obtain the trigger time of the current micro-batch ; Step9.2: Combine the instance of the feature engineering controller and the real-time network-level short-term inbound and outbound passenger flow , and successively determine whether initialization is required , and whether it is necessary to update the real-time passenger flow data cache in ; Step 9.3: Example based on feature engineering controller The trigger time of the current micro-batch Calculate the number of calculations within the current sliding window , if we calculate the number of , then with the help of The encapsulated method performs online feature engineering, constructs real-time multi-modal features, and passes them into the pre-trained model cached in the prediction control module for real-time prediction and outputs the prediction results; Step9.4: Perform loop control, jump to Step7, and start a new round of real-time short-term passenger flow prediction.

2. The real-time rail transit passenger flow prediction method based on a streaming computing engine according to claim 1, wherein The specific content of Step3 is as follows: Let represent any moment within the operation time period of the rail transit system, represent the time granularity of the short-term in-out passenger flow at the network level, represent the time window the short-term in-out passenger flow at the network level of the rail transit system within, represent the number of historical short-term passenger flow records used when constructing the features of a single pattern, represent the time interval between adjacent historical short-term passenger flow records when constructing the input features, represent 24 hours, represent the pattern of passenger flow, and represent the real-time pattern, daily pattern and weekly pattern respectively. The relationship between the short-term in-out passenger flow at the network level in different patterns and is as follows: ; Among them, , , respectively represent the short-term inbound and outbound passenger flows at the network level in real-time mode, daily mode, and weekly mode; Specified mode Corresponding passenger flow characteristics Are shown as follows: ; Let represent the multi-modal feature sample of the short-term inbound and outbound passenger flow at the network level at the predicted time ; Among them, , , respectively represent the passenger flow characteristics in real-time mode, daily mode, and weekly mode; Let the total operating time of the rail transit system be , the sliding window size for calculating short-term passenger flow is , the sliding interval is ,have , , the size is , the interval is The sliding time window will Divide into There are time slices with overlapping segments; Indicating the multi-modal feature sample of the nth time slice, , the sample set obtained by the multi-modal feature sample construction method based on multiple rolling windows is represented as follows: ; Since the rail transit system shuts down at night, within the time range before every morning it is impossible to construct a complete multi-modal feature. Therefore, the historical passenger flow records during a period before the shutdown of the rail transit system on the previous day are used to supplement the multi-modal feature. Then, the construction formula of the multi-modal feature is defined as follows: ; Among them, is the total daily outage time of the rail transit system, represents the time that the rail transit system has been in operation at a certain moment.

3. A real-time prediction method for rail transit passenger flow based on a streaming computing engine according to claim 1, characterized in that, The specific content of Step7 is as follows: Step7.1: Set multiple parameters such as the date and start and end times of simulating historical card-swiping data, the load expansion multiple, the communication URL of the message middleware, and the delay simulation function, and start the real-time data simulation module; Step7.2: Read the card-swiping data within the specified date and time range; Step7.3: Expand the total amount of data in the dataset according to the load expansion multiple. Traverse each record in the "swiping time" column and add generated by the simulated delay function to generate a new column "writing time". Step7.4: Based on the new column "write time", index the sub-data set of the card-swiping data loaded in Step7.2 second by second, and write it into the message middleware Kafka.

4. A real-time rail transit passenger flow prediction method based on a streaming computing engine according to claim 3, characterized in that The specific delay simulation function is as follows: Use the delay simulation function to simulate the total transmission delay of the card swiping data from the turnstile to the site proxy server and then forwarded to the data center , and the formula of the delay simulation function is as follows: ; Among them, U(a, b) represents a uniform distribution function with a lower bound of a and an upper bound of b, which means generating a random number with a probability of that is uniformly distributed between 5. A real-time prediction method for rail transit passenger flow based on a streaming computing engine according to claim 1, characterized in that, The specific content of Step8 is as follows: Step8.1: The streaming computing engine obtains the real-time card-swiping data stream from the message middleware at a certain time interval; Step8.2: Perform data preprocessing on the real-time card-swiping data stream and group it based on a sliding time window; Step8.3: Calculate the real-time network-level short-term inbound and outbound passenger flows for the data within each group; Step8.4: Write the calculated real-time network-level short-term inbound and outbound passenger flows into the persistent key-value storage service.

6. The real-time rail transit passenger flow prediction method based on a streaming computing engine according to claim 5, wherein The specific content of Step8.3 is as follows: Using a parallel algorithm Calculate the short-term inbound and outbound passenger flow at the network level. The input of the parallel algorithm is the card-swipe data set generated by stations in the rail transit system within a time window . The card-swipe data set has partitions distributed on different nodes, and is used to represent the partition number. The specific implementation steps of the parallel algorithm are as follows: Step8.3.1: Initialize a long integer array with a length of and an initial value of 0 on each partition , which is used to store intermediate calculation results. The first elements store the inbound passenger flow, and the last elements store the outbound passenger flow; Step8.3.2: Traverse all the card-swipe data records in each partition in parallel, and calculate and update the long integer array according to the in / out station identifier and station number of the card-swipe record ; Step8.3.3: Wait for all partitions to finish traversing the dataset of the local partition, and then collect each partition's to the list storing long integer arrays ; Step8.3.4: Merge the lists of long integer arrays in each partition in the in a pairwise manner using element-wise addition of the , and finally obtain the network-level short-term inbound and outbound passenger flows within the time window. long integer arrays in each partition to finally obtain the network-level short-term inbound and outbound passenger flows within the time window.

7. A real-time prediction method for rail transit passenger flow based on a streaming computing engine according to claim 1, characterized in that The feature engineering controller is specifically implemented as follows: Design a custom type As the feature engineering controller, when deploying the feature engineering control module, create an instance of the feature engineering controller on the application's daemon process Driver for subsequent control processes, specifically: ​ Feature Engineering Controller encapsulates a real-time passenger flow data cache of type TreeMap[Long,Array[Long]] and a historical passenger flow data cache for caching historical and real-time network-level short-term in-out passenger flow data, where the TreeMap is a collection of key-value pairs sorted by key. Here, the key is the short-term passenger flow timestamp of type Long, and the value is a long integer array Array[Long] storing the network-level short-term in-out passenger flow; in addition, the Feature Engineering Controller also encapsulates other member variables that assist in the control of the online feature engineering process, such as the current date that can determine whether initialization is to be performed , and the current event time that stores the time of the latest short-term passenger flow event ; Feature Engineering Controller Encapsulates 4 core functions for abstracting and standardizing the process of online feature engineering, specifically: Initialization function : Based on the incoming timestamp , that is, the event time of the latest short-term in-out passenger flow at the network level, update the current date and other auxiliary variables, calculate and initialize the real-time passenger flow cache and the historical passenger flow cache The time range of the short-term in-out passenger flow at the network level required, and query and load the corresponding data records from the persistent key-value storage service to initialize the two passenger flow caches; Real-time Passenger Flow Data Cache Update Function : Pass in the real-time short-term inbound and outbound passenger flows at the network level to update the real-time passenger flow cache , if there is a new short-term passenger flow record appended, remove the oldest record in the real-time passenger flow cache ; Real-time Feature Time Series Calculation Function : Pass in the current event time , calculate the time index sequence of the network-level short-term inbound and outbound passenger flow records required for real-time mode features , let be , then corresponds to ; Multi - mode Feature Constructor : Input , that is , calculate and , and use to index data records from the real - time passenger flow cache to construct real - time mode passenger flow features , use and to index data records from the historical passenger flow cache to construct daily - mode and weekly - mode passenger flow features and .

Citation Information

Patent Citations

  • Urban rail transit station short-time passenger flow prediction method

    CN110276474A

  • Rail transit section passenger flow short-time prediction method and system based on big data technology

    CN110782060A

  • Traffic operation situation prediction method and system based on cellular automaton

    CN116311922A

  • Electric vehicle load prediction method considering meteorological factors and dynamic traffic

    CN117096869A