A multi-scale time series bus passenger flow prediction method based on IC card data

By constructing a multi-scale time series bus passenger flow prediction method, and using ARMA, SARIMA and ARIMA-GARCH models combined with dynamic entropy integral probability fusion and Kalman filtering techniques, the multi-scale characteristics and dynamic adaptability issues in bus passenger flow prediction are solved. This enables adaptive adjustment and continuous optimization of the model, thereby improving prediction accuracy and stability.

CN121435195BActive Publication Date: 2026-03-24FUZHOU PLANNING DESIGN & RES INST
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202512015144.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-03-24
Estimated Expiration
2045-12-30

AI Technical Summary

Technical Problem

Existing methods for predicting public transport passenger flow cannot effectively handle the multi-scale characteristics of passenger flow data, lack a dynamic model interaction mechanism, cannot adaptively adjust model weights and parameters, and fail to achieve closed-loop optimization of prediction accuracy and model updates.

Method used

The ARMA model, SARIMA model, and ARIMA-GARCH composite model are used to process time series matrices at multiple time scales. A dynamic state transition probability matrix is ​​constructed through a dynamic entropy integral probability fusion mechanism and Kalman filtering technology to achieve adaptive adjustment of model weights and parameter updates, thereby generating mixed passenger flow prediction data.

Benefits of technology

It improves the time-period adaptability and overall accuracy of public transport passenger flow forecasting, realizes model self-optimization and continuous learning, and enhances the accuracy and stability of forecast results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121435195B_ABST
    Figure CN121435195B_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-scale time series bus passenger flow prediction methods based on IC card data, it is related to intelligent transportation technical field, including, based on the passenger flow prediction data of fusion, the effectiveness probability vector of each model is obtained by dynamic entropy integral probability fusion mechanism, dynamic state transition probability matrix is constructed, and the state vector corresponding to each model and observation vector are established by state space mapping relationship, generate standardization state space framework;Kalman filter is used to condition filtering and probability updating to standardization state space framework, and weighted fusion is carried out, to generate mixed passenger flow prediction data;Multi-dimensional error analysis is carried out to mixed passenger flow prediction data and actual passenger flow observation data, to generate precision reward signal and carry out parameter updating by interactive multi-model algorithm, simultaneously update multi-time scale time series matrix.The application realizes self-optimization and continuous learning of prediction process, and reaches the effect of parameter dynamic optimization and guaranteeing continuous updating of data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent transportation technology, and in particular to a multi-scale time series bus passenger flow prediction method based on IC card data. Background Technology

[0002] With the acceleration of urbanization and the popularization of intelligent transportation, public transport passenger flow forecasting technology has become a core support for urban public transport management and optimization. Traditional passenger flow forecasting methods mainly rely on historical statistical data and simple time series models, such as the Autoregressive Integral Moving Average (ARIMA) model and its seasonal extension (SARIMA). These methods have shown some effectiveness in long-term trend forecasting under stable conditions. In recent years, with the development of machine learning technology, data-driven methods such as Support Vector Machines (SVM), Random Forests, and Neural Networks have been introduced into the field of passenger flow forecasting, improving forecast accuracy by capturing nonlinear features. In addition, with the widespread application of Internet of Things (IoT) technology, public transport IC cards have accumulated massive amounts of spatiotemporal data, providing a data foundation for refined passenger flow forecasting.

[0003] While existing technologies have made some progress in passenger flow forecasting, significant shortcomings remain. First, most methods employ static model structures, failing to effectively handle the multi-scale characteristics of passenger flow data. Furthermore, their fixed-weight fusion methods cannot adapt to the dynamic changes in passenger flow patterns during operating hours (e.g., peak / off-peak). For example, the ARIMA model excels at capturing short-term linear trends but lags in responding to periodic changes and sudden fluctuations; while neural networks can fit complex nonlinear relationships, their interpretability is poor and they are sensitive to data quality. More critically, existing methods lack dynamic model interaction mechanisms, making it impossible to adaptively adjust model weights and parameters based on real-time data. Second, traditional methods fail to achieve closed-loop optimization between prediction accuracy and model updates, and they do not establish a logical correlation between prediction results and public transport scheduling parameters. Prediction errors are merely used as evaluation indicators and are not converted into feedback signals for real-time model adjustments, preventing the learning and continuous improvement from historical errors. Summary of the Invention

[0004] In view of the aforementioned existing problems, the present invention is proposed.

[0005] Therefore, this invention provides a multi-scale time series bus passenger flow prediction method based on IC card data to solve the problems that existing technologies cannot effectively handle the multi-scale characteristics of passenger flow and lack continuous learning capabilities.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0007] This invention provides a multi-scale time-series bus passenger flow prediction method based on IC card data. The method includes: collecting IC card data streams and external data sources; constructing a multi-time-scale time-series matrix through data cleaning and aggregation; processing the multi-time-scale time-series matrix using ARMA, SARIMA, and ARIMA-GARCH composite models to generate fused passenger flow prediction data; constructing model probability vectors and state transition matrices based on the fused passenger flow prediction data and the statistics of historical prediction errors during different operating periods; obtaining the effectiveness probability vectors of each model through a dynamic entropy integral probability fusion mechanism to construct a dynamic state transition probability matrix; establishing the state vectors and observation vectors corresponding to each model through state-space mapping relationships to generate a standardized state-space framework; performing conditional filtering and probability updates on the standardized state-space framework using Kalman filtering, followed by weighted fusion to generate mixed passenger flow prediction data; performing multi-dimensional error analysis on the mixed passenger flow prediction data and actual passenger flow observation data to generate an accuracy reward signal; updating parameters through an interactive multi-model algorithm, while simultaneously updating the multi-time-scale time-series matrix.

[0008] The core innovation of this invention lies in constructing model effectiveness probabilities based on operating time periods. Specifically, traditional multi-model fusion uses a fixed weight mechanism, which cannot cope with the differentiated passenger flow patterns during different operating time periods, such as large fluctuations in passenger flow during peak hours and stable passenger flow during off-peak hours. This invention, however, by statistically analyzing the historical prediction errors of each sub-model under different operating time periods (morning peak, off-peak, evening peak, nighttime off-peak, etc.), constructs time-segmented model effectiveness probability vectors and state transition matrices, and adaptively allocates model weights according to real-time operating time periods, allowing the fusion prediction results to accurately match the passenger flow characteristics of the current time period, thus completely solving the adaptability shortcomings of fixed-weight fusion.

[0009] As a preferred embodiment of the multi-scale time series bus passenger flow prediction method based on IC card data described in this invention, the IC card data stream includes boarding station identifier, vehicle identifier information, precise timestamp, and anonymized card number;

[0010] The external data sources include time-series context data, meteorological and environmental data, and traffic and event data.

[0011] As a preferred embodiment of the multi-scale time series bus passenger flow prediction method based on IC card data described in this invention, the specific steps for constructing the multi-time-scale time series matrix are as follows:

[0012] Collect IC card data streams and external data sources, and initialize the corresponding heterogeneous intelligent agents to generate multi-dimensional feature vectors;

[0013] A multi-agent confidence-weighted negotiation game is performed on the multi-dimensional feature vector to generate a fusion weight vector;

[0014] The weighted fusion vector, IC card data stream, and external data source are weighted and fused to generate a unified passenger flow estimate.

[0015] By aggregating the unified passenger flow estimate across multiple time scales, a multi-time-scale time series matrix is ​​constructed.

[0016] As a preferred embodiment of the multi-scale time series bus passenger flow prediction method based on IC card data described in this invention, the multi-time scale time series matrix includes a weekly scale series, a daily scale series, and a time interval series.

[0017] As a preferred embodiment of the multi-scale time series public transport passenger flow prediction method based on IC card data described in this invention, the specific steps for generating fused passenger flow prediction data are as follows:

[0018] The weekly-scale sequence is input into the ARMA model and processed by the autoregressive moving average algorithm to output weekly-scale passenger flow forecast data.

[0019] The daily-scale sequence is input into the SARIMA model and processed by the seasonal autoregressive integral moving average algorithm to output daily-scale passenger flow forecast data.

[0020] The time interval series is input into the ARIMA-GARCH composite model, and the time interval scale passenger flow prediction data is output by co-processing autoregressive integral moving average and generalized autoregressive conditional heteroscedasticity.

[0021] The weekly, daily, and time-interval passenger flow forecast data are dynamically fused, with weight adjustments and weighted fusion processes to generate fused passenger flow forecast data.

[0022] As a preferred embodiment of the multi-scale time series bus passenger flow prediction method based on IC card data described in this invention, the specific steps for constructing the dynamic state transition probability matrix are as follows:

[0023] Error pattern recognition was performed on the fused passenger flow prediction data by comparing and analyzing the prediction errors of multiple models and extracting their distribution features, and the prediction error sequences of each model were obtained.

[0024] The prediction error sequence of each model is input into the entropy estimation unit. Through sliding window probability distribution calculation and information entropy analysis, the prediction error entropy value of each model is output.

[0025] The prediction error entropy value and prediction error sequence of each model are integrated by a dynamic entropy integral probability fusion mechanism to generate the effectiveness probability vector of each model.

[0026] The validity probability vectors of each model are input into the state transition analysis unit, and a dynamic state transition probability matrix is ​​constructed through historical probability sequence analysis and Markov chain modeling.

[0027] As a preferred embodiment of the multi-scale time series bus passenger flow prediction method based on IC card data described in this invention, the specific steps for generating the standardized state space framework are as follows:

[0028] Based on the dynamic state transition probability matrix, the state vector and observation vector corresponding to each model are established through the state space mapping relationship;

[0029] The state vector and the observation vector are standardized and integrated to generate a standard state space form.

[0030] The standard state-space form is integrated with the validity probability vectors of each model to generate a standardized state-space framework.

[0031] As a preferred embodiment of the multi-scale time series public transport passenger flow prediction method based on IC card data described in this invention, the specific steps of using Kalman filtering to perform conditional filtering and probability updates on the standardized state-space frame are as follows:

[0032] Kalman filtering is applied to the standardized state-space framework to generate prior state estimates and prior error covariance matrices for each model.

[0033] Based on the prior state estimates and prior error covariance matrices of each model, Kalman filtering is performed using actual passenger flow observation data to generate the posterior state estimates and updated error covariance matrices of each model.

[0034] As a preferred embodiment of the multi-scale time series public transport passenger flow prediction method based on IC card data described in this invention, the specific steps for weighted fusion to generate mixed passenger flow prediction data are as follows:

[0035] Dynamic weight allocation is performed on the posterior state estimate, the updated error covariance matrix, and the effectiveness probability vector of each model to generate dynamic fusion weights for each model.

[0036] A mixed state estimate is generated by weighted summation based on dynamic fusion weights and posterior state estimates.

[0037] The mixed state estimate is transformed by observation equations to generate mixed passenger flow prediction data.

[0038] As a preferred embodiment of the multi-scale time series bus passenger flow prediction method based on IC card data described in this invention, the steps of performing multi-dimensional error analysis on the mixed passenger flow prediction data and actual passenger flow observation data, generating an accuracy reward signal, updating parameters through an interactive multi-model algorithm, and simultaneously updating the multi-time-scale time series matrix are as follows:

[0039] Multi-dimensional error analysis is performed on the mixed passenger flow forecast data and the actual passenger flow observation data to generate instantaneous error sequences and historical error statistics;

[0040] Dynamic precision reward processing is performed based on instantaneous error sequences and historical error statistics to generate precision reward signals;

[0041] Based on the accuracy reward signal, the state transition probability matrix and the effectiveness probability vector of each model are updated through an interactive multi-model algorithm. At the same time, the collected IC card data stream and external data source are dynamically weighted and fused to update the multi-time-scale time series matrix.

[0042] The beneficial effects of this invention are:

[0043] This invention breaks through the limitations of traditional fixed-weight fusion by constructing model effectiveness probabilities based on operating time periods. It achieves adaptive adjustment of model weights under different operating time periods (peak / off-peak / valley), accurately matches time-specific passenger flow pattern changes, solves the industry pain point that fixed weights cannot dynamically adapt to passenger flow characteristics, and significantly improves the time-specific adaptability and overall accuracy of prediction results.

[0044] By combining dynamic entropy integral probability fusion with Markov chain modeling, we can achieve accurate quantification of the effectiveness probability of multiple models and dynamic characterization of state transition relationships, thereby improving the intelligence of model selection and enhancing the adaptability of the prediction process.

[0045] By generating accuracy reward signals and constructing dual feedback loops, the prediction process achieves self-optimization and continuous learning, resulting in dynamic parameter optimization and ensuring continuous data updates. Simultaneously, the logic for converting prediction results into scheduling parameters is supplemented, enhancing the engineering practicality of the technical solution.

[0046] By using the RLS algorithm to perform online parameter estimation and updates on the adapted model, the real-time performance and accuracy of the model parameters are improved, ensuring the predictive stability of the model during long-term operation. Attached Figure Description

[0047] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0048] Figure 1 This is a flowchart of a multi-scale time series bus passenger flow prediction method based on IC card data.

[0049] Figure 2 A flowchart for constructing a multi-time-scale time series matrix.

[0050] Figure 3 A flowchart for generating integrated passenger flow forecast data.

[0051] Figure 4 A flowchart for generating a standardized state-space framework. Detailed Implementation

[0052] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0053] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0054] Secondly, the term "an embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0055] Reference Figures 1-4 This is one embodiment of the present invention, which provides a multi-scale time series bus passenger flow prediction method based on IC card data, including the following steps:

[0056] S1. Collect IC card data streams and external data sources, and construct a multi-time-scale time series matrix through data cleaning and aggregation.

[0057] Collect IC card data streams and external data sources, and initialize the corresponding heterogeneous intelligent agents to generate multi-dimensional feature vectors.

[0058] The specific process includes: IC card data streams are recorded in real time by the bus's onboard terminal equipment when passengers swipe their cards to board. Each swipe generates a data record containing the boarding station identifier, vehicle identification information, precise timestamp, and anonymized card number, which is continuously transmitted to the data processing platform in the form of a data stream; time-series context data from external data sources is obtained through historical operation logs of urban public transportation scheduling; meteorological environmental data is obtained by accessing real-time weather forecasts released by meteorological departments; and traffic and event data are collected through channels such as real-time road condition information, announcements of large-scale events, and emergency notifications released by traffic management departments. All of the above types of data are kept synchronized in time during the collection process, and corresponding heterogeneous intelligent agents are initialized respectively. The heterogeneous intelligent agents perform feature parsing and vectorization based on the structural characteristics and semantic content of their respective data types to generate multi-dimensional feature vectors.

[0059] A multi-agent confidence-weighted negotiation game is performed on the multi-dimensional feature vector to generate a fusion weight vector.

[0060] The specific process includes performing a multi-agent confidence-weighted negotiation game on the multi-dimensional feature vector. This means that each heterogeneous agent evaluates the credibility of its own features in the current context based on the multi-dimensional feature vector it generates, and exchanges confidence information with other heterogeneous agents through a negotiation mechanism. During the game, each agent dynamically adjusts its influence in the fusion process according to the confidence level, and finally reaches a consensus on the weight allocation scheme to generate a fusion weight vector.

[0061] It should be noted that the negotiation mechanism is a process in which heterogeneous intelligent agents exchange their confidence assessments of multi-dimensional feature vectors and iteratively adjust the weights and influences to reach a consensus.

[0062] The weighted fusion vector, IC card data stream, and external data source are weighted and fused to generate a unified passenger flow estimate.

[0063] The specific process includes using each component in the fusion weight vector as a weighting coefficient for the IC card data stream processed by the corresponding heterogeneous intelligent agent and the external data source, and performing linear weighted combination of the IC card data stream and the external data source at the feature level or passenger flow counting level, so that the data from high confidence sources occupies a larger proportion in the fusion result, thereby generating a unified passenger flow estimate.

[0064] By aggregating the unified passenger flow estimate across multiple time scales, a multi-time-scale time series matrix is ​​constructed.

[0065] The specific process includes summarizing and statistically analyzing the unified passenger flow estimate according to three different granularity time windows: weekly, daily, and time interval series. The weekly scale accumulates or averages the unified passenger flow estimate over a seven-day period. The daily scale groups and aggregates the unified passenger flow estimate in 24-hour units. The time interval series divides and statistically analyzes the unified passenger flow estimate into finer-grained time periods, such as hours or half-hours, forming a multi-time-scale time series matrix that includes weekly, daily, and time interval series.

[0066] S2. The ARMA model, SARIMA model and ARIMA-GARCH composite model are used to process the time series matrix at multiple time scales and generate fused passenger flow prediction data.

[0067] The weekly-scale sequence is input into the ARMA model and processed by the autoregressive moving average algorithm to output weekly-scale passenger flow forecast data.

[0068] The specific process includes using the weekly time series as the input time series for the ARMA model, performing a stationarity test on the weekly time series to confirm that it meets the application conditions of the ARMA model, establishing an autoregressive relationship between the current weekly passenger flow value and several previous weekly passenger flow values ​​within the ARMA model framework, and combining a linear combination of prediction errors from several past periods as the moving average component. The autoregressive coefficients and moving average coefficients in the ARMA model are determined through maximum likelihood estimation or least squares method. Finally, the fitted ARMA model is used to recursively predict the passenger flow values ​​for subsequent weeks, outputting weekly time series passenger flow prediction data.

[0069] Furthermore, the training process of the ARMA model includes a stationarity test on the weekly-scale sequence. If the weekly-scale sequence does not meet the stationarity requirement, it is made stationary through differencing or other transformations. Then, based on the truncation or tailing characteristics of the autocorrelation function and partial autocorrelation function, the autoregressive order and moving average order of the ARMA model are initially determined. Using historical weekly-scale sequence data, the autoregressive coefficients and moving average coefficients of the ARMA model are estimated using maximum likelihood estimation or conditional least squares. The adequacy of the ARMA model's fit is then verified through a residual white noise test. If the residual sequence is white noise, the ARMA model training is complete; otherwise, the autoregressive order or moving average order is adjusted, and parameter estimation is performed again until an ARMA model that meets the statistical test requirements is obtained. The RLS algorithm is used to perform online parameter estimation and updates on the adapted model to ensure real-time parameter adaptation.

[0070] It should be noted that the autoregressive moving average algorithm is a time series forecasting method. It models and forecasts stationary time series by representing the current observation value as the sum of a linear combination of the observation values ​​at several past times (autoregressive part) and a linear combination of the prediction errors at several past times (moving average part).

[0071] The autocorrelation function measures the linear correlation between a time series and itself at different time lags, specifically representing the correlation coefficient between the current observation and an observation at a certain lag in the past. The partial autocorrelation function, on the other hand, measures the direct correlation between the current observation and an observation at a specific lag, while controlling for the influence of intermediate lags. The difference lies in that the autocorrelation function reflects the overall correlation, including indirect effects, while the partial autocorrelation function eliminates the interference of intermediate lags, retaining only the net correlation between the two time points.

[0072] The daily-scale sequence is input into the SARIMA model and processed by the seasonal autoregressive integral moving average algorithm to output daily-scale passenger flow forecast data.

[0073] The specific process includes using the daily-scale sequence as the input time series for the SARIMA model, performing stationarity and seasonality analysis on the daily-scale sequence, performing differencing if a non-stationary trend exists, and applying seasonal differencing if periodic fluctuations exist. Within the SARIMA model framework, non-seasonal autoregressive terms, non-seasonal moving average terms, seasonal autoregressive terms, and seasonal moving average terms are constructed simultaneously, and the difference order and seasonal difference order are combined to form a complete SARIMA model structure. Using historical daily-scale sequence data, the coefficients of each term in the SARIMA model are determined through methods such as maximum likelihood estimation. Then, the rationality of the model is verified through residual diagnosis. Finally, based on the fitted SARIMA model, the passenger flow values ​​for future dates are recursively predicted, and daily-scale passenger flow prediction data is output.

[0074] Further, the training process of the SARIMA model involves: performing a stationarity test on the daily-scale sequence; if a trend is observed in the daily-scale sequence, non-seasonal differencing is performed; if a periodic seasonal pattern exists, seasonal differencing is applied to bring the sequence to a stationary state. The autocorrelation function and partial autocorrelation function of the differencing sequence are analyzed. Combined with the seasonal cycle length, the non-seasonal autoregressive order, non-seasonal moving average order, seasonal autoregressive order, seasonal moving average order, and the corresponding non-seasonal and seasonal differencing orders of the SARIMA model are initially determined. Using historical daily-scale sequence data, all parameters in the SARIMA model are jointly estimated using maximum likelihood estimation or conditional least squares. The fitted residuals are subjected to white noise and normality tests to determine whether the residuals are uncorrelated random errors. If the residuals satisfy the white noise assumption, the SARIMA model training is complete; otherwise, the model order is adjusted and the parameters are re-estimated based on the diagnostic results until a statistically reasonable SARIMA model is obtained.

[0075] It should be noted that the seasonal autoregressive integral moving average algorithm is a method for modeling and predicting time series with trend and periodic characteristics. By combining non-seasonal autoregressive terms, difference terms, and moving average terms with seasonal autoregressive terms, seasonal difference terms, and seasonal moving average terms corresponding to the seasonal cycle, it simultaneously characterizes and predicts the long-term trend and recurring seasonal patterns of the original series.

[0076] The time interval series is input into the ARIMA-GARCH composite model, and the time interval scale passenger flow prediction data is output through the combined processing of autoregressive integral moving average and generalized autoregressive conditional heteroscedasticity.

[0077] The specific process includes: modeling the time interval series using autoregressive integral moving average; stationarizing the series through differencing; establishing a linear relationship between the current passenger flow value, historical passenger flow value, and historical prediction error to obtain the conditional mean equation; using the residual series of the autoregressive integral moving average model as input to the generalized autoregressive conditional heteroscedasticity (GHP); modeling the squared term of the residuals to capture the heteroscedasticity characteristics of passenger flow fluctuations over time to form the conditional variance equation; jointly estimating the parameters of the autoregressive integral moving average and the GHP to enable the ARIMA-GARCH composite model to simultaneously characterize the temporal dependence and volatility clustering of the time interval series; and finally predicting the passenger flow values ​​for future time intervals based on the ARIMA-GARCH composite model to output time interval-scale passenger flow prediction data.

[0078] Further, the training process of the ARIMA-GARCH composite model is as follows: The time interval series is subjected to a unit root test to determine its stationarity. If the time interval series is non-stationary, differencing is performed until a stationary series is obtained, and the differencing order in the autoregressive integral moving average is determined accordingly. Based on the stationary time interval series, the autoregressive order and moving average order of the autoregressive integral moving average are initially set by analyzing the autocorrelation function and partial autocorrelation function. The autoregressive integral moving average model is then fitted using maximum likelihood estimation or conditional least squares to obtain the residual series. The ARCH effect is tested on the residual series. If heteroscedasticity exists, the residual series is input into the generalized autoregressive conditional heteroscedasticity (GARCH). The autoregressive order and moving average order of the GARCH are determined based on the autocorrelation characteristics of the squared residuals, and the parameters of the GARCH are estimated. Finally, the parameters of the autoregressive integral moving average and the GARCH are jointly optimized. The model's rationality is verified through information criteria and residual diagnosis, completing the pre-training of the ARIMA-GARCH composite model.

[0079] Heteroscedasticity refers to the variation of the error term or fluctuation amplitude of a time series over time, meaning that the prediction error has different variances at different times. In passenger flow data, heteroscedasticity manifests as drastic fluctuations and large variances in passenger flow during certain periods (such as morning and evening peak hours), while passenger flow is stable and has smaller variances during other periods (such as late at night). Unlike the homoscedasticity assumption of constant variance, heteroscedasticity reflects the non-uniformity and clustering of data fluctuations.

[0080] Information criteria are statistical indicators used for model selection. They help avoid overfitting by balancing the goodness of fit of the model with the number of parameters. In the training of the ARIMA-GARCH composite model, commonly used information criteria include the Akaike Information Criterion (AIC) and the Bayesian Information Criterion (BIC). The smaller the value, the better the model. These criteria are used to determine the optimal combination of autoregression order, moving average order, and generalized autoregressive conditional heteroscedasticity order.

[0081] The weekly, daily, and time-interval passenger flow forecast data are dynamically fused, with weight adjustments and weighted fusion processes to generate fused passenger flow forecast data.

[0082] The specific process includes obtaining the corresponding dynamic fusion weights based on the confidence level or historical prediction accuracy of weekly, daily, and time-interval passenger flow forecast data at the current prediction time. The weekly, daily, and time-interval passenger flow forecast data are then weighted and summed with their respective dynamic fusion weights. The result is the fused passenger flow forecast data, which simultaneously reflects long-term cycle trends, short-to-medium-term daily variation patterns, and fine-grained fluctuation characteristics.

[0083] S3. Based on the fusion of passenger flow prediction data, the effectiveness probability vectors of each model are obtained through the dynamic entropy integral probability fusion mechanism, a dynamic state transition probability matrix is ​​constructed, and the state vectors and observation vectors corresponding to each model are established through the state space mapping relationship to generate a standardized state space framework.

[0084] Error pattern recognition was performed on the fused passenger flow prediction data by comparing and analyzing the prediction errors of multiple models and extracting their distribution features, and the prediction error sequences of each model were obtained.

[0085] The specific process includes comparing the fused passenger flow forecast data with the actual passenger flow observation data point by point in time, obtaining the prediction errors corresponding to the weekly-scale passenger flow forecast data output by the ARMA model, the daily-scale passenger flow forecast data output by the SARIMA model, and the time-interval-scale passenger flow forecast data output by the ARIMA-GARCH composite model, forming the prediction error sequence of each model. Further statistical distribution characteristic analysis is performed on these prediction error sequences, including the extraction of features such as the mean, variance, skewness, kurtosis, and temporal correlation of the errors, thereby identifying the error patterns of different models under different scenarios and obtaining the prediction error sequence of each model.

[0086] The prediction error sequences of each model are input into the entropy estimation unit. Through sliding window probability distribution calculation and information entropy analysis, the prediction error entropy value of each model is output, expressed as:

[0087] ;

[0088] in, Representation Model In time The prediction error entropy value, Indicates the index of the model. Indicates the current time point, Indices representing the discrete intervals of the error. This represents the total number of discrete intervals of the error. This represents the time variable within the sliding window. Indicates the size of the sliding window. Indicates a point in time The time decay weight value at that point, Indicates the first The lower boundary value of each error discrete interval Indicates the first The upper boundary value of a discrete interval of error. Representation Model In time The prediction error This represents the time variable within the sliding window.

[0089] The specific process includes sending the prediction error sequences of each model corresponding to the weekly-scale passenger flow prediction data output by the ARMA model, the daily-scale passenger flow prediction data output by the SARIMA model, and the time-interval-scale passenger flow prediction data output by the ARIMA-GARCH composite model into the entropy estimation unit for processing. In the entropy estimation unit, a fixed-length sliding window is set and slides sequentially along the time axis. Each time point within the coverage of the sliding window is assigned a weight value that decays over time to reflect the greater impact of recent errors on the current uncertainty assessment. All weighted error values ​​within the sliding window are divided into discrete intervals, and the weighted frequency of error values ​​within each discrete interval is counted and normalized to obtain the probability distribution of errors within the window. Based on the probability distribution, the prediction error entropy value of the corresponding model at the current time point is calculated according to the definition of information entropy. The prediction error entropy value quantitatively characterizes the randomness or uncertainty level of the model prediction error.

[0090] The dynamic entropy integral probability fusion mechanism integrates the prediction error entropy values ​​and prediction error sequences of each model to generate the effectiveness probability vector of each model.

[0091] The specific process includes inputting the prediction error entropy values ​​of the ARMA model, SARIMA model, and ARIMA-GARCH composite model, along with their respective prediction error sequences, into a dynamic entropy integral probability fusion mechanism. In this mechanism, the uncertainty of each model's prediction result at the current moment is assessed based on the magnitude of its prediction error entropy value; a lower entropy value indicates a more stable and reliable model prediction. Furthermore, the error amplitude and volatility are quantified using a time-decay weighted method, taking into account the historical performance of each model's prediction error sequence. Based on this, a confidence score for each model is constructed using the reciprocal or negative exponential transformation of the entropy value, and all model confidence scores are normalized to give each model a probability value between zero and one. This probability value reflects the model's relative effectiveness in the current prediction context. Finally, the effectiveness probabilities of the ARMA model, SARIMA model, and ARIMA-GARCH composite model are arranged sequentially to generate an effectiveness probability vector for each model.

[0092] It should be noted that the dynamic entropy integral probability fusion mechanism is a fusion method that dynamically generates the effectiveness probabilities of ARMA models, SARIMA models, and ARIMA-GARCH composite models by converting the entropy values ​​of each model's prediction error and the prediction error sequence of each model into model confidence values ​​and performing normalization processing.

[0093] The validity probability vectors of each model are input into the state transition analysis unit, and a dynamic state transition probability matrix is ​​constructed through historical probability sequence analysis and Markov chain modeling.

[0094] The specific process includes: forming a historical probability sequence from the validity probability vectors of each model at the current time and several past times; performing statistical analysis on the historical probability sequence in the state transition analysis unit to identify the validity transition patterns of the ARMA model, SARIMA model, and ARIMA-GARCH composite model at different time points; and estimating the state transition frequency using the historical probability sequence based on the Markov chain modeling assumption, i.e., the validity state of each model at the current time depends only on the state at the previous time, thereby constructing a dynamic state transition probability matrix that updates over time. The dynamic state transition probability matrix describes the probability of mutual transition of validity states between the ARMA model, SARIMA model, and ARIMA-GARCH composite model.

[0095] It should be noted that Markov chain modeling is a method for mathematically describing stochastic processes with state transition characteristics. The core assumption is that the state at the next moment depends only on the state at the current moment and is independent of earlier historical states. By defining a set of discrete states and the transition probabilities between states, a state transition probability matrix is ​​constructed to characterize the possibility of transitioning from one state to another, and can be used to analyze long-term behavior, steady-state distribution, or make state predictions.

[0096] Based on the dynamic state transition probability matrix, the state vector and observation vector corresponding to each model are established through the state space mapping relationship.

[0097] The specific process involves defining a state vector for each model based on the transition rules of the effective states between the ARMA model, SARIMA model, and ARIMA-GARCH composite model described by the dynamic state transition probability matrix. The state vector represents the internal predicted state of the model at the current moment. At the same time, based on the actual available passenger flow observation information, a corresponding observation vector is constructed for each model. The observation vector reflects the relationship between the model's predicted output and the actual passenger flow observation data. The state vector and the observation vector are associated through a state space mapping relationship, thereby transforming the multi-model prediction problem into a standard state space expression.

[0098] The state vector and observation vector are standardized and integrated to generate a standard state space form.

[0099] The specific process includes unifying the dimensions and normalizing the values ​​of the state and observation vectors corresponding to the ARMA model, the SARIMA model, and the ARIMA-GARCH composite model, respectively, to eliminate the bias caused by scale differences between different models. Based on state space theory, the state evolution equations and observation equations of each model are organized into matrix expressions with a unified structure. The state equations describe the dynamic evolution of the state vectors over time, and the observation equations describe the linear mapping relationship between the observation vectors and the state vectors, forming a standard state space form that meets the requirements of Kalman filtering.

[0100] The standard state-space form is integrated with the validity probability vectors of each model to generate a standardized state-space framework.

[0101] The specific process involves combining the standard state-space forms corresponding to the ARMA model, SARIMA model, and ARIMA-GARCH composite model with the probability values ​​corresponding to the validity probability vectors of each model. This allows the state equations and observation equations of each model to be weighted and embedded with their current validity probabilities, thereby expressing the dynamic state evolution characteristics and relative reliability of each model in a unified mathematical structure. This forms a standardized state-space framework that integrates multi-model prediction capabilities and real-time validity assessment.

[0102] S4. Kalman filtering is used to perform conditional filtering and probability updates on the standardized state-space framework, and weighted fusion is performed to generate mixed passenger flow prediction data.

[0103] Kalman filtering is applied to the standardized state-space framework to generate prior state estimates and prior error covariance matrices for each model.

[0104] The specific process includes: based on the state equations defined by the ARMA model, SARIMA model, and ARIMA-GARCH composite model in the standardized state space framework, using the posterior state estimate of the previous time step as the initial condition, and recursively obtaining the state prediction value of the current time step through the state transition matrix, thus obtaining the prior state estimates of the ARMA model, SARIMA model, and ARIMA-GARCH composite model. At the same time, based on the posterior error covariance matrix, state transition matrix, and process noise covariance of the previous time step, the uncertainty is propagated according to the prediction equation of Kalman filtering, thus obtaining the prior error covariance matrix of the ARMA model, SARIMA model, and ARIMA-GARCH composite model.

[0105] Based on the prior state estimates and prior error covariance matrices of each model, Kalman filtering is performed using actual passenger flow observation data to generate the posterior state estimates and updated error covariance matrices of each model.

[0106] The specific process includes obtaining the Kalman gain for each model. The Kalman gain is jointly determined by the prior error covariance matrix and the observation noise covariance of each model. The residuals between the actual passenger flow observation data and the prior state estimates of each model are weighted and fed back to the prior state estimates using the Kalman gain, thereby correcting and obtaining the posterior state estimates of the ARMA model, the SARIMA model, and the ARIMA-GARCH composite model. At the same time, the uncertainty measure is updated based on the Kalman gain and the prior error covariance matrix, generating the updated error covariance matrix of the ARMA model, the updated error covariance matrix of the SARIMA model, and the updated error covariance matrix of the ARIMA-GARCH composite model.

[0107] Dynamic weight allocation is performed on the posterior state estimate, the updated error covariance matrix, and the validity probability vector of each model to generate dynamic fusion weights for each model.

[0108] The specific process includes combining the current posterior state estimation accuracy of each model, the degree of state uncertainty reflected by the updated error covariance matrix, and the validity probability of each model in the validity probability vector of each model. By using the trace or determinant of the updated error covariance matrix as an uncertainty measure, and combining it with the corresponding component in the validity probability vector of each model, a confidence score that changes over time is constructed. Then, the confidence scores of all models are normalized to generate the dynamic fusion weights of the ARMA model, the dynamic fusion weights of the SARIMA model, and the dynamic fusion weights of the ARIMA-GARCH composite model.

[0109] A mixed state estimate is generated by weighted summation based on dynamic fusion weights and posterior state estimates.

[0110] The specific process includes weighted summation of the dynamic fusion weights of the ARMA model, the SARIMA model, and the ARIMA-GARCH composite model with the posterior state estimates of the ARMA model, the SARIMA model, and the ARIMA-GARCH composite model. The posterior state estimate of each model is then multiplied by its corresponding dynamic fusion weight and summed to obtain a vector that comprehensively reflects the current prediction confidence and state information of each model, which is the hybrid state estimate.

[0111] The mixed state estimate is transformed by observation equations to generate mixed passenger flow prediction data.

[0112] The specific process includes substituting the mixed state estimate into the observation equation defined in the standardized state space framework, mapping the mixed state estimate from the state space to the observation space through the observation equation, thereby obtaining a predicted value with the same dimension as the actual passenger flow observation data. The predicted value is the mixed passenger flow prediction data that integrates the advantages of the ARMA model, SARIMA model and ARIMA-GARCH composite model and has been dynamically weighted.

[0113] S5. Perform multi-dimensional error analysis on the mixed passenger flow forecast data and the actual passenger flow observation data, generate an accuracy reward signal, and update the parameters through an interactive multi-model algorithm, while updating the multi-timescale time series matrix.

[0114] Multi-dimensional error analysis is performed on the mixed passenger flow forecast data and the actual passenger flow observation data to generate instantaneous error sequences and historical error statistics.

[0115] The specific process includes obtaining the difference between the mixed passenger flow forecast data and the actual passenger flow observation data at each time point to form an instantaneous error sequence reflecting the current forecast deviation. Based on this, the errors over a period of time are statistically summarized, including calculating indicators such as the mean, variance, mean absolute error, and root mean square error of the errors, which constitute historical error statistics.

[0116] Dynamic precision reward processing is performed based on instantaneous error sequences and historical error statistics to generate precision reward signals.

[0117] The specific process includes obtaining a numerical index that adjusts with the prediction accuracy based on the magnitude of the prediction error at the current moment in the instantaneous error sequence and the long-term prediction stability reflected by the historical error statistics. When the instantaneous error is small and the historical error statistics show good prediction performance, this numerical index takes a higher value, indicating a higher accuracy reward; otherwise, it takes a lower value. The numerical index is the accuracy reward signal.

[0118] Based on the accuracy reward signal, the state transition probability matrix and the effectiveness probability vector of each model are updated through an interactive multi-model algorithm. At the same time, the collected IC card data stream and external data source are dynamically weighted and fused to update the multi-time-scale time series matrix.

[0119] The specific process includes: using the accuracy reward signal to quantify the performance of the ARMA model, SARIMA model, and ARIMA-GARCH composite model in the interactive multi-model algorithm; adjusting the transition frequency estimation of the effective states between the models to correct the dynamic state transition probability matrix; and re-obtaining the effectiveness probability of each model based on the corrected state transition relationship to generate an updated effectiveness probability vector for each model. Simultaneously, the corresponding heterogeneous agents are initialized for the newly collected IC card data stream and external data source, generating new multi-dimensional feature vectors. A multi-agent confidence-weighted negotiation game is performed to obtain a new fusion weight vector. This fusion weight vector is then weighted and fused with the newly collected IC card data stream and external data source to obtain an updated unified passenger flow estimate. Finally, the estimate is re-aggregated according to three time granularities: weekly, daily, and time interval series, completing the update of the multi-time-scale time series matrix.

[0120] In summary, this invention achieves precise quantification of the effectiveness probability of multiple models and dynamic characterization of state transition relationships by combining dynamic entropy integral probability fusion with Markov chain modeling, thereby improving the intelligence of model selection and enhancing the adaptability of the prediction process. By generating accuracy reward signals and constructing dual feedback loops, the invention achieves self-optimization and continuous learning of the prediction process, resulting in dynamic parameter optimization and ensuring continuous data updates.

[0121] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A multi-scale time series bus passenger flow prediction method based on IC card data, characterized in that: include, The process involves collecting IC card data streams and external data sources, and then cleaning and aggregating the data to construct a multi-time-scale time series matrix. The specific steps are as follows: Collect IC card data streams and external data sources, and initialize the corresponding heterogeneous intelligent agents to generate multi-dimensional feature vectors; A multi-agent confidence-weighted negotiation game is performed on the multi-dimensional feature vector to generate a fusion weight vector; The weighted fusion vector, IC card data stream, and external data source are weighted and fused to generate a unified passenger flow estimate. A multi-time-scale time series matrix is ​​constructed by aggregating unified passenger flow estimates across multiple time scales. The ARMA model, SARIMA model, and ARIMA-GARCH composite model are used to process time series matrices at multiple time scales and generate fused passenger flow forecast data. Based on the fusion of passenger flow prediction data and the statistics of historical prediction errors during different operating periods, a model probability vector and a state transition matrix are constructed. Then, the effectiveness probability vector of each model is obtained through the dynamic entropy integral probability fusion mechanism, a dynamic state transition probability matrix is ​​constructed, and the state vector and observation vector corresponding to each model are established through the state space mapping relationship to generate a standardized state space framework. Kalman filtering is used to perform conditional filtering and probability updates on the standardized state-space framework, and then weighted fusion is performed to generate mixed passenger flow prediction data. Multi-dimensional error analysis is performed between the mixed passenger flow forecast data and the actual passenger flow observation data to generate an accuracy reward signal. Parameters are then updated using an interactive multi-model algorithm, simultaneously updating the multi-timescale time series matrix. The specific steps are as follows: Multi-dimensional error analysis is performed on the mixed passenger flow forecast data and the actual passenger flow observation data to generate instantaneous error sequences and historical error statistics; Dynamic precision reward processing is performed based on instantaneous error sequences and historical error statistics to generate precision reward signals; Based on the accuracy reward signal, the state transition probability matrix and the effectiveness probability vector of each model are updated through an interactive multi-model algorithm. At the same time, the collected IC card data stream and external data source are dynamically weighted and fused to update the multi-time-scale time series matrix.

2. The multi-scale time series bus passenger flow prediction method based on IC card data as described in claim 1, characterized in that: The IC card data stream includes boarding station identifier, vehicle identifier information, precise timestamp, and anonymized card number; The external data sources include time-series context data, meteorological and environmental data, and traffic and event data.

3. The multi-scale time series bus passenger flow prediction method based on IC card data as described in claim 2, characterized in that: The multi-timescale time series matrix includes weekly-scale sequences, daily-scale sequences, and time interval sequences.

4. The multi-scale time series bus passenger flow prediction method based on IC card data as described in claim 3, characterized in that: The specific steps for generating the integrated passenger flow prediction data are as follows: The weekly-scale sequence is input into the ARMA model and processed by the autoregressive moving average algorithm to output weekly-scale passenger flow forecast data. The daily-scale sequence is input into the SARIMA model and processed by the seasonal autoregressive integral moving average algorithm to output daily-scale passenger flow forecast data. The time interval series is input into the ARIMA-GARCH composite model, and the time interval scale passenger flow prediction data is output by co-processing autoregressive integral moving average and generalized autoregressive conditional heteroscedasticity. The weekly, daily, and time-interval passenger flow forecast data are dynamically fused, with weight adjustments and weighted fusion processes to generate fused passenger flow forecast data.

5. The multi-scale time series bus passenger flow prediction method based on IC card data as described in claim 4, characterized in that: The specific steps for constructing the dynamic state transition probability matrix are as follows: Error pattern recognition was performed on the fused passenger flow prediction data by comparing and analyzing the prediction errors of multiple models and extracting their distribution features, and the prediction error sequences of each model were obtained. The prediction error sequence of each model is input into the entropy estimation unit. Through sliding window probability distribution calculation and information entropy analysis, the prediction error entropy value of each model is output. The prediction error entropy value and prediction error sequence of each model are integrated by a dynamic entropy integral probability fusion mechanism to generate the effectiveness probability vector of each model. The validity probability vectors of each model are input into the state transition analysis unit, and a dynamic state transition probability matrix is ​​constructed through historical probability sequence analysis and Markov chain modeling.

6. The multi-scale time series bus passenger flow prediction method based on IC card data as described in claim 5, characterized in that: The specific steps for generating the standardized state-space framework are as follows. Based on the dynamic state transition probability matrix, the state vector and observation vector corresponding to each model are established through the state space mapping relationship; The state vector and the observation vector are standardized and integrated to generate a standard state space form. The standard state-space form is integrated with the validity probability vectors of each model to generate a standardized state-space framework.

7. The multi-scale time series bus passenger flow prediction method based on IC card data as described in claim 6, characterized in that: The Kalman filter is used to perform conditional filtering and probability updates on the standardized state-space framework. The specific steps are as follows. Kalman filtering is applied to the standardized state-space framework to generate prior state estimates and prior error covariance matrices for each model. Based on the prior state estimates and prior error covariance matrices of each model, Kalman filtering is performed using actual passenger flow observation data to generate the posterior state estimates and updated error covariance matrices of each model.

8. The multi-scale time series bus passenger flow prediction method based on IC card data as described in claim 7, characterized in that: The weighted fusion process to generate mixed passenger flow prediction data involves the following steps: Dynamic weight allocation is performed on the posterior state estimate, the updated error covariance matrix, and the effectiveness probability vector of each model to generate dynamic fusion weights for each model. The mixed state estimate is generated by weighted summation based on dynamic fusion weights and posterior state estimates; the mixed state estimate is then transformed by observation equations to generate mixed passenger flow prediction data.

Citation Information

Patent Citations

  • Method for predicting short-time passenger flow of bus

    CN104517159A