Subway passenger flow volume prediction method and system based on deep learning model
By generating time-frequency domain fusion data and combining CNN and Transformer models, the problem of temporal and spatial dependence in subway passenger flow prediction in the prior art is solved, and the prediction accuracy and prediction ability of long-term trends are improved.
Patent Information
- Application Number
- CN202510268818.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art is difficult to effectively capture complex space-time dependencies in subway passenger flow prediction, and the model training is complex and the computational complexity is high. It is difficult to fully consider the complexity and multi-level characteristics of the data using CNN or RNN alone.
The subway passenger flow prediction method based on deep learning model is adopted. By obtaining passenger flow data and off-site weather data, time-frequency domain fusion data is generated, and feature extraction and modeling is combined with CNN and Transformer models, the station type of the subway station is considered to improve prediction accuracy.
It improves the accuracy of the prediction of passenger flow in the subway station for a long period of time, overcomes the limitations of RNN in long-sequence data processing, avoids gradient vanishing or gradient explosion problems, and enhances the prediction ability of long-term passenger flow trends.
Smart Images

Figure CN120217080A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical fields of deep learning and traffic passenger flow prediction, and particularly relates to a subway passenger flow prediction method and system based on a deep learning model. Background Art
[0002] With the continuous expansion of the urban subway system and the gradual saturation of ground traffic, the congestion of ground traffic has become increasingly prominent, leading more and more people to choose the subway for long-distance travel. Therefore, how to effectively predict subway passenger flow has become an important issue in urban traffic management. Accurate passenger flow prediction can help relevant departments optimize operation and dispatching, avoid excessive congestion during peak hours, improve the operation efficiency of the subway, and ensure the travel experience of passengers.
[0003] Currently, subway passenger flow prediction mainly includes traditional statistical analysis methods, machine learning models, and deep learning models. For traditional statistical methods (Autoregressive Integrated Moving Average (ARMA) model, Seasonal Autoregressive Integrated Moving Average (SARIMA) model), they are often vulnerable to external factors (such as weather, holidays, etc.) when dealing with complex non-linear relationships and high-dimensional features, thus failing to effectively capture these impacts and resulting in a decline in prediction accuracy; machine learning methods (Support Vector Machine (SVM), Decision Tree (DT), Random Forest (RF)) have improved prediction accuracy compared to traditional statistical methods when facing relatively complex non-linear relationships, but when dealing with long time series data, model training may become more complex, and a large amount of data annotation and feature engineering are required, increasing the cost of model construction and maintenance; nowadays, deep learning methods (Deep Neural Network (DNN), Convolutional Neural Network (CNN), and Recurrent Neural Network (RNN)) can effectively capture complex spatio-temporal dependence relationships when dealing with passenger flow data with time series features, but when dealing with multi-dimensional, large-scale time series data, they are prone to problems such as difficult model training and high computational complexity. Secondly, when using CNN or RNN alone to deal with subway passenger flow prediction, it is often difficult to fully consider the complexity and multi-level features of the data; the Transformer model can effectively capture global features through the multi-head attention mechanism and has strong parallel computing capabilities, overcoming the limitations of RNN in processing long sequence data, but it is not as accurate as CNN in dealing with local features and has high requirements for prior knowledge of the data, often resulting in poor prediction effects. Summary of the Invention
[0004] To address the above technical problems, this application provides a subway passenger flow prediction method and system based on a deep learning model, which obtains passenger flow data for future time steps through a trained deep learning model according to multi-dimensional time series data, improving the accuracy of subway passenger flow prediction over a long time period.
[0005] In a first aspect, an embodiment of the present application provides a subway passenger flow prediction method based on a deep learning model, including:
[0006] Obtain the passenger flow data and off-station weather data of the target subway station in the first time period;
[0007] Generate a time-domain data sequence and a frequency-domain data sequence according to the passenger flow data and the off-station weather data, and splice the time-domain data sequence and the frequency-domain data sequence to generate time-frequency domain fusion data;
[0008] Input the time-frequency domain fusion data and the station type of the target subway station into a preset passenger flow prediction model, so that after the passenger flow prediction model extracts features and performs feature mapping on the time-frequency domain fusion data, generate the passenger flow prediction data for the second time period, and the station type is a double-peak type station, a full-peak type station or a non-peak type station;
[0009] Among them, the passenger flow prediction model is obtained by training an initial prediction model according to a historical time series data set, and the initial prediction model is jointly constructed by an initial CNN model and an initial Transformer model.
[0010] The present application provides a subway passenger flow prediction method based on a deep learning model, which predicts the passenger flow data of the target subway station in a future period of time according to the passenger flow data and the off-station weather data. Among them, considering that the passenger flow data and the off-station weather data have time series attributes and change with time, and since traditional passenger flow prediction methods all analyze the time domain, the time domain has certain advantages in capturing local correlations, but the frequency domain is also more effective in capturing the global correlations of the non-periodic sequence of subway passenger flow. Therefore, these two types of data are preprocessed to generate a time-domain data sequence and a frequency-domain data sequence, and then spliced to obtain time-frequency domain fusion data, which improves the accuracy of subsequent passenger flow prediction. Regarding model construction, the embodiment of the present application combines a CNN model and a Transformer model to jointly predict the model, integrating the local feature extraction ability of the CNN and the global information modeling ability of the Transformer, which can effectively predict the subway passenger flow and improve the accuracy of the long-term passenger flow prediction of the subway station. In addition, the embodiment of the present application also combines the station type of the subway station for prediction. Considering that the passenger flow change trends of different station types are different, the station type of the target subway station and the time-frequency domain fusion data are input into the passenger flow prediction model together during the prediction process, further improving the accuracy of the model prediction.
[0011] Further, generating a time-domain data sequence and a frequency-domain data sequence according to the passenger flow data and the off-station weather data, and performing data splicing on the time-domain data sequence and the frequency-domain data sequence to generate time-frequency domain fusion data, including:
[0012] Performing data merging and normalization processing on the passenger flow data and the off-station weather data according to the time sequence to obtain the time-domain data sequence;
[0013] Performing Fourier transform on the time-domain data sequence to generate the frequency-domain data sequence;
[0014] Performing data splicing on the time-domain data sequence and the frequency-domain data sequence along the feature channels of a preset dimension to generate the time-frequency domain fusion data.
[0015] The embodiment of the present application provides a method for generating time-frequency domain fusion data, introducing the fast Fourier transform to solve the problem of discrete frequency offset, using the fast Fourier transform to transform the tiled time-domain data sequence into the frequency domain to generate the corresponding frequency-domain data sequence, and fully excavating the feature information in the data. Then, the time-domain data sequence and the frequency-domain data sequence are merged to form a new feature layer, obtaining a time-frequency domain fusion data that contains both local features and global features, improving the accuracy of subsequent passenger flow prediction.
[0016] In a possible implementation manner, after the passenger flow prediction model performs feature extraction and feature mapping on the time-frequency domain fusion data, generating passenger flow prediction data for a second time period, including:
[0017] Inputting the station type of the target subway station and the time-frequency domain fusion data into a preset CNN model, so that the CNN model performs convolution, activation, and pooling operations on the time-frequency domain fusion data, and then generating feature maps at different time steps;
[0018] Inputting each of the feature maps into a preset Transformer model, so that the Transformer model performs multi-layer encoding on each of the feature maps to generate high-dimensional features, and then mapping the high-dimensional features to the target prediction space, and then generating passenger flow prediction data for the second time period.
[0019] Further, the CNN model performs convolution, activation, and pooling operations on the time-frequency domain fusion data, and then generating feature maps at different time steps, including:
[0020] Using one-dimensional convolution to extract the local time sequence features and local frequency domain features of the time-frequency domain fusion data to obtain a number of single-step feature maps;
[0021] Perform activation and pooling operations on each of the single-step feature maps through a preset activation function and pooling window to generate feature maps for different time steps.
[0022] An embodiment of the present application provides a method for generating passenger flow prediction data. First, the CNN model performs convolution, activation, and pooling operations on the time-frequency domain fusion data according to the station type to extract local features in the data and generate feature maps for different time steps. Then, the Transformer model captures the global temporal dependence relationship of the data, extracts global features from each feature map and performs multi-layer encoding to obtain high-dimensional features. Finally, the high-dimensional features are mapped to the target prediction space to compress complex feature information and generate the final prediction result. Through this process, the dimensions of the input and output are adjusted so that the features extracted from the Transformer model are accurately transformed into the passenger flow prediction data for the second time period, thereby completing the prediction task of the entire model. By combining the CNN model and the Transformer model, the embodiment of the present application overcomes the limitations of the RNN in processing long sequence data and avoids the problem of gradient disappearance or gradient explosion existing in traditional RNN or LSTM models when modeling long sequences, thereby improving the prediction ability of the long-term passenger flow trend.
[0023] In a possible implementation manner, the training of the initial prediction model according to the historical time series data set to obtain the passenger flow prediction model includes:
[0024] Obtain the historical passenger flow data and historical off-station weather data of each subway station on the target subway line;
[0025] Analyze the passenger flow change trend of each subway station according to each of the historical passenger flow data, and determine the station type of each subway station;
[0026] Generate corresponding historical time domain data sequences and historical frequency domain data sequences according to each of the historical passenger flow data and each historical off-station weather data, and splice each of the historical time domain data sequences and the corresponding historical frequency domain data sequences with the subway station as the basic unit to generate corresponding historical time-frequency domain fusion data for each subway station;
[0027] Construct a training data set according to each of the time-frequency domain fusion data, the station type of each subway station, and each of the historical passenger flow data;
[0028] Use the training data set to perform several iterative trainings on the initial prediction model in a supervised learning manner until the preset number of training times is reached to obtain training parameters;
[0029] Import the training parameters into the initial prediction model to obtain the passenger flow prediction model.
[0030] An embodiment of the present application provides a method for training an initial prediction model. First, the station types of each subway station are determined through the analysis of historical passenger flow data. The passenger flow change trends of different station types are different. Determining the station type in advance and constructing it as part of the training data can assist the model in understanding the passenger flow change characteristics of different types of stations, and improve the training efficiency and prediction accuracy of the model. Then, according to the training data, the initial prediction model is iteratively trained several times in a supervised learning manner, and the training duration of the model is controlled by a preset number of training times, avoiding overfitting of the model due to excessive training time, and improving the training efficiency of the model.
[0031] Furthermore, during the process of iteratively training the initial prediction model several times, it is judged whether to terminate the training in advance through an early stopping mechanism.
[0032] The embodiment of the present application further introduces an early stopping mechanism to prevent the model from falling into overfitting. Specifically, the early stopping mechanism is that after each training round, the validation set is evaluated and its loss value is calculated. If the validation set loss continues to decrease, it means that the model is still effectively learning; if the validation set loss stops decreasing or starts to increase, it indicates that the model may start to overfit. At this time, even if the preset number of training times is not reached, the training will be terminated in advance, avoiding ineffective overfitting training of the model, and improving the training efficiency and robustness of the model.
[0033] Furthermore, during the process of iteratively training the initial prediction model several times, a preset proportion of neurons in the initial prediction model are randomly discarded through Dropout regularization technology.
[0034] The embodiment of the present application further introduces Dropout regularization technology to reduce overfitting. Specifically, Dropout regularization technology is to randomly discard some neurons in the neural network at each training stage, forcing the network to learn more robust feature representations, gradually reducing the training parameters during the training process, and improving the training efficiency and robustness of the model.
[0035] In a second aspect, an embodiment of the present application provides a subway passenger flow prediction system based on a deep learning model, including an acquisition module, a data processing module, and a prediction module;
[0036] Among them, the acquisition module is used to acquire the passenger flow data and off-station weather data of the target subway station in the first time period;
[0037] The data processing module is used to generate a time-domain data sequence and a frequency-domain data sequence according to the passenger flow data and off-station weather data, and splice the time-domain data sequence and the frequency-domain data sequence to generate time-frequency domain fusion data;
[0038] The prediction module is used to input the time-frequency domain fusion data and the station type of the target subway station into a preset passenger flow prediction model, so that after the passenger flow prediction model extracts features and performs feature mapping on the time-frequency domain fusion data, passenger flow prediction data for a second time period is generated, and the station type is a double-peak station, a full-peak station or a non-peak station;
[0039] Among them, the passenger flow prediction model is obtained by training an initial prediction model according to a historical time series data set, and the initial prediction model is jointly constructed by an initial CNN model and an initial Transformer model.
[0040] In a possible implementation manner, after the passenger flow prediction model extracts features and performs feature mapping on the time-frequency domain fusion data, generating the passenger flow prediction data for the second time period includes:
[0041] Input the station type of the target subway station and the time-frequency domain fusion data into a preset CNN model, so that the CNN model performs convolution, activation, and pooling operations on the time-frequency domain fusion data, and then generates feature maps for each different time step;
[0042] Input each of the feature maps into a preset Transformer model, so that the Transformer model performs multi-layer encoding on each of the feature maps to generate high-dimensional features, and then maps the high-dimensional features to a target prediction space, and then generates passenger flow prediction data for the second time period.
[0043] In a possible implementation manner, obtaining the passenger flow prediction model by training the initial prediction model according to the historical time series data set includes:
[0044] Obtain the historical passenger flow data and historical off-station weather data of each subway station on the target subway line;
[0045] Analyze the passenger flow change trend of each subway station according to each of the historical passenger flow data, and determine the station type of each subway station;
[0046] Generate corresponding historical time domain data sequences and historical frequency domain data sequences according to each of the historical passenger flow data and each historical off-station weather data, and splice each of the historical time domain data sequences and the corresponding historical frequency domain data sequences with the subway station as the basic unit to generate corresponding historical time-frequency domain fusion data for each subway station;
[0047] Construct a training data set according to each of the time-frequency domain fusion data, the station type of each subway station, and each of the historical passenger flow data;
[0048] Iteratively train the initial prediction model several times using supervised learning based on the training dataset until a preset number of training times is reached to obtain training parameters;
[0049] Import the training parameters into the initial prediction model to obtain the passenger flow prediction model. Description of the Drawings
[0050] Figure 1 It is a schematic flowchart of a subway passenger flow prediction method based on a deep learning model provided by an embodiment of the present application.
[0051] Figure 2 It is a schematic overall flowchart of model training in a subway passenger flow prediction method based on a deep learning model provided by an embodiment of the present application.
[0052] Figure 3 It is a schematic diagram of the long-term passenger flow change of a certain subway line in a subway passenger flow prediction method based on a deep learning model provided by an embodiment of the present application.
[0053] Figure 4 It is a schematic diagram of the passenger flow change of a non-peak type station in a subway passenger flow prediction method based on a deep learning model provided by an embodiment of the present application.
[0054] Figure 5 It is a schematic diagram of the passenger flow change of a double-peak type station in a subway passenger flow prediction method based on a deep learning model provided by an embodiment of the present application.
[0055] Figure 6 It is a schematic diagram of the passenger flow change of a full-peak type station in a subway passenger flow prediction method based on a deep learning model provided by an embodiment of the present application.
[0056] Figure 7 It is a schematic flowchart of iteratively training the initial prediction model in a subway passenger flow prediction method based on a deep learning model provided by an embodiment of the present application.
[0057] Figure 8 It is a schematic structural diagram of a subway passenger flow prediction system based on a deep learning model provided by an embodiment of the present application. Detailed Embodiments
[0058] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0059] It should be noted that the step numbers in the text are only for the convenience of explaining specific embodiments and do not serve to limit the order of execution of the steps. In the description of the present application, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features.
[0060] Embodiment 1:
[0061] As Figure 1 shown, Embodiment 1 provides a subway passenger flow prediction method based on a deep learning model, including steps S1 - S3:
[0062] Step S1, obtain the passenger flow data and off - station weather data of the target subway station in the first time period;
[0063] Step S2, generate a time - domain data sequence and a frequency - domain data sequence according to the passenger flow data and the off - station weather data, and splice the time - domain data sequence and the frequency - domain data sequence to generate a time - frequency domain fusion data;
[0064] Step S3, input the time - frequency domain fusion data and the station type of the target subway station into a preset passenger flow prediction model, so that after the passenger flow prediction model extracts features and performs feature mapping on the time - frequency domain fusion data, generate the passenger flow prediction data for the second time period, where the station type is a double - peak type station, a full - peak type station or a non - peak type station;
[0065] Among them, the passenger flow prediction model is obtained by training an initial prediction model according to a historical time - series data set, and the initial prediction model is jointly constructed by an initial CNN model and an initial Transformer model.
[0066] This application provides a subway passenger flow prediction method based on a deep learning model. According to passenger flow data and off-station weather data, the passenger flow data of the target subway station for a period of time in the future can be predicted. Among them, considering that the passenger flow data and off-station weather data have time series attributes and are data that change over time, and since traditional passenger flow prediction methods all analyze the time domain, the time domain has certain advantages in capturing local correlations, but the frequency domain is also more effective in capturing the global correlations of the non-periodic sequence of subway passenger flow. Therefore, these two types of data are preprocessed to generate a time domain data sequence and a frequency domain data sequence, and then the data is spliced to obtain time-frequency domain fusion data, which improves the accuracy of subsequent passenger flow prediction. Regarding model construction, the embodiment of this application combines a CNN model and a Transformer model to form a joint prediction model, which integrates the local feature extraction ability of the CNN and the global information modeling ability of the Transformer, and can effectively predict the subway passenger flow and improve the accuracy of long-term passenger flow prediction at subway stations. In addition, the embodiment of this application also combines the station type of the subway station for prediction. Considering that the passenger flow change trends of different station types are different, during the prediction process, the station type of the target subway station and the time-frequency domain fusion data are input into the passenger flow prediction model together to further improve the accuracy of model prediction.
[0067] In a preferred embodiment, in step S1, the passenger flow data includes the number of people entering and leaving a certain subway station every 15 minutes, and the off-station weather data includes data such as temperature, air pressure, humidity, precipitation, and precipitation type.
[0068] Further, the generating a time domain data sequence and a frequency domain data sequence according to the passenger flow data and the off-station weather data, and splicing the time domain data sequence and the frequency domain data sequence to generate time-frequency domain fusion data includes:
[0069] Performing data merging and normalization processing on the passenger flow data and the off-station weather data according to the time sequence to obtain the time domain data sequence;
[0070] Performing a Fourier transform on the time domain data sequence to generate the frequency domain data sequence;
[0071] Splicing the time domain data sequence and the frequency domain data sequence along the feature channels of a preset dimension to generate the time-frequency domain fusion data.
[0072] The embodiment of the present application provides a method for generating time-frequency domain fusion data, which introduces the fast Fourier transform to solve the problem of discrete frequency offset. The fast Fourier transform is used to transform the tiled time-domain data sequence into the frequency domain to generate the corresponding frequency-domain data sequence, so as to fully extract the feature information in the data. Then, the time-domain data sequence and the frequency-domain data sequence are merged to form a new feature layer, and a time-frequency domain fusion data containing both local features and global features is obtained, thereby improving the accuracy of subsequent passenger flow prediction.
[0073] In a preferred embodiment, after obtaining the passenger flow data and the off-station weather data, data preprocessing is performed on the data. After removing the null data, the data formats are unified. The key item for integrating the two data sets is the time item. First, the time item format is unified, and then the two data sets are integrated according to the time item. Next, the hour, minute timestamps, as well as the weekday, weekend, and adjusted holiday time data are extracted from the date data, and the data set is normalized to obtain the time-domain data sequence.
[0074] In order to obtain the frequency-domain feature data of the original data, it is necessary to perform a fast Fourier transform (FFT) on the time-domain sequence of passenger flow at a single station with a length of. The radix-2 algorithm selected by time is adopted. First, the data with a length of is divided into odd and even groups, that is:
[0075]
[0076] where r = 0, 1,..., Q, x1(r) is the even time sequence, and x2(r) is the odd time sequence. According to the above formula, the first half of the frequency-domain data X(k) is calculated as:
[0077]
[0078] where k = 0, 1,..., N / 2 - 1, M ≥ Q. When M ≥ Q, M - Q zero-value points are supplemented. is the rotation factor, which is obtained by the following formula:
[0079]
[0080] The second half of X(k) is
[0081]
[0082] where k = N / 2,..., N - 1. After merging the first half and the second half of X(k), it is the frequency-domain data obtained after the fast Fourier transform.
[0083] A channel with a dimension of 1 is set, and the time-domain feature and the frequency-domain feature are spliced along this channel to obtain a 2D feature matrix, that is, the time-frequency domain fusion data.
[0084] In a possible implementation manner, in step S3, after the passenger flow prediction model extracts features and performs feature mapping on the time-frequency domain fusion data, passenger flow prediction data for a second time period is generated, including:
[0085] Input the station type of the target subway station and the time-frequency domain fusion data into a preset CNN model, so that the CNN model performs convolution, activation, and pooling operations on the time-frequency domain fusion data, and then generates feature maps for each different time step;
[0086] Input each of the feature maps into a preset Transformer model, so that the Transformer model performs multi-layer encoding on each of the feature maps to generate high-dimensional features, and then maps the high-dimensional features to the target prediction space, and then generates passenger flow prediction data for the second time period.
[0087] Further, the CNN model performs convolution, activation, and pooling operations on the time-frequency domain fusion data, and then generates feature maps for each different time step, including:
[0088] Use one-dimensional convolution to extract the local time series features and local frequency domain features of the time-frequency domain fusion data, and obtain a number of single-step feature maps;
[0089] Perform activation and pooling operations on each of the single-step feature maps through a preset activation function and pooling window to generate feature maps for each different time step.
[0090] The embodiment of the present application provides a method for generating passenger flow prediction data. First, the CNN model performs convolution, activation, and pooling operations on the time-frequency domain fusion data according to the station type to extract each local feature in the data and generate feature maps for each different time step. Then, the Transformer model is used to capture the global time series dependence relationship of the data, extract global features from each feature map and perform multi-layer encoding to obtain high-dimensional features. Finally, the high-dimensional features are mapped to the target prediction space to compress complex feature information and generate the final prediction result. Through this process, the dimensions of the input and output are adjusted, so that the features extracted from the Transformer model are accurately converted into passenger flow prediction data for the second time period, thereby completing the prediction task of the entire model. The embodiment of the present application combines the CNN model and the Transformer model, overcomes the limitations of the RNN in processing long sequence data, and avoids the problems of gradient disappearance or gradient explosion existing in traditional RNN or LSTM models when modeling long sequences, thereby improving the prediction ability of long-term passenger flow trends.
[0091] In a preferred embodiment, the time-frequency domain fusion data after the FFT framework is passed to the CNN module. First, one-dimensional convolution is used to extract the local time-series and frequency-domain features to extract the short-term trend of the passenger flow. The input sequence length (time step) is set to the data sample size of two days, the input features are the data features selected by the preprocessing, the prediction output is single-step prediction, and the size of the convolution kernel is set to 3 to facilitate capturing short-term dependencies. To reduce the computational overhead, the convolutional feature map (the feature map obtained after extracting the local time-frequency domain features through the CNN convolutional layer) is activated (through an activation function used to introduce non-linear features) and then pooled to reduce the size of the feature map (the pooling window size is set to 2, that is, the maximum value of every 2 consecutive elements is taken, and the feature dimension is reduced through max pooling). The pooled feature map is imported into the Transformer model, and positional encoding is added to it to ensure that the features of each time step carry the positional information of the sequence position (converted into a tensor format of [the number of samples used in the iteration, sequence length, feature dimension]).
[0092] The Transformer model is stacked by multiple Transformer Encoder Layer layers, and each layer is composed of several modules: First, the multi-head self-attention mechanism can effectively capture the global time dependencies in the sequence and simultaneously model the dynamic associations between external features to help the model learn the correlations between different positions in the input sequence. Second, the feed-forward neural network processes the features of each time step step by step to further enhance the feature representation ability. In addition, layer normalization and residual connections play a key role in each layer, which can stabilize the training process of the model, avoid the problems of gradient disappearance or explosion, and at the same time promote the optimization of the deep network. After multiple layers of processing by the Transformer encoder, the model not only retains the time-series information of the sequence but also successfully learns the feature representation of long-term dependencies.
[0093] These high-dimensional features output from the Transformer encoder are then input into the fully connected layer for further processing. The main function of the fully connected layer is to map the high-dimensional feature space to the target prediction space. Its core role is to compress complex feature information and generate the final prediction result. Through this process, the fully connected layer can adjust the dimensions of the input and output, so that the features extracted from the Transformer module are accurately converted into target prediction values, thus completing the prediction task of the entire model.
[0094] In a possible implementation manner, training the initial prediction model according to the historical time-series data set to obtain the passenger flow prediction model includes:
[0095] Obtain the historical passenger flow data and historical off-station weather data of each subway station on the target subway line;
[0096] Analyze the passenger flow change trends of each subway station based on the respective historical passenger flow data, and determine the station types of each subway station;
[0097] Generate corresponding historical time-domain data sequences and historical frequency-domain data sequences according to the respective historical passenger flow data and historical weather data outside the stations, and splice each of the historical time-domain data sequences and the corresponding historical frequency-domain data sequences with the subway station as the basic unit to generate historical time-frequency domain fusion data corresponding to each subway station;
[0098] Construct a training dataset based on each of the time-frequency domain fusion data, the station types of each subway station, and the respective historical passenger flow data;
[0099] Perform several iterative trainings on the initial prediction model in a supervised learning manner according to the training dataset until the preset number of training times is reached to obtain training parameters;
[0100] Import the training parameters into the initial prediction model to obtain the passenger flow prediction model.
[0101] The embodiment of the present application provides a training method for an initial prediction model. First, determine the station types of each subway station through the analysis of each historical passenger flow data. The passenger flow change trends of different station types are different. Determining the station type in advance and constructing it as a part of the training data can assist the model in understanding the passenger flow change characteristics of different types of stations, and improve the training efficiency and prediction accuracy of the model. Then, according to the training data, perform several iterative trainings on the initial prediction model in a supervised learning manner, and control the training duration of the model through the preset number of training times to avoid the model falling into overfitting due to too long training time, and improve the training efficiency of the model.
[0102] Further, during the process of performing several iterative trainings on the initial prediction model, judge whether to terminate the training in advance through an early stopping mechanism.
[0103] The embodiment of the present application further introduces an early stopping mechanism to prevent the model from falling into overfitting. Specifically, the early stopping mechanism is that after each training round, the validation set is evaluated and its loss value is calculated. If the validation set loss continues to decrease, it means that the model is still effectively learning; if the validation set loss stops decreasing or starts to increase, it indicates that the model may start to overfit. At this time, even if the preset number of training times is not reached, the training will be terminated in advance to avoid the model performing ineffective overfitting training and improve the training efficiency and robustness of the model.
[0104] Further, during several iterative trainings of the initial prediction model, a preset proportion of neurons in the initial prediction model are randomly discarded through the Dropout regularization technique.
[0105] The embodiment of the present application further introduces the Dropout regularization technique to reduce overfitting. Specifically, the Dropout regularization technique randomly discards some neurons in the neural network at each training stage, forcing the network to learn more robust feature representations, gradually reducing the training parameters during the training process, and improving the training efficiency and robustness of the model.
[0106] In a preferred embodiment, the training process of the initial prediction model is as Figure 2 shown. First is the data preparation. Import the passenger flow data and weather conditions of a subway line in Guangzhou (the passenger flow data includes the number of people entering and leaving each subway station on a certain subway line every 15 minutes, and the weather conditions include temperature, air pressure, humidity, precipitation, precipitation type), and perform data preprocessing. Generate a weekly passenger flow change chart of subway stations based on the passenger flow data. The long-term passenger flow change charts of 7 subway stations on the subway line are as Figure 3 shown. According to the passenger flow data, the 7 stations are divided into 3 types of station types: double-peak type stations, where the passenger flow will rise sharply during the morning rush hour, reaching 10,000 person-times. This type of station is mainly the stations within the working area; full-peak type stations, where the passenger flow basically remains above 3,000 person-times and will rise sharply during the morning rush hour, reaching 25,000 person-times. This type of station is mostly transfer stations, integrated stations of the working area and the business district; non-peak type stations, where the passenger flow basically remains below 5,000 person-times. The passenger flow trend charts of different station types are as Figures 4 - 6 shown, where Figure 4 represents non-peak type stations, Figure 5 represents double-peak type stations, Figure 6 represents full-peak type stations. Extract the hour, minute timestamps, as well as the weekday, weekend, and adjusted holiday time data for the date data to cope with the double-peak and full-peak types of the subway; then transfer the processed data to the FFT model to generate historical time-frequency domain fusion data.
[0107] Then, use the historical time-frequency domain fusion data to perform several iterative trainings on the initial prediction model constructed based on the CNN model and the Transformer model. The training step process is as Figure 7As shown, the model adopts the supervised learning method (the model is trained under the guidance of labeled data. By learning the mapping relationship between input features and corresponding labels, accurate prediction of unseen data is achieved). The input features are the historical time-domain data sequence of the passenger flow of a single subway station, time features (hour, day of the week, whether it is a holiday), weather features (temperature, humidity, precipitation, precipitation type, etc.), and the historical frequency-domain data sequence obtained through fast Fourier transform, and the data is normalized. The output label is the passenger flow of the next time step of the subway station, and the number of output time steps can be set according to the actual situation. A single time step is 15 minutes. During the training process, the Min-Max Normalization method is adopted Compress the input features into the interval [0,1]; use the sliding window technique to construct a fixed-size window on the data, extract segments from the data in units of the window, train the input of each window, and output the training results of this time step; define the loss function to evaluate the difference between the predicted value and the true value of the model; use the objective function (loss function) to guide the model optimization.
[0108] Specifically, the loss function is defined as the mean square error (MSE); the Adam optimizer is adopted to dynamically adjust the learning rate and improve the training efficiency. The key formula of Adam is:
[0109] 1. Calculate the gradient of the current parameter θ t :
[0110]
[0111] where g t is the gradient of the t-th iteration, and J(θ t ) is the loss function.
[0112] 2. Bias correction, first-order momentum correction:
[0113] m t =β1m t-1 +(1 - β1)g t
[0114]
[0115] where β1 is the decay rate of the first-order momentum, and m t is the exponential weighted average of the gradient.
[0116] 3. Second-order momentum correction:
[0117]
[0118] where β2 is the decay rate of the second-order momentum, and v tIs the exponentially weighted average of the squared gradient.
[0119] 4. Parameter update:
[0120]
[0121] Where η is the learning rate and ∈ is an extremely small number used to prevent the denominator from being zero.
[0122] Furthermore, during the training process of the model, this embodiment introduces an early stopping mechanism to prevent overfitting and improve training efficiency. Specifically, the validation set is not immediately used for validation at the beginning of training. Usually, after completing the first full training epoch, the loss of the validation set begins to be calculated. After each training epoch, the validation set is evaluated and its loss value is calculated. If the validation set loss continues to decrease, it indicates that the model is still effectively learning; if the validation set loss stops decreasing or starts to increase, it indicates that the model may start to overfit. To avoid premature stopping of training, a patience value is set, allowing the training to continue when the validation set loss has not improved for several training epochs. The patience value in this article is 10 epochs, that is, if the validation set loss does not significantly decrease in 10 consecutive training epochs, the training process will stop early, thus avoiding overfitting and saving computing resources. In this way, the early stopping mechanism can effectively ensure that the model stops training under optimal conditions, avoid overfitting while improving training efficiency, and ultimately enhance the generalization ability of the model.
[0123] Furthermore, during the model training process, the embodiments of the present application also introduce the Dropout regularization technique to reduce overfitting. Specifically, by randomly discarding some neurons in the neural network at each training stage, the network is forced to learn more robust feature representations. During the training stage, the output of each neuron is multiplied by a randomly generated binary matrix (mask), where 0 indicates that the neuron is discarded and 1 indicates that the neuron is retained. The discarded proportion is controlled by the hyperparameter p, and common p values are usually set between 0.1 and 0.5, that is, 10% to 50% of the neurons are randomly discarded each time training. The present invention selects to discard 10%, which can effectively reduce the model's dependence on certain specific neurons, thereby reducing the risk of overfitting. The dropout formula during training is output = input × mask, where mask is a binary matrix indicating which neurons are discarded; in order to maintain consistency during the test stage, neurons are not discarded, but the output of each neuron is multiplied by a factor 1 - p to compensate for the impact of discarding neurons during the training stage and ensure more accurate prediction results during testing. During testing, it is desired to use all neurons for inference, and discarding neurons may lead to unstable prediction results. Therefore, no neurons are selected to be discarded, but the activation values are scaled to avoid the bias caused by discarding neurons. The dropout formula during testing is output = input × (1 - p) ((1 - p) is the scaling factor), and it is set to 10% as consistent with the above. In this way, Dropout can effectively prevent overfitting during the training stage and ensure the stability of the model and the accuracy of prediction during the test stage.
[0124] Furthermore, after importing the training parameters into the initial prediction model to obtain the passenger flow prediction model, a preset test set is used to comprehensively evaluate the prediction performance of the passenger flow prediction model, including:
[0125] Predict the test set and inverse normalize the normalized values to restore the actual passenger flow; then calculate evaluation metrics (such as RMSE, MAE, MAPE, R 2 etc.). The specific calculation formulas are as follows:
[0126] RMSE (Root Mean Square Error), used to measure the prediction error of the regression model (the smaller the RMSE, the closer the prediction result of the model is to the true value, and the better the model performance). Calculation formula:
[0127]
[0128] MAE (Mean Absolute Error), measures the average of the absolute values of the errors between the predicted values and the true values, and is used to evaluate the size of the prediction error of the model. Calculation formula:
[0129]
[0130] MAPE (Mean Absolute Percentage Error) measures the percentage of the prediction error of the model in relation to the true value, expressed as the percentage of the absolute value of the error in relation to the true value. Calculation formula:
[0131]
[0132] R 2 (Coefficient of determination), an indicator that measures the goodness of fit of a regression model, representing the correlation between the prediction results of the model and the actual observed values. R 2 Ranges between 0 and 1. The closer the value is to 1, the better the fitting effect of the model. Calculation formula:
[0133]
[0134] In summary, compared with the prior art, the embodiments of the present application consider that the frequency domain is also more effective in capturing the global correlation of the non-periodic sequence of subway passenger flow. By combining time-domain convolution and frequency-domain convolution feature extraction, introducing the Fast Fourier Transform to extract frequency-domain convolution features, and combining time-frequency domain features through a convolutional neural network, local patterns and short-term trends can be automatically extracted from sequence data, such as the peak, trough, or short-term fluctuation characteristics of passenger flow. Through the Transformer module, the embodiments of the present application can model long-term dependencies and global features, effectively identify the long-term trends and complex time dependencies of subway passenger flow. By using the sliding window technique, the time series data is converted into a supervised learning problem, enabling the model to focus more on input data with strong time dependencies, thereby having higher prediction accuracy; strong data adaptability. The embodiments of the present application adopt a normalization preprocessing method to enable the model to more stably process input features with different dimensions and numerical ranges, and are applicable to diverse subway passenger flow data; support multi-dimensional inputs (such as multi-modal features like time, weather, date, etc.), and flexibly fuse these features through fully connected layers and the Transformer module to further improve the generalization ability of the model; introduce the Adam optimizer and combine the early stopping mechanism to effectively control the training time and complexity of the model, avoiding the overfitting problem. Moreover, compared with traditional machine learning methods, the training efficiency is significantly improved; the technical solution of the present invention can not only perform single-step prediction but also be extended to multi-step prediction tasks, suitable for the common multi-time-step requirements in subway passenger flow prediction.
[0135] Embodiment 2:
[0136] As Figure 8 shown, Embodiment 2 provides a subway passenger flow prediction system based on a deep learning model, including an acquisition module 10, a data processing module 20, and a prediction module 30;
[0137] Among them, the obtaining module 10 is used to obtain the passenger flow data and the off-station weather data of the target subway station in the first time period;
[0138] The data processing module 20 is used to generate a time-domain data sequence and a frequency-domain data sequence according to the passenger flow data and the off-station weather data, and splice the time-domain data sequence and the frequency-domain data sequence to generate a time-frequency domain fusion data;
[0139] The prediction module 30 is used to input the time-frequency domain fusion data and the station type of the target subway station into a preset passenger flow prediction model, so that after the passenger flow prediction model extracts and maps the features of the time-frequency domain fusion data, it generates the passenger flow prediction data for the second time period, and the station type is a double-peak type station, a full-peak type station or a non-peak type station;
[0140] Among them, the passenger flow prediction model is obtained by training an initial prediction model according to a historical time series data set, and the initial prediction model is jointly constructed by an initial CNN model and an initial Transformer model.
[0141] Further, the data processing module 20 generates a time-domain data sequence and a frequency-domain data sequence according to the passenger flow data and the off-station weather data, and splices the time-domain data sequence and the frequency-domain data sequence to generate a time-frequency domain fusion data, including:
[0142] Merge and normalize the passenger flow data and the off-station weather data according to the time order to obtain the time-domain data sequence;
[0143] Perform Fourier transform on the time-domain data sequence to generate the frequency-domain data sequence;
[0144] Splice the time-domain data sequence and the frequency-domain data sequence along the feature channels of a preset dimension to generate the time-frequency domain fusion data.
[0145] In a possible implementation manner, after the passenger flow prediction model extracts and maps the features of the time-frequency domain fusion data, it generates the passenger flow prediction data for the second time period, including:
[0146] Input the station type of the target subway station and the time-frequency domain fusion data into a preset CNN model, so that the CNN model performs convolution, activation and pooling operations on the time-frequency domain fusion data, and then generates feature maps at different time steps;
[0147] Input each of the feature maps into a preset Transformer model, so that the Transformer model performs multi-layer encoding on each of the feature maps to generate high-dimensional features, and then maps the high-dimensional features to a target prediction space, thereby generating passenger flow prediction data for the second time period.
[0148] Further, the CNN model performs convolution, activation, and pooling operations on the time-frequency domain fusion data, and then generates feature maps for each different time step, including:
[0149] Use one-dimensional convolution to extract the local time series features and local frequency domain features of the time-frequency domain fusion data to obtain a number of single-step feature maps;
[0150] Through a preset activation function and pooling window, perform activation and pooling operations on each of the single-step feature maps to generate feature maps for each different time step.
[0151] In a possible implementation manner, the training of the initial prediction model according to the historical time series data set to obtain the passenger flow prediction model includes:
[0152] Obtain the historical passenger flow data and historical off-station weather data of each subway station on the target subway line;
[0153] Analyze the passenger flow change trend of each subway station according to each of the historical passenger flow data, and determine the station type of each subway station;
[0154] Generate corresponding historical time domain data sequences and historical frequency domain data sequences according to each of the historical passenger flow data and each historical off-station weather data, and splice each of the historical time domain data sequences and the corresponding historical frequency domain data sequences with the subway station as the basic unit to generate each historical time-frequency domain fusion data corresponding to each subway station;
[0155] Construct a training data set according to each of the time-frequency domain fusion data, the station type of each subway station, and each of the historical passenger flow data;
[0156] Use the training data set to perform several iterative trainings on the initial prediction model in a supervised learning manner until a preset number of training times is reached to obtain training parameters;
[0157] Import the training parameters into the initial prediction model to obtain the passenger flow prediction model.
[0158] Further, during the several iterative trainings of the initial prediction model, judge whether to terminate the training in advance through an early stopping mechanism.
[0159] Further, during several iterative trainings of the initial prediction model, a preset proportion of neurons in the initial prediction model are randomly discarded through the Dropout regularization technique.
[0160] This application provides a subway passenger flow prediction system based on a deep learning model, which predicts the passenger flow data of a target subway station in a future period according to the passenger flow data and the off-station weather data. Among them, considering that the passenger flow data and the off-station weather data have time series attributes and are data that change over time, and since traditional passenger flow prediction methods all analyze the time domain, the time domain has certain advantages in capturing local correlations, but the frequency domain is also more effective in capturing the global correlations of the non-periodic sequence of subway passenger flow. Therefore, these two types of data are preprocessed to generate a time domain data sequence and a frequency domain data sequence, and then data splicing is performed to obtain time-frequency domain fusion data, improving the accuracy of subsequent passenger flow prediction. Regarding model construction, the embodiment of this application combines a CNN model and a Transformer model to form a joint prediction model, integrating the local feature extraction ability of the CNN and the global information modeling ability of the Transformer, which can effectively predict the subway passenger flow and improve the accuracy of long-term passenger flow prediction at subway stations. In addition, the embodiment of this application also combines the station type of the subway station for prediction. Considering that the passenger flow change trends of different station types are different, the station type of the target subway station and the time-frequency domain fusion data are input into the passenger flow prediction model together during the prediction process, further improving the accuracy of model prediction.
[0161] The more detailed working principle and step flow of this embodiment can but are not limited to referring to the relevant records in Embodiment 1.
[0162] The above specific embodiments further elaborate on the purpose, technical solutions, and beneficial effects of this application. It should be understood that the above are only specific embodiments of this application and are not used to limit the protection scope of this application. In particular, it is pointed out that for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this application shall be included in the protection scope of this application.
Claims
1. A subway passenger flow prediction method based on a deep learning model, characterized in that: include: Obtain passenger flow data and weather data outside the target subway station in the first time period; Generate a time domain data sequence and a frequency domain data sequence according to the passenger flow data and the off-station weather data, and perform data splicing on the time domain data sequence and the frequency domain data sequence to generate time-frequency domain fusion data; Inputting the time-frequency domain fusion data and the station type of the target subway station into a preset passenger flow prediction model, so that the passenger flow prediction model generates passenger flow prediction data for a second time period after performing feature extraction and feature mapping on the time-frequency domain fusion data, wherein the station type is a bimodal station, a full-peak station, or a non-peak station; The passenger flow prediction model is obtained by training an initial prediction model based on a historical time series data set, and the initial prediction model is jointly constructed by an initial CNN model and an initial Transformer model.
2. A subway passenger flow prediction method based on a deep learning model as claimed in claim 1, characterized in that: The step of generating a time domain data sequence and a frequency domain data sequence according to the passenger flow data and the off-station weather data, and splicing the time domain data sequence and the frequency domain data sequence to generate time-frequency domain fusion data includes: Merging and normalizing the passenger flow data and the off-station weather data according to time sequence to obtain the time domain data sequence; Performing Fourier transform on the time domain data sequence to generate the frequency domain data sequence; The time domain data sequence and the frequency domain data sequence are concatenated along a feature channel of a preset dimension to generate the time-frequency domain fusion data.
3. The method for predicting subway passenger flow based on a deep learning model according to claim 1, characterized in that: The passenger flow prediction model performs feature extraction and feature mapping on the time-frequency domain fusion data to generate passenger flow prediction data for the second time period, including: Inputting the station type of the target subway station and the time-frequency domain fusion data into a preset CNN model, so that the CNN model performs convolution, activation and pooling operations on the time-frequency domain fusion data, thereby generating feature maps at different time steps; Each of the feature maps is input into a preset Transformer model so that the Transformer model performs multi-layer encoding on each of the feature maps to generate high-dimensional features, and then maps the high-dimensional features to a target prediction space, thereby generating passenger flow prediction data for a second time period.
4. A subway passenger flow prediction method based on a deep learning model as claimed in claim 3, characterized in that: The CNN model performs convolution, activation and pooling operations on the time-frequency domain fusion data to generate feature maps at different time steps, including: Using one-dimensional convolution to extract local time series features and local frequency domain features of the time-frequency domain fusion data, to obtain a number of single-step feature maps; Through the preset activation function and pooling window, each of the single-step feature maps is activated and pooled to generate feature maps of different time steps.
5. The method for predicting subway passenger flow based on a deep learning model according to claim 1, characterized in that: The step of training the initial prediction model according to the historical time series data set to obtain the passenger flow prediction model includes: Obtain the historical passenger flow data and historical weather data outside each subway station on the target subway line; Analyzing the passenger flow change trend of each subway station according to each of the historical passenger flow data, and determining the station type of each subway station; Generate corresponding historical time domain data sequences and historical frequency domain data sequences according to each of the historical passenger flow data and each of the historical weather data outside the station, and splice each of the historical time domain data sequences and the corresponding historical frequency domain data sequences with the subway station as the basic unit to generate each of the historical time and frequency domain fusion data corresponding to each subway station; Constructing a training data set according to each of the time-frequency domain fusion data, the station type of each subway station and each of the historical passenger flow data; Performing several iterative training on the initial prediction model using a supervised learning method according to the training data set until a preset number of training times is reached to obtain training parameters; The training parameters are imported into the initial prediction model to obtain the passenger flow prediction model.
6. A subway passenger flow prediction method based on a deep learning model as claimed in claim 5, characterized in that: During several iterations of training of the initial prediction model, an early stopping mechanism is used to determine whether to terminate the training in advance.
7. A subway passenger flow prediction method based on a deep learning model as claimed in claim 5, characterized in that: During several iterations of training of the initial prediction model, a preset proportion of neurons in the initial prediction model are randomly discarded by using the Dropout regularization technique.
8. A subway passenger flow prediction system based on a deep learning model, characterized in that: It includes an acquisition module, a data processing module and a prediction module; Wherein, the acquisition module is used to acquire the passenger flow data and weather data outside the target subway station in the first time period; The data processing module is used to generate a time domain data sequence and a frequency domain data sequence according to the passenger flow data and the weather data outside the station, and perform data splicing on the time domain data sequence and the frequency domain data sequence to generate time-frequency domain fusion data; The prediction module is used to input the time-frequency domain fusion data and the site type of the target subway station into a preset passenger flow prediction model, so that the passenger flow prediction model generates passenger flow prediction data for the second time period after performing feature extraction and feature mapping on the time-frequency domain fusion data, and the site type is a bimodal site, a full-peak site or a non-peak site; The passenger flow prediction model is obtained by training an initial prediction model based on a historical time series data set, and the initial prediction model is jointly constructed by an initial CNN model and an initial Transformer model.
9. A subway passenger flow prediction system based on a deep learning model as claimed in claim 8, characterized in that: The passenger flow prediction model performs feature extraction and feature mapping on the time-frequency domain fusion data to generate passenger flow prediction data for the second time period, including: Inputting the station type of the target subway station and the time-frequency domain fusion data into a preset CNN model, so that the CNN model performs convolution, activation and pooling operations on the time-frequency domain fusion data, thereby generating feature maps at different time steps; Each of the feature maps is input into a preset Transformer model so that the Transformer model performs multi-layer encoding on each of the feature maps to generate high-dimensional features, and then maps the high-dimensional features to a target prediction space, thereby generating passenger flow prediction data for a second time period.
10. A subway passenger flow prediction system based on a deep learning model as claimed in claim 8, characterized in that: The step of training the initial prediction model according to the historical time series data set to obtain the passenger flow prediction model includes: Obtain the historical passenger flow data and historical weather data outside each subway station on the target subway line; Analyzing the passenger flow change trend of each subway station according to each of the historical passenger flow data, and determining the station type of each subway station; Generate corresponding historical time domain data sequences and historical frequency domain data sequences according to each of the historical passenger flow data and each of the historical weather data outside the station, and splice each of the historical time domain data sequences and the corresponding historical frequency domain data sequences with the subway station as the basic unit to generate each of the historical time and frequency domain fusion data corresponding to each subway station; Constructing a training data set according to each of the time-frequency domain fusion data, the station type of each subway station and each of the historical passenger flow data; Performing several iterative training on the initial prediction model using a supervised learning method according to the training data set until a preset number of training times is reached to obtain training parameters; The training parameters are imported into the initial prediction model to obtain the passenger flow prediction model.
Citation Information
Cited By
Lightweight transportation junction passenger flow prediction system and method
CN120633955A