Rail transit scheduling method and system based on GA optimization Transform

By using a Transformer model optimized based on genetic algorithms, combined with multidimensional features and a self-attention mechanism, the accuracy and adaptability issues of rail transit passenger flow prediction were solved. This achieved high-precision and high-stability prediction in complex scenarios, meeting the dynamic scheduling requirements of large-scale networked operations.

CN121094431APending Publication Date: 2025-12-09SHANGHAI INST OF TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511230355.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2025-12-09

AI Technical Summary

Technical Problem

Existing methods for predicting passenger flow in rail transit are unable to guarantee high accuracy and dynamic adaptability when faced with complex and ever-changing passenger flow patterns. This results in scheduling schemes lacking accuracy and generalization ability, and failing to meet the dynamic scheduling needs under large-scale networked operations.

Method used

We employ a Transformer model optimized based on a genetic algorithm. By constructing multi-dimensional features and utilizing a multi-head self-attention mechanism, we capture long-distance spatiotemporal dependencies. Combined with K-means clustering technology, we classify stations into commercial centers, transfer stations, and ordinary stations. We introduce a time-differentiated fitness evaluation mechanism and optimize hyperparameters to improve prediction accuracy and applicability.

Benefits of technology

It significantly improves the accuracy and applicability of passenger flow forecasting, provides minute-level precision support under complex passenger flow patterns, meets the dynamic scheduling needs under large-scale networked operations, and enhances the accuracy and generalization capability of rail transit scheduling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121094431A_ABST
    Figure CN121094431A_ABST
Patent Text Reader

Abstract

The invention relates to a rail transit scheduling method and a rail transit scheduling system based on GA (Genetic Algorithm) optimization Transform. The method comprises the steps of firstly collecting station passenger flow information and preprocessing to obtain a passenger flow sequence; constructing multi-dimensional features including time, space and external weather, encoding the multi-dimensional features, and fusing the encoded multi-dimensional features with the passenger flow sequence to obtain a feature fusion vector; then, constructing a Transform model based on a multi-head self-attention mechanism, inputting the feature fusion vector, and optimizing the model by using a genetic algorithm; and finally, inputting the real-time data after preprocessing and feature extraction into the optimization model to obtain a passenger flow prediction result, and carrying out rail transit scheduling according to the passenger flow prediction result. Compared with the prior art, the method has the advantages that the accuracy and generalization ability of rail transit dispatching are improved, the defects in the aspects of long-time sequence modeling, automatic hyper-parameter optimization and rail transit complex topological structure adaptability in the prior art are effectively overcome, and the dynamic dispatching requirement under large-scale network operation is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of rail transit scheduling technology, and in particular to a rail transit scheduling method and system based on GA-optimized Transformer. Background Technology

[0002] Rail transit passenger flow forecasting is a crucial component of intelligent transportation systems, playing a key role in the planning, scheduling, and management of urban rail transit. Accurate passenger flow forecasting helps subway operators optimize train schedules, improve transportation efficiency, reduce operating costs, and effectively respond to peak passenger flows and unforeseen circumstances. The accuracy of passenger flow forecasting directly determines the rationality of rail transit scheduling. Significant deviations between forecasted and actual passenger flows can lead to insufficient train capacity during peak hours or wasted capacity during off-peak hours, potentially impacting passenger flow management efficiency at transfer stations and increasing passenger organization risks.

[0003] However, due to the complexity of urban rail transit systems, passenger flow data has nonlinear, time-varying, and multidimensional characteristics, making it difficult for traditional passenger flow forecasting methods to guarantee high-precision forecasting results when faced with complex and ever-changing passenger flow patterns. This results in scheduling schemes lacking dynamic adaptability and having low scheduling accuracy.

[0004] With the rapid development of artificial intelligence and deep learning technologies, deep learning methods based on time series modeling have gradually become a research hotspot. These include Recurrent Neural Networks (RNNs), Long Short-Term Memory Networks (LSTMs), and Gated Recurrent Units (GRUs), which have demonstrated good performance in capturing time dependencies and have been initially applied to passenger flow prediction to assist scheduling decisions. However, these methods still have certain limitations: for example, the vanishing gradient problem during training makes it difficult to capture long-term passenger flow patterns; insufficient parallel computing power results in a lag in prediction response speed in real-time scheduling scenarios; and limited ability to model long-sequence dependencies makes it impossible to accurately correlate passenger flow coupling relationships across multiple lines and transfer points. These shortcomings directly affect the timeliness and reliability of passenger flow prediction results, thus restricting the level of precision in rail transit scheduling.

[0005] In summary, the accuracy of current rail transit scheduling still needs to be improved, and the current scheduling methods lack generalization ability, making it difficult to meet the dynamic scheduling needs under large-scale network operation. Summary of the Invention

[0006] The purpose of this invention is to overcome the shortcomings of the existing technology and provide a rail transit scheduling method and system based on GA-optimized Transformer.

[0007] The objective of this invention can be achieved through the following technical solutions:

[0008] According to one aspect of the present invention, a rail transit scheduling method based on GA-optimized Transformer is provided, the method steps including:

[0009] S1. Collect station passenger flow information and preprocess it to obtain passenger flow sequence;

[0010] S2. Construct a multi-dimensional feature that includes time dimension features, spatial dimension features and external weather features, and then encode the multi-dimensional feature and fuse it with the passenger flow sequence to obtain a feature fusion vector;

[0011] S3. Construct a Transformer model based on a multi-head self-attention mechanism;

[0012] S4. Input the feature fusion vector into the constructed Transformer model, optimize the model based on the genetic algorithm, and output the optimized Transformer model.

[0013] S5. After preprocessing and feature extraction, the real-time data is input into the optimized Transformer model, and the passenger flow prediction result is output. Based on the prediction result, rail transit scheduling is carried out.

[0014] As a preferred technical solution, the station passenger flow information in S1, i.e., rail transit AFC data, includes at least the card swiping time, station number, card swiping device number, entry and exit status, and card swiping type.

[0015] As a preferred technical solution, data preprocessing in S1 includes data cleaning, time format conversion, and normalization. Data cleaning includes filling in missing values, removing duplicate values, and correcting outliers. Time series format conversion uses a sliding window mechanism to divide the sample sequence, with a window length set to 60 minutes. Normalization uses the Min-Max normalization method, as shown in the following formula:

[0016]

[0017] Among them, X min and X max X and X′ represent the minimum and maximum values ​​of historical data for a certain site, respectively. X is the original data value of the site, and X′ is the standard normalized value of the site.

[0018] As a preferred technical solution, among the multidimensional features in S2:

[0019] Time-related features include weekday, non-working day, and holiday features;

[0020] The spatial dimension features are defined by using the K-means clustering algorithm to classify stations into commercial center stations, transfer stations, and ordinary stations based on historical passenger flow, volatility, and time-of-day changes.

[0021] External weather characteristics include rainfall, wind chill index, and overall comfort index.

[0022] The feature fusion vector in S2 is in matrix form:

[0023]

[0024] The time step T = 4, which is a 60-minute window containing four 15-minute time intervals; the feature dimension d = 32.

[0025] As a preferred technical solution, the Transformer model in S3 includes an encoder and a decoder; the encoder is used to input the sequence and model it, and the decoder is used to complete the prediction of future passenger flow and output it; the input of the decoder layer is the historical passenger flow, which passes through the attention mechanism in the Transformer to calculate and output the temporal constraint feature matrix. The temporal constraint feature matrix is ​​used to calculate and output the fusion feature matrix through the activation function. After multiple training and normalization, the fusion feature matrix outputs the final predicted passenger flow.

[0026] As a preferred technical solution, the specific steps for optimizing the model in S4 based on the genetic algorithm include:

[0027] S41. Initialize the population and generate multiple different hyperparameter combinations. Each chromosome represents a set of hyperparameter combinations, and each chromosome contains 5 genes. These 5 genes correspond to the hyperparameters of the model: batch size, learning rate, number of encoder layers, number of decoder layers, and matrix initialization gain.

[0028] S42. Based on the fitness evaluation mechanism with time-time differences, construct a weighted root mean square error function, and construct the fitness function from the weighted root mean square error function;

[0029] S43. Selection, crossover, and mutation processes based on fitness functions;

[0030] S44. When the fitness improvement effect is less than 1% for 10 consecutive generations or the maximum number of iterations is reached, output the chromosome with the highest fitness as the final configuration, that is, output the current batch size, learning rate, number of encoder and decoder layers and initial gain.

[0031] As a preferred technical solution, in S41, the range of hyperparameter values ​​meets the constraints of rail transit. Specifically, the batch size is 16, 32, or 64, the learning rate is 0.0001-0.005, the number of encoder or decoder layers is 3-5, and the initialization gain is 1.0-1.5.

[0032] As a preferred technical solution, the specific formulas for the weighted root mean square error function and fitness function constructed in S42 are as follows:

[0033]

[0034] Where RMSE is the weighted root mean square error function, y i This represents the actual passenger flow value. For predicted values, from Inverse normalization is used to represent passenger flow, 4 represents the number of chromosomes, Fitness is the fitness function, and W 时段 (i) represents the time period weighting coefficient. For transfer stations, the weights for morning peak, evening peak, and ordinary time periods are 1.8, 1.5, and 1.0, respectively, while for commercial stations, they are 1.3, 1.7, and 1.0, respectively.

[0035] As a preferred technical solution, the rail transit scheduling in S5 based on the prediction results specifically includes:

[0036] Determine the passenger flow threshold for the predicted station, then select the peak passenger flow forecast and the maximum passenger capacity per train from the forecast results, and calculate whether to add more trains or keep the original number of trains. The specific formula is as follows:

[0037]

[0038] In the formula, N add This represents the number of trains that need to be added. For the prediction results, T is the passenger flow threshold, C is the maximum number of passengers on a single train, and only when N add If the value is ≥1 and is an integer, the additional train operation will be performed.

[0039] According to another aspect of the present invention, a rail transit scheduling system based on GA-optimized Transformer is provided, the system comprising a data acquisition and preprocessing module, a feature extraction module, a model building module, a model optimization module, and a rail transit scheduling module;

[0040] The data acquisition and preprocessing module is used to collect station passenger flow information and preprocess it to obtain passenger flow sequences;

[0041] The feature extraction module is used to construct multidimensional features including time dimension features, spatial dimension features and external weather features, and then encodes the multidimensional features and fuses them with the passenger flow sequence to obtain a feature fusion vector;

[0042] The model building module is used to build Transformer models based on the multi-head self-attention mechanism;

[0043] The model optimization module is used to input the feature fusion vector into the constructed Transformer model, optimize the model based on the genetic algorithm, and output the optimized Transformer model.

[0044] The rail transit scheduling module is used to input real-time data into the optimized Transformer model after preprocessing and feature extraction, and output passenger flow prediction results. Based on the prediction results, rail transit scheduling is carried out.

[0045] Compared with the prior art, the present invention has the following beneficial effects:

[0046] 1. In this invention, passenger flow information at stations is collected and preprocessed to obtain passenger flow sequences. By constructing and fusing multi-dimensional features, the ability to model complex passenger flow patterns is enhanced. A Transformer model is constructed based on a multi-head self-attention mechanism, which captures long-distance spatiotemporal dependencies, solving the bottlenecks of gradient vanishing and long sequence modeling. The model is optimized using a genetic algorithm, significantly reducing the cost of manual parameter tuning and improving parallel computing efficiency. Finally, scheduling is performed based on real-time data prediction results. This method integrates multi-dimensional information and, based on the Transformer model constructed using a multi-head self-attention mechanism, improves the accuracy and applicability of passenger flow prediction, thereby enhancing the precision and generalization ability of rail transit scheduling. It effectively addresses the shortcomings of existing technologies in long-term sequence modeling, automated hyperparameter optimization, and adaptability to complex rail transit topologies, meeting the dynamic scheduling needs under large-scale network operation.

[0047] 2. The multidimensional features of this invention include: time-dimensional features such as weekday, non-working day, and holiday features; spatial-dimensional features such as K-means clustering algorithm, which classifies stations into commercial center stations, transfer stations, and ordinary stations based on historical passenger flow, volatility, and time-period variation features; and external weather features such as rainfall, wind chill index level, and comprehensive comfort index level. By constructing multidimensional features that include time, space, and external weather, and introducing multidimensional features and K-means clustering technology, the ability to model complex passenger flow patterns under large-scale networked operations is enhanced. After encoding, these features are fused with passenger flow sequences into a feature fusion vector in a specific matrix form. This achieves the effect of fully exploring multiple factors affecting passenger flow, enriching the information in the input model, and improving the accuracy of passenger flow prediction. At the same time, the introduction of multidimensional features makes this method adaptable to the needs of different urban road networks, possessing high practical value and promotion potential.

[0048] 3. In this invention, the range of hyperparameter values ​​meets the constraints of rail transit. By setting the rail transit constraints, such as batch sizes of 16, 32, or 64, the selection of hyperparameters becomes more targeted and reasonable, ensuring that model optimization is carried out within the effective range. This significantly improves the prediction accuracy of rail transit passenger flow and the adaptability of the model in the field of rail transit operation, providing minute-level accurate support for rail transit operation scheduling.

[0049] 4. In this invention, a weighted root mean square error function is constructed based on a time-time differentiated fitness evaluation mechanism, and a fitness function is constructed from the weighted root mean square error function. By constructing specific weighted root mean square error functions and fitness functions, different weights are set for different time periods of different types of stations, so that the model pays more attention to the prediction accuracy of key time periods and stations, and improves the model's ability to adapt to rail transit scheduling scenarios.

[0050] 5. In this invention, rail transit scheduling is carried out based on the prediction results, specifically including: determining the passenger flow threshold of the predicted station, then selecting the predicted peak passenger flow and the full load of a single train from the prediction results, and calculating whether to add trains or keep the original train schedule; by determining the passenger flow threshold based on the prediction results, calculating the number of trains to be added or keeping the original train schedule, and clarifying the conditions for adding trains, train resources are rationally allocated to improve operational efficiency. Attached Figure Description

[0051] Figure 1 This is a schematic diagram of the steps of the rail transit scheduling method based on GA-optimized Transformer in this invention;

[0052] Figure 2 This is a flowchart of the rail transit scheduling process based on GA-optimized Transformer in this invention;

[0053] Figure 3 This is a flowchart of the overall passenger flow prediction model in this invention;

[0054] Figure 4 This is a flowchart of the real-time rail transit dispatching system in the embodiment;

[0055] Figure 5 This is a visualization of passenger flow prediction in the example.

[0056] Figure 6a This is a passenger flow prediction and analysis curve for station one in the example;

[0057] Figure 6b This is a passenger flow prediction and analysis curve for station two in the example;

[0058] Figure 6c This is a graph showing the passenger flow prediction and analysis curves for three stations in the example.

[0059] Figure 7a This is a weekday passenger flow forecast analysis curve for the same station in the example;

[0060] Figure 7b This is a curve showing the predicted passenger flow at the same station during holidays, as shown in the example. Detailed Implementation

[0061] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0062] This proposal suggests a rail transit scheduling method based on GA (Genetic Algorithm) optimization of Transformer (Transformer Neural Network). Although general GA-Transformer-based methods have been applied in other time series forecasting fields, significant technological gaps remain in rail transit passenger flow forecasting scenarios. This proposal addresses these gaps by improving upon the specific characteristics of rail transit.

[0063] 1) By capturing long-distance spatiotemporal dependencies through self-attention mechanisms, the bottlenecks of gradient vanishing and long sequence modeling are solved;

[0064] 2) The use of genetic algorithms to automatically optimize hyperparameters significantly reduces the cost of manual parameter tuning and improves the efficiency of parallel computing;

[0065] 3) Introduce multidimensional features and K-means clustering technology to enhance the ability to model complex passenger flow patterns.

[0066] And the following scenario-based adaptation design will be carried out:

[0067] 1) Apply specialization constraints to the GA hyperparameter search space (such as layer / learning rate range compression);

[0068] 2) Design a fitness assessment mechanism that differentiates by time period to overcome the shortcomings of the traditional RMSE (Root Mean Square Error) which treats all time periods equally.

[0069] It significantly improves prediction accuracy and model adaptability, providing minute-level precise support for operation scheduling; it forms a dedicated optimization system for rail transit, breaking free from the constraints of a general framework; it can adapt to the needs of different urban road networks and has high practical value and promotion potential.

[0070] Example 1

[0071] In this embodiment, a rail transit scheduling method based on GA-optimized Transformer is adopted, and the method steps are as follows: Figure 1 As shown, it specifically includes:

[0072] S1. Collect station passenger flow information and preprocess it to obtain passenger flow sequence;

[0073] S2. Construct a multi-dimensional feature that includes time dimension features, spatial dimension features and external weather features, and then encode the multi-dimensional feature and fuse it with the passenger flow sequence to obtain a feature fusion vector;

[0074] S3. Construct a Transformer model based on a multi-head self-attention mechanism;

[0075] S4. Input the feature fusion vector into the constructed Transformer model, optimize the model based on the genetic algorithm, and output the optimized Transformer model.

[0076] S5. After preprocessing and feature extraction, the real-time data is input into the optimized Transformer model, and the passenger flow prediction result is output. Based on the prediction result, rail transit scheduling is carried out.

[0077] This paper proposes a rail transit scheduling method based on GA-optimized Transformer. It fully utilizes the advantages of Transformer in time series modeling and combines genetic algorithm to automatically optimize model hyperparameters, thereby significantly improving the accuracy and stability of passenger flow prediction. Figure 2 The overall process of this technical solution is demonstrated. The overall solution includes the following five steps: data preprocessing; multi-dimensional feature construction and fusion; Transformer model construction and implementation; genetic algorithm optimization of Transformer hyperparameters; and construction of a real-time rail transit scheduling system.

[0078] The data preprocessing steps mainly include data collection, cleaning, format conversion, and normalization. Data collection utilizes card swipe records from subway stations using the AFC (Automatic Fare Collection) system. Fields include swipe time, station ID, device ID, entry / exit status, and payment type. Passenger flow is aggregated into 15-minute intervals to match operational scheduling cycles.

[0079] Data cleaning steps include:

[0080] 1) Add missing values ​​and remove duplicate values ​​to ensure data integrity and uniqueness;

[0081] 2) Outlier correction: The Z-score method is used to identify and replace abnormal passenger flow values ​​with historical average values. The Z-score threshold is set at ±3.5 (to accommodate large passenger flow fluctuations), specifically expressed as follows:

[0082] Anomaly detection: |X t -μ 7d |>3.5σ 7d ;

[0083] Correction value: X t ′=μ 24h ;

[0084] When the original data X at time t t Compared with the 7-day average μ 7d The absolute deviation is greater than the 7-day standard deviation σ 7d When the value is 3.5 times the normal value, it is considered abnormal data. Abnormal data X t Correction value X t Take the 24-hour average μ 24h .

[0085] The time series format conversion uses a sliding window mechanism to divide the sample sequence, with the window length set to 60 minutes.

[0086] The normalization process uses the Min-Max normalization method, whose expression is as follows:

[0087]

[0088] Among them, X min and X max These are the minimum and maximum values ​​of historical data for a certain site, respectively. X is the original data value of the site, and X′ is the standard normalized value of the site.

[0089] Among them, multidimensional features include time dimension features, spatial features, and external weather:

[0090] 1) Extract time-dimensional features for weekdays, non-working days, and holidays to capture passenger flow fluctuation characteristics;

[0091] 2) Spatial Dimension Features: Based on historical passenger flow, volatility, and time-of-day variation characteristics, the K-means clustering algorithm is used to classify stations into commercial center stations, transfer stations, and ordinary stations. The specific method is as follows:

[0092] Clustering input features include three dimensions:

[0093]

[0094] Clustering outputs three labels: 0 = ordinary station (e.g., residential area), 1 = transfer station (≥3 lines intersect), 2 = commercial center station (evening peak period > 35%).

[0095] 3) External weather characteristics include rainfall, wind chill index level, and overall comfort index level.

[0096] All discrete features are encoded using one-hot encoding, concatenated with the normalized passenger flow sequence, and finally a multi-dimensional vector is generated as the input to the subsequent Transformer model.

[0097] The Transformer model consists of an encoder and a decoder. The encoder is mainly used for input sequence modeling, and the decoder outputs the future passenger flow prediction. The input is a multi-dimensional feature sequence of historical passenger flow generated by S1 preprocessing and S2 feature fusion, and the output is a predicted passenger flow sequence for the future target time period (e.g., 1-4 hours in the future, with a 15-minute granularity).

[0098] 1. Input Layer Design

[0099] The input layer receives the feature fusion output from S2, forming a matrix X∈R. T×d :

[0100]

[0101] Among them: time steps T = 4 (60-minute window, 4 15-minute time intervals), feature dimension d = 32 (including normalized entry / exit volume, time stamp, station type code, rainfall level, wind chill level, comfort index and other features);

[0102] 2. Embedded layer processing

[0103] The original feature matrix X is input and converted into a high-dimensional vector:

[0104] The original features are mapped to a high-dimensional space through a linear transformation E = Embedding(X), and the output E ∈ R is obtained. 4×64 This solves the problem of linear inseparability of original features. While meeting the feature dimension requirements of the Transformer's attention mechanism, the 64-dimensional space can also carry richer periodic basis functions (24-hour periodic components), meeting the needs of long-term passenger flow modeling.

[0105] 3. Add location encoding

[0106] Add location coding to provide time and location information, and change the standard Transformer's 10,000-cycle base to 24 (to adapt to the 24-hour daily cycle of passenger flow). This adjustment makes the location coding phase precisely aligned with the morning and evening peak hours (e.g., t=7 corresponds to the start of the morning peak), as detailed below:

[0107]

[0108] Where t represents the time step index, i represents the dimension index, and PE is the position encoding matrix. The final input encoding is represented as:

[0109] H0=E+PE∈R 4×64

[0110] 4. Spatiotemporal modeling of the encoder

[0111] 1) Multi-head attention layer (MHSA)

[0112] First, add positional encoding to the matrix H0∈R 4×64 As input, generate the query matrix, key matrix, and value matrices Q, K, V:

[0113] Q,K,V=H0·W Q,K,V ;

[0114] Where the weight matrix W Q,K,V ∈R 64×8 The values ​​were obtained from Gaussian and Xavier distributions, respectively, and W... V The initial weights are increased by 20% to enhance weather-specific response capabilities and improve adaptability for rail transit.

[0115] W Q ∈R 64×8 ~N(0,0.01) 2 );

[0116] W K ∈R 64×8 =XavierInit(64,8);

[0117] W V ∈R 64×8 =XavierInit(64,8)×(1+0.2);

[0118] Secondly, the Multi-Head Self-Attention Module (MHSA) is used, and its expression is as follows:

[0119]

[0120] In the formula, Q, K, and V are the query matrix, key matrix, and value matrix, respectively. These are scaling factors for K and Q, fixed at 1. Scaling is to avoid QK T An excessively large value can cause the Softmax gradient to explode.

[0121] The matrix A∈R of the output time-dependent relationships is calculated using the Softmax activation function formula. 4×64 :

[0122]

[0123] 2) Feedforward Neural Network Layer (FFN)

[0124] The attention output matrix A∈R 4×64 As input, the formula for the feedforward neural network is:

[0125] Z=FFN(H)=ReLU(HW1+b1)W2+b2∈R 4×64 ;

[0126] In the formula, ReLU is the activation function, b1 and b2 are the hidden layer bias vector and the output layer bias vector, respectively, W1 and W2 are the weight matrices from the input layer to the hidden layer and from the hidden layer to the output layer, respectively; the encoder output matrix Z, whose element values ​​represent the intensity of the fused spatiotemporal features, positive values ​​indicate the trend of passenger flow growth, and negative values ​​indicate the trend of passenger flow dissipation;

[0127] 3) Residual connectivity and layer normalization

[0128] The encoder has a residual connection part, that is, the output of each layer is the result of the attention processing of this layer plus the input; the layer normalization step is to standardize the 64-dimensional features of a single time period, in order to stabilize the values.

[0129] 5. Decoder Prediction Output

[0130] Additionally, the decoder, similar in structure to the encoder, is used to output future passenger flow predictions.

[0131] 1) Masked Self-Attention Layer

[0132] The input for this layer is the historical passenger flow Y. hist In the self-attention mechanism, the mask matrix M ensures that the model can only see past passenger flow data. The output temporal constraint feature matrix M is calculated through the Transformer model's unique attention mechanism. out ,in:

[0133] Y hist = [Inbound / Outbound Volume (t-45min), ... Inbound / Outbound Volume t];

[0134]

[0135] 2) Encoder-decoder attention layer

[0136] This layer is the core of spatiotemporal feature fusion, and the fused feature matrix C∈R is calculated through the activation function Softmax. 4 ×64 :

[0137]

[0138] Among them W Q W K W V The weight matrix consists of parameters that the model learns automatically during training;

[0139] 3) Predicting the output layer

[0140] After multiple training and normalization processes, the final predicted passenger flow is output.

[0141] The initial model parameters are set as follows:

[0142] 1) Batch size is 32;

[0143] 2) The encoder-decoder layer count is 5;

[0144] 3) The number of training rounds is 70;

[0145] 4) The learning rate is 0.0004;

[0146] 5)W V The initial gain is 1.4.

[0147] Genetic algorithms (GA) are used to automatically search for and optimize the key hyperparameters of a Transformer model; the optimized parameters are then used to train the Transformer model. Specifically, combined with... Figure 3 As shown, the optimization process is as follows:

[0148] 1. Initialize the population

[0149] Multiple different hyperparameter combinations are generated. Each "chromosome" represents a set of hyperparameter combinations, and each "chromosome" contains 5 "genes," corresponding to the key parameters of the S3 model: [batch size, learning rate, number of encoder layers, number of decoder layers, W]. V [Matrix initialization gain];

[0150] Scenario-based adaptation design means imposing professional constraints on the GA hyperparameter search space, with the parameter range strictly following the constraints in the table below.

[0151] Table 1 Rail Transit Constraints

[0152]

[0153] 2. Fitness Calculation

[0154] The fitness function of a genetic algorithm is used to evaluate the fitness of an individual in the problem space and determine its probability of being selected during the evolution process. This paper uses the weighted root mean square error (RMSE) as the evaluation metric and sets up a time-differentiated fitness evaluation mechanism based on the rail transit scenario, namely:

[0155]

[0156] Among them, y i This represents the actual passenger flow value. For predicted values, from Inverse normalization to passenger flow, 4 represents the number of chromosomes, W 时段 (i) represents the time period weighting coefficient. For transfer stations, the weights for morning peak, evening peak, and ordinary time periods are 1.8, 1.5, and 1.0, respectively, while for commercial stations, they are 1.3, 1.7, and 1.0, respectively.

[0157] 3. Genetic manipulation

[0158] 1) Select

[0159] The selection mechanism uses a roulette wheel algorithm, where higher fitness individuals have a greater probability of being selected, and individuals with fitness below the average are eliminated, with the following probabilities:

[0160]

[0161] 2) Cross

[0162] The selected chromosomes are subjected to crossover operation at a preset crossover probability (0.75), and some hyperparameter values ​​are randomly exchanged to generate new offspring. For example, parent A [32,0.00015,4,1.5] and parent B [64,0.0003,5,1] are crossed to generate offspring [32,0.0003,4,1].

[0163] 3) Variation

[0164] A subset of genes is randomly adjusted with a low mutation probability (0.1) to maintain population diversity and prevent local optima. Specifically, the mutation range of the learning rate is set to ±0.0001, the number of layers is ±1, and W... V The gain is ±0.1;

[0165] 4. Termination and Output

[0166] When the fitness improvement is less than 1% for 10 consecutive generations or the maximum number of iterations is reached, the chromosome with the highest fitness is output as the final configuration, i.e., batch size is 32, learning rate is 0.00013, encoder-decoder layers are 4, and W... V The initial gain is 1.2.

[0167] Among them, the construction of the real-time rail transit dispatching system and the integration of system processes Figure 4 As shown, a visualization and prediction system for operation scheduling was designed, consisting of the following modules:

[0168] 1) Data acquisition module: Connects to the AFC (Automatic Fare Collection) system interface to collect card swipe records in real time;

[0169] 2) Data preprocessing module: Implements missing value handling, anomaly detection, and sliding window partitioning;

[0170] 3) Feature extraction module: Dynamically extracts input features based on current time, site category, and historical traffic;

[0171] 4) Prediction module: Inputs features into the GA-optimized Transformer model for real-time prediction;

[0172] 5) Results Display Module: Combined with Figure 5 As shown, the site-level prediction results are displayed in chart form, and scheduling suggestions are provided, specifically:

[0173] The passenger flow threshold T for each station is determined based on practical factors such as its size, area, and type. Then, the peak passenger flow forecast is selected. Let C be the maximum passenger capacity of a single train (taking a Type A subway train in a certain city as an example, with a maximum capacity of 930 people). The formula for calculating whether to add more trains or maintain the original train number is as follows:

[0174]

[0175] In the formula, N add This represents the number of trains that need to be added, and only if N... add It will only be executed when the value is ≥1.

[0176] This system can be deployed in the command system of rail transit operators to achieve minute-level updates and hourly-level predictive accuracy control, significantly enhancing scheduling efficiency and service response capabilities.

[0177] The experiment was conducted based on partial AFC (Average Facing) data from a subway system in a certain city in a certain year. The passenger flow sampling period was 15 minutes. As shown in Figures 6 and 7, ... Figure 6a A passenger flow forecast analysis curve for a single station; Figure 6b The passenger flow forecast analysis curve for Station 2; Figure 6c A graph showing the predicted passenger flow at the station. Figure 7a A weekday passenger flow forecast analysis curve for the same station; Figure 7bThe graph shows the passenger flow forecast analysis curves for the same station during holidays. In the graph, the red line represents the actual value, and the blue line represents the predicted value. The method is compared with ARIMA (Autoregressive Integrated Moving Average), LSTM (Long Short-Term Memory), and GRU (Gated Recurrent Unit) using actual test data. Evaluation is performed using MAE (Mean Absolute Error) and RMSE (Root Mean Square Error).

[0178] Under different station types, the results show that at a certain transfer station, the traditional ARIMA model has a mean absolute error (MAE) of 13.121 and a root mean square error (RMSE) of 21.215, while the GA-Transformer model in this proposal reduces the MAE to approximately 10.215 and the RMSE to approximately 17.966, respectively, achieving a reduction of approximately 22% in MAE and approximately 15% in RMSE. Similar improvements have been verified at other station types, fully demonstrating the significant advantages of this method in capturing complex passenger flow fluctuations and handling nonlinear and time-varying characteristics.

[0179] On weekdays and holidays at the same station, the results show that the GA-Transformer model can accurately capture passenger flow fluctuations during morning and evening peak hours and off-peak periods, and maintains high prediction accuracy even during off-peak periods on holidays. These results further validate the model's applicability in diverse scenarios.

[0180] This method introduces a genetic algorithm to achieve adaptive optimization of hyperparameters and combines it with multi-dimensional data feature modeling, enabling the Transformer model to still have high accuracy and high stability in complex time-varying scenarios of rail transit, and has good prospects for practical application.

[0181] Example 2

[0182] In this embodiment, a rail transit scheduling system based on GA-optimized Transformer is adopted. The system includes a data acquisition and preprocessing module, a feature extraction module, a model building module, a model optimization module, and a rail transit scheduling module.

[0183] The data acquisition and preprocessing module is used to collect station passenger flow information and preprocess it to obtain passenger flow sequences;

[0184] The feature extraction module is used to construct multidimensional features including time dimension features, spatial dimension features and external weather features, and then encodes the multidimensional features and fuses them with the passenger flow sequence to obtain a feature fusion vector;

[0185] The model building module is used to build Transformer models based on the multi-head self-attention mechanism;

[0186] The model optimization module is used to input the feature fusion vector into the constructed Transformer model, optimize the model based on the genetic algorithm, and output the optimized Transformer model.

[0187] The rail transit scheduling module is used to input real-time data into the optimized Transformer model after preprocessing and feature extraction, and output passenger flow prediction results. Based on the prediction results, rail transit scheduling is carried out.

[0188] The specific implementation process of this system includes:

[0189] Data preprocessing; construction and fusion of multidimensional features; construction and implementation of Transformer model; optimization of Transformer hyperparameters using genetic algorithm; construction of a real-time rail transit passenger flow prediction system.

[0190] Preprocessing includes data acquisition, cleaning, format conversion, and normalization. Data acquisition uses AFC data from rail transit, including: card swipe time, station number, card reader number, entry / exit status, and card swipe type; data cleaning includes filling in missing values, removing duplicate values, and correcting outliers; time series format conversion uses a sliding window mechanism to divide the sample sequence, with the window length set to 60 minutes.

[0191] The normalization process uses the Min-Max normalization method, as shown in the following formula:

[0192]

[0193] Among them, X min and X max These are the minimum and maximum values ​​of historical data for a certain site, respectively. X is the original data value of the site, and X′ is the standard normalized value of the site.

[0194] Multidimensional features include time-dimensional features, spatial features, and external weather.

[0195] 1) Extract time-dimensional features for weekdays, non-working days, and holidays to capture passenger flow fluctuation characteristics;

[0196] 2) Spatial dimension features: Based on historical passenger flow, volatility and time-period variation characteristics, the K-means clustering algorithm is used to classify the stations into commercial center stations, transfer stations and ordinary stations;

[0197] 3) External weather characteristics include rainfall, wind chill index level, and overall comfort index level;

[0198] The Transformer model is constructed and implemented using an encoder and decoder structure. The encoder is used for input sequence modeling, and the decoder outputs future passenger flow predictions. Its input comes from the feature fusion output of S2, forming a matrix X∈R. T×d :

[0199]

[0200] Among them: time steps T = 4 (60-minute window, 4 15-minute time intervals), feature dimension d = 32 (including normalized entry / exit volume, time stamp, station type code, rainfall level, wind chill level, comfort index and other features);

[0201] After processing by the embedding layer, adding positional encoding, and performing spatiotemporal modeling, the encoder layer generates a query matrix, a key matrix, and value matrices Q, K, V. Finally, the matrix A∈R representing the time-dependent relationships is calculated using the Softmax activation function formula. 4×64 The formula is as follows:

[0202]

[0203] In the formula, Q, K, and V are the query matrix, key matrix, and value matrix, respectively. These are scaling factors for K and Q, fixed at 1. Scaling is to avoid QK T An excessively large value can cause the Softmax gradient to explode.

[0204] The input to the decoder layer is the historical passenger flow Y. hist = [Inbound / Outbound Volume (t-45min), ... Inbound / Outbound Volume t], and the output temporal constraint feature matrix M is calculated using the attention mechanism unique to the Transformer model. out The fused feature matrix C∈R is calculated using the activation function Softmax. 4×64 After multiple training and normalization processes, the matrix outputs the final predicted passenger flow. The calculation expression is as follows:

[0205]

[0206] Among them W Q W K W VThe weight matrix is ​​a set of parameters that the model learns automatically during training.

[0207] Genetic algorithm optimization of hyperparameters involves using a genetic algorithm (GA) to optimize key hyperparameters of a Transformer model: batch size, learning rate, number of encoder layers, number of decoder layers, and W. V Matrix initialization gain;

[0208] Specialization constraints are imposed on the search space of hyperparameters to define the parameter range. The genetic algorithm first initializes the population and generates multiple different "chromosomes". Each "chromosome" represents a set of hyperparameter combinations, containing 5 "genes" that correspond to the key parameters of the Transformer model to be optimized.

[0209] The fitness function of a genetic algorithm is used to evaluate the fitness of an individual in the problem space and determine its probability of being selected during the evolutionary process. This paper uses the weighted root mean square error (RMSE) as the evaluation metric and sets up a time-differentiated fitness evaluation mechanism based on the rail transit scenario, namely:

[0210]

[0211] Among them, y i This represents the actual passenger flow value. For predicted values, from Inverse normalization to passenger flow, 4 represents the number of chromosomes, W 时段 (i) represents the time period weighting coefficient. For transfer stations, the weights for morning peak, evening peak, and ordinary time periods are 1.8, 1.5, and 1.0, respectively, while for commercial stations, they are 1.3, 1.7, and 1.0, respectively.

[0212] The chromosome with the highest fitness is selected, crossovered, and mutated as the final configuration. That is, the batch size is 32, the learning rate is 0.00013, the encoder-decoder layers are 4, and W... V The initial gain is 1.2.

[0213] This system integrates with the AFC (Automatic Fare Collection) system to collect data in real time. After preprocessing and feature extraction, the data is input into an optimized Transformer model to predict future passenger flow. The results are displayed in visual formats such as line charts, and operational scheduling suggestions are provided, such as adjusting train frequency or service capacity.

[0214] The system also includes a visualization module, which enables minute-level updates and hour-level prediction accuracy control.

[0215] The scheduling information is displayed through a graphical interface, including future passenger flow forecasts, predictions for peak hours, off-peak hours, and abnormal situations. It also features a decision support mechanism that provides real-time operational scheduling suggestions based on the forecast results.

[0216] In summary, this system is designed with modules for data acquisition and preprocessing, each performing its corresponding function. This system makes the rail transit dispatching process systematic and modular, ensuring the orderly connection of each link and improving dispatching efficiency and accuracy.

[0217] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A rail transit scheduling method based on GA-optimized Transformer, characterized in that, The method steps include: S1. Collect station passenger flow information and preprocess it to obtain passenger flow sequence; S2. Construct a multi-dimensional feature that includes time dimension features, spatial dimension features and external weather features, and then encode the multi-dimensional feature and fuse it with the passenger flow sequence to obtain a feature fusion vector; S3. Construct a Transformer model based on a multi-head self-attention mechanism; S4. Input the feature fusion vector into the constructed Transformer model, optimize the model based on the genetic algorithm, and output the optimized Transformer model. S5. After preprocessing and feature extraction, the real-time data is input into the optimized Transformer model, and the passenger flow prediction result is output. Based on the prediction result, rail transit scheduling is carried out.

2. The rail transit scheduling method based on GA-optimized Transformer according to claim 1, characterized in that, The station passenger flow information in S1, i.e. rail transit AFC data, includes at least the card swiping time, station number, card swiping device number, entry / exit status, and card swiping type.

3. The rail transit scheduling method based on GA-optimized Transformer according to claim 2, characterized in that, The data preprocessing in S1 includes data cleaning, time format conversion, and normalization. Data cleaning includes filling in missing values, removing duplicates, and correcting outliers. The time series format conversion uses a sliding window mechanism to divide the sample sequence, with a window length set to 60 minutes. The normalization process uses the Min-Max normalization method, as shown in the following formula: Among them, X min and X max Let X be the minimum and maximum historical data values ​​of a certain site, and let X be the original data value of that site. ′ This is the standard normalized value for this site.

4. The rail transit scheduling method based on GA-optimized Transformer according to claim 1, characterized in that, Among the multidimensional features in S2: Time-related features include weekday, non-working day, and holiday features; The spatial dimension features are defined by using the K-means clustering algorithm to classify stations into commercial center stations, transfer stations, and ordinary stations based on historical passenger flow, volatility, and time-of-day changes. External weather characteristics include rainfall, wind chill index, and overall comfort index. The feature fusion vector in S2 is in matrix form: The time step T = 4, which is a 60-minute window containing four 15-minute time intervals; the feature dimension d = 32.

5. A rail transit scheduling method based on GA-optimized Transformer according to claim 1, characterized in that, The Transformer model in S3 includes an encoder and a decoder; the encoder is used to input sequences and model them, and the decoder is used to complete future passenger flow prediction and output them; the input of the decoder layer is historical passenger flow, which passes through the attention mechanism in the Transformer to calculate and output a temporal constraint feature matrix. This temporal constraint feature matrix is ​​used to calculate and output a fusion feature matrix through an activation function. After multiple training and normalization, the fusion feature matrix outputs the final predicted passenger flow.

6. The rail transit scheduling method based on GA-optimized Transformer according to claim 1, characterized in that, In S4, when optimizing the model based on a genetic algorithm, the specific steps include: S41. Initialize the population and generate multiple different hyperparameter combinations. Each chromosome represents a set of hyperparameter combinations, and each chromosome contains 5 genes. These 5 genes correspond to the hyperparameters of the model: batch size, learning rate, number of encoder layers, number of decoder layers, and matrix initialization gain. S42. Based on the fitness evaluation mechanism with time-time differences, construct a weighted root mean square error function, and construct the fitness function from the weighted root mean square error function; S43. Selection, crossover, and mutation processes based on fitness functions; S44. When the fitness improvement effect is less than 1% for 10 consecutive generations or the maximum number of iterations is reached, output the chromosome with the highest fitness as the final configuration, that is, output the current batch size, learning rate, number of encoder and decoder layers and initial gain.

7. A rail transit scheduling method based on GA-optimized Transformer according to claim 6, characterized in that, In S41, the range of hyperparameter values ​​satisfies the constraints of rail transit. Specifically, the batch size is 16, 32, or 64, the learning rate is 0.0001-0.005, the number of encoder or decoder layers is 3-5, and the initialization gain is 1.0-1.

5.

8. A rail transit scheduling method based on GA-optimized Transformer according to claim 6, characterized in that, In step S42, the specific formulas for the constructed weighted root mean square error function and fitness function are as follows: Where RMSE is the weighted root mean square error function, y i This represents the actual passenger flow value. For predicted values, from Inverse normalization is used to represent passenger flow, 4 represents the number of chromosomes, Fitness is the fitness function, and W 时段 (i) represents the time period weighting coefficient. For transfer stations, the weights for morning peak, evening peak, and ordinary time periods are 1.8, 1.5, and 1.0, respectively, while for commercial stations, they are 1.3, 1.7, and 1.0, respectively.

9. A rail transit scheduling method based on GA-optimized Transformer according to claim 1, characterized in that, The rail transit scheduling in S5 based on the prediction results specifically includes: Determine the passenger flow threshold for the predicted station, then select the peak passenger flow forecast and the maximum passenger capacity per train from the forecast results, and calculate whether to add more trains or keep the original number of trains. The specific formula is as follows: In the formula, N add This represents the number of trains that need to be added. For the prediction results, T is the passenger flow threshold, C is the maximum number of passengers on a single train, and only when N add If the value is ≥1 and is an integer, the additional train operation will be performed.

10. A rail transit scheduling system based on GA-optimized Transformer, characterized in that, The system operates using a rail transit scheduling method based on GA-optimized Transformer as described in any one of claims 1-9. The system includes a data acquisition and preprocessing module, a feature extraction module, a model building module, a model optimization module, and a rail transit scheduling module. The data acquisition and preprocessing module is used to collect station passenger flow information and preprocess it to obtain passenger flow sequences. The feature extraction module is used to construct multi-dimensional features including time dimension features, spatial dimension features and external weather features, and then encodes the multi-dimensional features and fuses them with the passenger flow sequence to obtain a feature fusion vector. The model building module is used to build a Transformer model based on a multi-head self-attention mechanism; The model optimization module is used to input the feature fusion vector into the constructed Transformer model, optimize the model based on the genetic algorithm, and output the optimized Transformer model. The rail transit scheduling module is used to input the preprocessed and feature-extracted real-time data into the optimized Transformer model, output passenger flow prediction results, and perform rail transit scheduling based on the prediction results.

Citation Information

Cited By

  • Urban rail transit abnormal large passenger flow real-time prediction method

    CN121638697A

  • Subway passenger flow prediction method and system based on optimized plant growth algorithm and weighted convolution TimeXer

    CN122288049A