Urban rail passenger flow and train collaborative organization optimization method based on deep reinforcement learning
Through deep reinforcement learning technology, combined with diffusion spatio-temporal graph model and DQN, the coordinated scheduling of passenger flow and trains in urban rail transit systems has been solved, and traditional methods are difficult to cope with complex passenger flow changes and energy efficiency balance, achieving more efficient operations and lower energy consumption.
Patent Information
- Application Number
- CN202510355798.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-25
AI Technical Summary
Traditional urban rail transit scheduling methods are difficult to effectively optimize complex passenger flow changes, train operation scheduling and energy efficiency, and are not adaptable to real-time data.
The optimization method of urban rail passenger flow and train collaborative organization of urban rail passenger flow based on deep reinforcement learning is adopted to predict passenger flow distribution through diffusion spatio-temporal graph model, and DQN is used to optimize train operation and passenger flow scheduling, combining real-time data dynamic adjustment strategies.
It improves the operating efficiency of urban rail transit systems, reduces energy consumption, and can better respond to changes in passenger waiting time and congestion, and improves the overall service level.
Smart Images

Figure CN120218544A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of urban rail transit operation dispatching, and specifically relates to an optimization method for coordinated organization of urban rail passenger flow and trains based on deep reinforcement learning. Background Art
[0002] With the acceleration of the urbanization process, the passenger flow of urban rail transit systems has been increasing year by year. How to efficiently optimize the passenger flow and train operation plan has become a key issue in improving the service quality of rail transit, enhancing the operation efficiency, and reducing energy consumption.
[0003] Traditional methods for optimizing rail transit operation plans mainly rely on mathematical modeling, heuristic algorithms, and empirical rules. However, with the increasing complexity of passenger flow distribution, passenger demand, and train operation, traditional methods face the following challenges: (1) complex passenger flow changes; (2) difficulty in optimizing train operation dispatching; (3) balance between energy consumption and efficiency; (4) insufficient adaptability to real-time data.
[0004] Currently, most existing studies focus on single - aspect optimization, such as train dispatching, station passenger flow management, or energy consumption optimization. Optimization methods that comprehensively consider multiple factors such as passenger flow, train operation, and energy efficiency are relatively scarce. With the rapid development of deep learning technology, methods based on deep reinforcement learning have gradually become an effective way to solve these problems. Reinforcement learning can automatically learn the optimal strategy from historical data through interaction with the environment, while deep neural networks can process a large amount of non - linear complex data to improve the performance of the model.
[0005] The differences between this application and the prior art are as follows:
[0006] Technical comparison with the patent CN115352502A "A method, device, electronic device, and storage medium for adjusting a train operation plan"
[0007] 1. Patent CN115352502A uses an adversarial neural network. The generator receives the normalized time input and outputs the passenger flow OD matrix. The discriminator judges the authenticity of the input OD matrix and also determines the time information, so that the generated OD matrix is as close as possible to the real matrix in statistical characteristics. In this study, a diffusion spatio-temporal graph model is adopted. The graph structure is used to describe the connection attributes between stations, and the adjacency matrix is used to describe the topological structure between stations. Taking the historical OD matrix as the input and combining the station network structure, the passenger flow distribution under multiple future time granularities is predicted. In the original scheme, the generator and the discriminator mainly rely on the input time data and the generated OD matrix for adversarial training. The spatial information is hidden in the OD matrix, lacking explicit modeling of the station network. In the new scheme, the diffusion spatio-temporal graph model can capture the rules of passenger flow between stations, making the capture of passenger flow diffusion and spatio-temporal dynamic characteristics more accurate, thus helping to improve the overall prediction accuracy.
[0008] 2. During the construction of the urban rail transit network simulation in Patent CN115352502A, the network structure between stations is not explicitly modeled. It mainly relies on the preset passenger travel paths to determine the boarding, alighting and transfer behaviors of passengers, and cannot truly reflect the asymmetry and individual differences in the process of passengers entering, transferring and leaving the station. In this study, after generating the initial passenger information using the OD matrix, the travel paths of passengers are assigned through the Logit model, effectively considering the sensitivity of passengers to travel costs, familiarity and induced information, and taking into account the randomness and decision-making rules of passenger behavior. Therefore, the urban rail transit network simulation model designed in this study can reflect the dynamic changes of train operation, station congestion and passenger behavior in real time, and is more suitable for the research on the coordinated optimization of passenger flow and trains in complex urban rail transit networks, laying a foundation for the subsequent coordinated organization and optimization of passenger flow and trains.
[0009] III. Patent CN115352502A focuses on train timetable adjustment. It constructs a deep reinforcement learning model based on the Advantage Actor-Critic (A2C) method, which consists of a policy network and a value network. Its reward function is calculated by the waiting times of inbound passengers and passenger transfers. This reward function can intuitively reflect the comprehensive performance of passenger waiting times under a specific timetable, thus guiding the policy network to adjust the train departure time to improve the passenger experience. In this study, the passenger flow scheduling problem and the train operation optimization problem are modeled as a Markov decision process in reinforcement learning, and Deep Q-Network (DQN) is used to find the optimal strategy. Mechanisms such as an experience pool and a target network are introduced, greatly improving the stability of training and reducing the risk brought by frequent changes in the target value. The reward function is based on the train operation cost and the passenger waiting time, combining the operation cost with the passenger experience, enabling the training process to dynamically respond to changes in passenger waiting time and crowding while taking into account economic benefits, thereby achieving the collaborative optimization of train timings and passenger flow scheduling. The original solution is applicable to continuous control problems based on policy gradients, with a simple structure, facilitating the preliminary verification of the impact of train time adjustment on passenger waiting time; while the new solution uses the DQN method based on value function approximation, which is more meticulous in state representation, action design, and reward construction, and can take into account both operation cost and passenger experience, providing a more comprehensive solution for the collaborative optimization of complex urban rail transit networks.
[0010] Therefore, combining deep reinforcement learning technology, a method that can comprehensively optimize passenger flow and train operation plans is proposed, which can not only improve the operation efficiency of the system but also reduce energy consumption, having important application prospects. Summary of the Invention
[0011] In view of the above problems, the present invention proposes an urban rail transit passenger flow and train collaborative organization optimization method based on deep reinforcement learning, which can not only improve the operation efficiency of the system but also reduce energy consumption, having important application prospects.
[0012] To achieve the above object, the technical solution adopted by the present invention is:
[0013] The urban rail transit passenger flow and train collaborative organization optimization method based on deep reinforcement learning includes the following steps:
[0014] Step 1: Utilize the AFC data of rail transit to analyze the arrival patterns and travel chains of passengers, and construct a passenger flow OD matrix;
[0015] Step 2: Use the passenger flow OD matrix constructed in Step 1 to train a diffusion model to generate passenger flow OD matrices for different time periods, providing rich data for train operation and passenger flow scheduling optimization;
[0016] Step 3: Using the OD matrix generated by the diffusion model in Step 2 as the input passenger flow, design the interaction events of passengers, trains, and stations based on the operation characteristics of the urban rail transit network system, simulate the state changes of the rail transit system, and construct a discrete-event-based urban rail transit network simulation model;
[0017] Step 4: On the basis of the urban rail transit network simulation model, study the collaborative scheduling optimization method for urban rail passenger flow and trains based on DQN. With the goal of minimizing the train traction energy consumption and the passenger travel time, construct the system state, action, and reward functions of DQN under the constraints of urban rail operation rules;
[0018] Step 5: Design a training process for the collaborative scheduling optimization method for urban rail passenger flow and trains based on DQN. Through training, obtain the optimized scheduling strategy and apply it to the train operation and passenger flow scheduling of the urban rail transit system, thereby improving the passenger travel experience and the service level of urban rail transit.
[0019] As a further improvement of the present invention, Step 1 is specifically as follows:
[0020] Using the processed rail transit AFC data, obtain the passenger entry and exit times, as well as the departure station and destination station information. By analyzing the passenger arrival patterns and travel chains, generate a passenger flow OD matrix with a time granularity Δt of ten minutes or half an hour, providing data support for passenger flow generation. The OD matrix at time t is denoted as X t ∈R N×N , where N is the number of stations. Combine all the OD matrices under Δt into a three-dimensional time-series matrix X h =[X Δt ,..., X hΔt , where j represents the number of historical time granularities, and the dimension is h×N×N.
[0021] As a further improvement of the present invention, Step 2 is specifically as follows:
[0022] Use the diffusion spatio-temporal graph model to predict the passenger flow OD matrix. Represent the graph structure with G=(V, E, A), where V is the set of stations, corresponding to the starting station i and the terminal station h of the OD matrix; E represents the connection attributes between stations; A is the adjacency matrix, used to describe the topological structure between stations, and X t [i, j] represents the passenger flow from station i to station j at time t;
[0023] The goal of passenger flow prediction is to input the historical OD matrix X H =[X Δt ,..., X HΔt , H≤h and the graph structure G, and predict the future OD matrix X F =[X H+Δt ,..., XH+FΔt , where F represents the number of prediction time granularities:
[0024] (X H ; G) → [X H+Δt , …, X H+FΔt : = X F .
[0025] As a further improvement of the present invention, step 3 is specifically as follows:
[0026] 3.1 In the discrete event-based urban rail transit network simulation model, there are three main entities, namely rail transit trains, rail transit stations, and passengers traveling by rail transit;
[0027] The train is the carrier for carrying passengers and needs to run on the track according to the requirements of the timetable. The train attributes include number, running direction, line number, originating station, stopping stations, and terminal station;
[0028] The station is the main place for passenger gathering and dispersal. The station attributes include station name, affiliated line, whether it is a transfer station, the numbers of the trains stopping at it, and the number of passengers on the current platform;
[0029] The passenger is the service object in the rail transit system. The passenger attributes include number, departure station, destination station, travel path, and entry time;
[0030] 3.2 In the discrete event-based urban rail transit network passenger flow simulation model, the train needs to trigger a series of events to update its state and interact with stations and passengers. The events include train initialization, train operation, and stopping;
[0031] 3.3 In the simulation model of train operation and passenger flow state, the events of passengers include initialization, entry, exit, waiting, transfer, and boarding / alighting.
[0032] As a further improvement of the present invention, in the simulation model of 3.3 train operation and passenger flow state, the events of passengers are specifically described as follows:
[0033] Passenger initialization: At the beginning of the simulation, according to the OD matrix generated in step 2, generate the initial information of passengers, including number, departure station, destination station, travel path, and entry time. Let λ = X nΔt [i, j] represent the passenger flow from station i to station j at the nth time granularity; the entry time of each passenger p obeys an exponential distribution:
[0034]
[0035] Denote the boarding time of passenger p from station i to station j during the time period from nΔt to (n + 1)Δt;
[0036] The boarding moment of passenger p is:
[0037]
[0038] According to the boarding moment of each passenger, relevant times during the passenger's journey are generated using the lognormal distribution The lognormal distribution is:
[0039]
[0040] In the formula, μ is the logarithmic mean of the generated time, σ is the standard deviation of the generated time, and represents the probability density of the lognormal distribution at Set the walking boarding time of passengers to follow the lognormal distribution The transfer time of passengers follows the lognormal distribution The walking alighting time of passengers follows the lognormal distribution
[0041] After generating the passenger boarding time, use the Logit model to allocate the passenger travel path:
[0042]
[0043] where represents the probability that the passenger selects path k within the ij path, K ij represents all the optional paths within the ij path in the OD matrix, represents the travel cost of path k in the OD matrix, θ represents the passenger's familiarity with the urban rail transit network, d represents the passenger's acceptance of the guiding information, and β represents the influence of the guiding information on path selection;
[0044] Passenger boarding, alighting, waiting, transferring, getting on and off: Passengers arrive at the station according to the generated boarding time, and the train arrives at the corresponding station within the specified time according to the set timetable, which triggers the getting on and off events of passengers. When passengers get off, it needs to be judged in combination with the path selection. If the getting off condition is not met, they will continue to stay on the train. If the condition is met, the transfer event or alighting event of passengers will be triggered. When passengers get on, it is restricted by the train direction, the current passenger capacity of the train and the passenger path. If the getting on condition is not met, they will continue to wait.
[0045] As a further improvement of the present invention, step 4 is specifically:
[0046] Model the passenger flow scheduling problem and the train operation optimization problem as Markov decision problems in reinforcement learning, and use DQN to find the optimal strategy. The learning process includes the following:
[0047] System state: The state includes platform congestion, section congestion, and train position;
[0048] Action: The train departure interval Δt train , the running time in the interval The station stop time The section recommendation a m , a m is a 0-1 variable, which is 1 when the m section is recommended, otherwise 0, m ∈ N;
[0049] Reward: Designed based on the optimization objective, and the specific description is as follows:
[0050] Train operation cost:
[0051] C op =C f D t
[0052] C op represents the train operation cost, C f represents the fuel cost per unit distance, D t represents the train running distance;
[0053] The waiting time of passenger p:
[0054]
[0055] In the formula represents the waiting time of passenger p, represents the arrival time of train x at platform y, which is obtained from the urban rail transit network simulation model;
[0056] To synthesize these objectives, they are weighted and combined into a total reward function R(t), and the goal is to maximize this function:
[0057]
[0058] α1 is the weight of the train operation cost, α2 is the weight of the passenger waiting time, η is the overall scaling coefficient, and the values of α1, α2, and η can be adjusted according to the emphasis of different actual requirements to balance the influence between different objectives.
[0059] As a further improvement of the present invention, step 5 is specifically as follows:
[0060] 5.1 The discrete event-based urban rail transit network simulation model serves as the simulation environment for deep reinforcement learning. By generating different passenger flow and train operation scenarios, it simulates the situation of passengers between platforms and carriages as well as the train operation. In this simulation environment, the DQN network continuously iterates and trains to optimize the train operation plan and passenger flow scheduling strategy, thereby improving the overall operation efficiency and passenger experience of the system;
[0061] 5.2 The specific steps during the training process include:
[0062] Initialize network parameters: Initialize the experience pool D, set the maximum number of samples M that the experience pool can store: Initialize the parameters w of the prediction network and the parameters w of the target network, and let w = w - Initialize, let w = w - . The prediction network outputs the Q value of each possible action by inputting the current state s z , that is, Q(s z , a z ; w); The target network has the same structure as the prediction network but different parameters. The role of the target network is to provide a stable target value;
[0063] Environment interaction: Based on the current state s z Select the action a according to the optimal action value function z And observe the new state s z+1 And the reward r z , obtain the sample (s z , a z , s z+1 ), and store it in the experience pool D;
[0064] Calculate the target value y z : Batch sample the saved samples from the experience pool D and calculate the target value. is the maximum Q value of the target network in the next state s z+1 and all actions a';
[0065]
[0066] Calculate the loss function L(w): The loss function measures the gap between the predicted Q value and the target Q value. Define the loss function as:
[0067]
[0068] Take the gradient of the loss function L(w) with respect to w: Calculate the gradient of the loss function with respect to the network parameters through backpropagation, and use the gradient descent method to update the network parameters;
[0069]
[0070] Update the target network w- : Every fixed number of steps, copy the prediction network parameter w to the target network w - , to stabilize the calculation of the target value and avoid instability caused by the rapid change of the target value during the training process;
[0071] 5.3 Continuously repeat the above training process to gradually approximate the optimal action value function, and an updated train operation plan and passenger flow scheduling plan can be obtained; in addition, it is necessary to evaluate the model performance and optimize the weight coefficient by comparing the passenger waiting time, car congestion degree, and operation cost under the influence of different parameters, and adjust hyperparameters such as the learning rate and discount factor γ to improve the effect and stability of policy learning.
[0072] Compared with the prior art, the beneficial effects of the present invention at least include:
[0073] The present invention proposes an optimization method for urban rail passenger flow and train collaborative organization based on deep reinforcement learning. In addition, it is necessary to evaluate the model performance and optimize the weight coefficient by comparing the passenger waiting time, car congestion degree, and operation cost under the influence of different parameters, and adjust hyperparameters such as the learning rate and discount factor γ to improve the effect and stability of policy learning.
[0074] Through the deep reinforcement learning framework adopted by the present invention, combined with the dynamic changes of the environment and real-time data, it is possible to comprehensively improve passenger flow distribution, train operation scheduling, and energy consumption optimization on the basis of multi-objective optimization. This method not only has high adaptability and flexibility, but also can continuously optimize the system strategy according to real-time feedback, improving the overall operation efficiency of urban rail transit. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] Figure 1 is the design flow chart of the optimization method for urban rail passenger flow and train collaborative organization based on deep reinforcement learning in the present invention;
[0076] Figure 2 is the schematic diagram of the relevant time for passengers to travel by rail transit in the present invention;
[0077] Figure 3 is the system composition block diagram of the optimization method for urban rail passenger flow and train collaborative organization based on deep reinforcement learning in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0078] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.
[0079] As a specific embodiment of the present invention, the present invention provides an optimization method for urban rail passenger flow and train collaborative organization based on deep reinforcement learning, and the design flow chart is as shown in Figure 1As shown, the schematic diagram of the relevant time is as Figure 2 As shown, the block diagram of the system composition is as Figure 3 shown;
[0080] Step 1: Utilize the AFC data of rail transit to analyze the arrival patterns and travel chains of passengers, and construct a passenger flow OD matrix.
[0081] Step 2: Use the passenger flow OD matrix constructed in Step 1 to train a diffusion model to generate passenger flow OD matrices for different time periods, providing rich data for the optimization of train operation and passenger flow scheduling.
[0082] Step 3: Take the OD matrix generated by the diffusion model in Step 2 as the input passenger flow, design interaction events for elements such as passengers, trains, and stations based on the operating characteristics of the urban rail transit network system, simulate the state changes of the urban rail system, and construct a discrete-event-based urban rail transit network simulation model.
[0083] Step 4: On the basis of the urban rail transit network simulation model, study the collaborative scheduling optimization method of urban rail passenger flow and trains based on DQN. With the goal of minimizing the train traction energy consumption and passenger travel time, construct the system state, action, and reward functions of DQN under the constraints of urban rail operation rules.
[0084] Step 5: Design a training process for the collaborative scheduling optimization method of urban rail passenger flow and trains based on DQN. Through training, obtain the optimized scheduling strategy and apply it to the train operation and passenger flow scheduling of the urban rail transit system, thereby improving the passenger travel experience and the service level of urban rail transit.
[0085] The discrete-event-based urban rail transit network simulation model serves as the simulation environment for deep reinforcement learning. By generating different passenger flow and train operation scenarios, it simulates the situation of passengers between platforms and carriages and the operation of trains.
[0086] In this simulation environment, the DQN network continuously iterates and trains to optimize the train operation plan and passenger flow scheduling strategy, thereby improving the overall operation efficiency and passenger experience of the system.
[0087] Initialize the experience pool D and set the maximum number of samples M that the experience pool can store: Initialize the parameters w of the prediction network and the parameters w - of the target network, and let w = w - .
[0088] The prediction network outputs the Q values of each possible action by inputting the current state s z , that is, Q(s z , a z ; w);
[0089] The target network and the prediction network have the same structure but different parameters. The role of the target network is to provide a stable target value;
[0090] Based on the current state s z Select the action a according to the optimal action-value function z And observe the new state s z+1 And the reward r z , to obtain the sample (s z , a z , r z , s z+1 ), and store it in the experience pool D;
[0091] Batch sample the saved samples from the experience pool D and calculate the target value, which is the maximum Q value of the target network in the next state s z+1 and all actions a';
[0092]
[0093] The loss function measures the gap between the predicted Q value and the target Q value, and the loss function is defined as:
[0094]
[0095] Calculate the gradient of the loss function with respect to the network parameters through backpropagation, and use the gradient descent method to update the network parameters.
[0096]
[0097] Every fixed number of steps, copy the prediction network parameter w to the target network w - , for stable target value calculation, to avoid instability caused by the too rapid change of the target value during the training process.
[0098] By repeatedly executing the above training process, gradually optimize the action-value function, so as to obtain an improved train operation plan and passenger flow scheduling plan.
[0099] The above is only a preferred embodiment of the present invention, and it is not any other form of limitation to the present invention. Any modification or equivalent change made according to the technical essence of the present invention still belongs to the scope protected by the present invention.
Claims
1. A method for optimizing the coordinated organization of urban rail passenger flow and trains based on deep reinforcement learning, characterized by: The following steps are involved: Step 1: Use the AFC data of rail transit to analyze the arrival patterns and travel chains of passengers and construct the passenger flow OD matrix; Step 2: Use the passenger flow OD matrix constructed in step 1 to train the diffusion model and generate passenger flow OD matrices for different time periods, providing rich data for train operation and passenger flow scheduling optimization; Step 3, taking the OD matrix generated by the diffusion model in step 2 as the input passenger flow, designing the interaction events of passengers, trains and station elements according to the operation characteristics of the urban rail transit network system, simulating the state changes of the urban rail transit system, and constructing an urban rail transit network simulation model based on discrete events; Step 4: Based on the urban rail transit network simulation model, the DQN-based urban rail passenger flow and train coordinated scheduling optimization method is studied. The system state, action and reward function of DQN are constructed under the constraints of urban rail operation rules with the goal of minimizing train traction energy consumption and passenger travel time. Step 5: Design a training process for the DQN-based urban rail passenger flow and train coordinated scheduling optimization method, obtain the optimized scheduling strategy through training, and apply it to the train operation and passenger flow scheduling of the urban rail transit system, thereby improving the passenger travel experience and the service level of urban rail transit.
2. The urban rail passenger flow and train coordinated organization optimization method based on deep reinforcement learning according to claim 1 is characterized in that: Step 1 is as follows: Using the processed rail transit AFC data, we can obtain the passengers’ entry and exit time, departure station and destination station information. By analyzing the passengers’ arrival patterns and travel chains, we can generate the passenger flow OD matrix with a time granularity of ten minutes or half an hour Δt, providing data support for passenger flow generation. The OD matrix at time t is represented by X t ∈R N×N , where N is the number of sites, all OD matrices under Δt are combined into a time series three-dimensional matrix X h =[X Δt , ..., X hΔt ], where h represents the number of historical time granularities, and the dimension is h×N×N.
3. The urban rail passenger flow and train coordinated organization optimization method based on deep reinforcement learning according to claim 1 is characterized in that: Step 2 is as follows: The passenger flow OD matrix is predicted using the diffusion spatiotemporal graph model. The graph structure is represented by G = (V, E, A), where V is the set of stations, corresponding to the starting station i and the terminal station j of the OD matrix; E represents the connection attribute between stations; A is the adjacency matrix used to describe the topological structure between sites, X t [i, j] represents the passenger flow from station i to station j at time t; The goal of passenger flow prediction is to input the historical OD matrix X H =[X Δt , ..., X HΔt ], H≤h and graph structure G, predict the future OD matrix X F =[X H+Δt , ..., X H+FΔt ], where F represents the number of prediction time granularities: (X H ;G)→[X H+Δt ,…X H+FΔt ]:=X F 。 4. The urban rail passenger flow and train coordinated organization optimization method based on deep reinforcement learning according to claim 1 is characterized in that: Step 3 is as follows: 3.1 In the urban rail transit network simulation model based on discrete events, three entities are involved, namely rail transit trains, rail transit stations, and passengers traveling by rail transit; A train is a vehicle that carries passengers and needs to run on a track according to the timetable. Train attributes include number, running direction, line number, departure station, stopover station, and terminal station. Stations are the main places where passengers gather and disperse. Station attributes include station name, line, whether it is a transfer station, number of the stopping train, and the number of passengers on the current platform; Passengers are the service objects in the rail transit system. Passenger attributes include number, departure station, destination station, travel path and arrival time; 3.2 In the urban rail transit network passenger flow simulation model based on discrete events, trains need to trigger a series of events to update their status and interact with stations and passengers. Events include train initialization, train operation, and stop; 3.3 In the simulation model of train operation and passenger flow status, passenger events include initialization, entry, exit, waiting, transfer, and boarding and alighting.
5. The urban rail passenger flow and train coordinated organization optimization method based on deep reinforcement learning according to claim 4 is characterized in that: In the simulation model of train operation and passenger flow status in 3.3, the passenger events are described in detail as follows: Passenger initialization: At the beginning of the simulation, according to the OD matrix generated in step 2, the initial information of the passengers is generated, including the number, departure station, destination station, travel path and arrival time, and λ = X nΔ t[i, j] represents the passenger flow from station i to station j at the nth time granularity; the arrival time of each passenger p is Following an exponential distribution: represents the arrival time of passenger p from station i to station j during the time period from nΔt to (n+1)Δt; Passenger p's arrival time for: Based on the arrival time of each passenger, the lognormal distribution is used to generate the relevant time of the passenger's travel process The lognormal distribution is: Where μ is the logarithmic mean of the generation time, σ is the standard deviation of the generation time, and The log-normal distribution is The probability density at the point is assumed to follow a log-normal distribution. The transfer time of passengers follows a log-normal distribution Passengers' walking time out of the station follows a log-normal distribution After generating the passenger entry time, the Logit model is used to allocate the passenger travel path: in represents the probability that a passenger chooses path k among paths ij, K ij Represents all optional paths in the OD matrix for the ij path, represents the travel cost of path k in the OD matrix, θ represents the passenger's familiarity with the urban rail transit network, d represents the passenger's acceptance of the inductive information, and β represents the impact of the inductive information on path selection; Passengers enter, exit, wait, transfer, get on and off the train: Passengers arrive at the station according to the generated entry time, and the train arrives at the corresponding station within the specified time according to the set timetable, which will trigger the passengers' boarding and disembarking events. Passengers' disembarking needs to be judged in combination with the path selection. If the disembarking conditions are not met, they will continue to stay on the train. If the conditions are met, the passenger's transfer event or exit event will be triggered. Passengers' boarding is constrained by the train direction, the current number of passengers on the train, and the passenger path. If the boarding conditions are not met, they will continue to wait.
6. The urban rail passenger flow and train coordinated organization optimization method based on deep reinforcement learning according to claim 1 is characterized in that: Step 4 is as follows: The passenger flow scheduling problem and train operation optimization problem are modeled as Markov decision problems in reinforcement learning, and DQN is used to find the optimal strategy. The learning process includes the following: System status: Status includes platform congestion, section congestion, and train position; Action: Train departure interval Δt train , interval running time Station stop time Section recommendation a m , a m It is a 0-1 variable. When the m segment is recommended, it is 1, otherwise it is 0, m∈N; Rewards: Designed based on optimization goals, as described below: Train operating costs: C op =C f D t C op represents the train operating cost, C f Denotes the fuel cost per unit distance, D t Indicates the train running distance; Passenger p's waiting time: In the formula represents the waiting time of passenger p, represents the arrival time of train x at platform y, which is obtained from the urban rail transit network simulation model; To combine these objectives, they are weighted together into a total reward function R(t), the goal of which is to maximize this function: α1 is the weight of train operating costs, α2 is the weight of passenger waiting time, and η is the overall scaling factor. The values of α1, α2, and η can be adjusted according to the emphasis of different actual needs to balance the impact of different objectives.
7. The urban rail passenger flow and train coordinated organization optimization method based on deep reinforcement learning according to claim 1 is characterized in that: Step 5 is as follows: 5.1 The urban rail transit network simulation model based on discrete events is used as a simulation environment for deep reinforcement learning. By generating different passenger flow and train operation scenarios, it simulates the situation of passengers on the platform and between carriages and the operation of trains. In this simulation environment, the DQN network optimizes the train operation plan and passenger flow scheduling strategy through continuous iterative training, thereby improving the overall operational efficiency of the system and passenger experience; 5.2 The specific steps in the training process include: Initialize network parameters: Initialize the experience pool D and set the maximum number of samples M that can be stored in the experience pool: Initialize the parameters w of the prediction network and w of the target network - Initialize, let w = w - The prediction network is fed with the current state s z To output the Q value of each possible action, that is, Q(s z , a z ; w); The target network and the prediction network have the same structure but different parameters. The role of the target network is to provide a stable target value; Environmental interaction: based on the current state z Select action a according to the optimal action value function z And observe the new state s z+1 and reward r z , get the sample (s z , a z , r z ,s z+1 ) and store it in the experience pool D; Calculate the target value y z : Sample the saved samples in batches from the experience pool D and calculate the target value. is the target network in the next state s z+1 and the maximum Q value among all actions a′; Calculate the loss function L(w): The loss function measures the gap between the predicted Q value and the target Q value. The loss function is defined as: Find the gradient of the loss function L(w) with respect to w: calculate the gradient of the loss function with respect to the network parameters through back propagation, and use the gradient descent method to update the network parameters; Update target network w - : Every fixed number of steps, copy the predicted network parameters w to the target network w - , calculated with stable target values, to avoid instability caused by rapid changes in target values during training; 5.3 By repeating the above training process and gradually approaching the optimal action value function, the updated train operation plan and passenger flow scheduling plan can be obtained; in addition, it is necessary to compare the passenger waiting time, carriage congestion and operating cost under the influence of different parameters, evaluate the model performance and optimize the weight coefficient to adjust the learning rate, discount factor γ and other hyperparameters to improve the effect and stability of strategy learning.
Citation Information
Patent Citations
Method and system for precisely inducing passenger flow in multiple scenes of urban rail transit
CN110428117A
Train operation scheme adjusting method and device, electronic equipment and storage medium
CN115352502A
Urban rail transit short-time passenger flow prediction method considering sudden factors
CN116128122A
Comprehensive optimization method for energy consumption and passenger waiting time of subway train
CN117350142A
Urban rail transit space-time passenger flow prediction system and method
CN117689080A
Cited By
Existing subway station transformation and transfer efficiency improving system based on simulation prediction
CN120764212A