Urban rail train group autonomous decision-making method

By constructing a dynamic passenger flow model and an autonomous decision-maker based on Sarsa reinforcement learning, combined with the vehicle-to-vehicle communication topology, the problems of dynamic passenger flow adaptive adjustment and distributed collaborative decision-making in the urban rail train operation control system were solved, thereby optimizing train intervals and improving operational efficiency and passenger service quality.

CN122009288APending Publication Date: 2026-05-12SOUTHWEST JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SOUTHWEST JIAOTONG UNIV
Filing Date
2026-01-30
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

The existing urban rail train operation control system lacks the ability to adaptively adjust to dynamic passenger flow. The centralized control architecture has the risk of single point of failure and lacks a distributed collaborative decision-making mechanism among multiple trains, making it difficult to achieve real-time matching of capacity and demand, resulting in excessively long passenger waiting times and high operating costs.

Method used

An autonomous decision-making method for urban rail train groups is adopted. By constructing a dynamic passenger flow model and an autonomous decision-maker based on Sarsa reinforcement learning, combined with the train-to-train communication topology, distributed collaborative optimization among trains is achieved. A multi-objective distributed optimization model is designed to optimize passenger dissatisfaction and operating costs.

Benefits of technology

It enables adaptive adjustment of operating intervals under dynamic passenger flow conditions, improves the matching degree between capacity and demand, reduces passenger waiting time and operating costs, enhances the robustness and scalability of the system, and adapts to different operating scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122009288A_ABST
    Figure CN122009288A_ABST
Patent Text Reader

Abstract

The invention provides an autonomous decision-making method for an urban rail train group, and belongs to the field of urban rail transit train operation control and intelligent scheduling. According to the method, the running interval self-adaptive adjustment and collaborative optimization of multiple trains under the dynamic passenger flow condition are realized by constructing a distributed collaborative framework based on train-train communication, establishing a dynamic passenger flow prediction model and designing an autonomous decision maker based on Sarsa reinforcement learning; therefore, the passenger service quality is improved and the operation cost is reduced on the premise of ensuring the operation safety.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of urban rail transit train operation control and intelligent scheduling, and particularly relates to an autonomous decision-making method for urban rail train groups. Background Technology

[0002] In daily operation, the efficiency and service quality of urban rail transit depend heavily on the matching degree between train schedules and actual passenger demand. Currently, urban rail systems generally use pre-determined fixed timetables for train dispatching. While this can meet the needs of regular passenger flow changes, it is difficult to achieve dynamic matching between passenger flow and train capacity when facing irregular passenger flow fluctuations (such as sudden large passenger flows or tidal passenger flows). This results in excessively high train occupancy rates and long waiting times for passengers during some periods, while other periods experience wasted capacity.

[0003] Traditional train operation optimization methods mainly include timetable optimization based on historical data, real-time scheduling based on centralized control, and rule-based emergency response strategies. For example, existing Automatic Train Control (ATO) systems primarily rely on preset operating curves and lack the ability to adapt to dynamic passenger flow; while centralized scheduling systems can achieve global optimization, they suffer from high computational complexity and large response delays, making it difficult to handle the real-time decision-making needs of large-scale train groups; rule-based methods rely on human experience, lack self-learning capabilities, and cannot adapt to complex and ever-changing operating environments. Existing technologies have the following drawbacks: Train operation plans lack the ability to adaptively adjust to dynamic passenger flow, making it difficult to achieve real-time matching of capacity and demand; centralized control architecture has the risk of single point of failure and is difficult to extend to large-scale train group collaborative optimization; lack of distributed collaborative decision-making mechanism among multiple trains makes it impossible to fully utilize train-to-train communication technology to achieve intelligent scheduling of train groups; traditional optimization methods are difficult to simultaneously consider multiple objectives of passenger service quality (waiting time, load factor) and operating costs (energy consumption, station dwell time). Summary of the Invention

[0004] To address the aforementioned shortcomings in existing technologies, this invention provides an autonomous decision-making method for urban rail train groups, which solves the problems of train operation plans lacking adaptive adjustment capabilities to dynamic passenger flow, centralized control architecture having single-point failure risks, and lacking a distributed collaborative decision-making mechanism among multiple trains.

[0005] To achieve the above objectives, the technical solution adopted by this invention is: an autonomous decision-making method for urban rail train groups, comprising the following steps: S1. Initialize the system and acquire real-time operating status data of each train and passenger flow data of each station on the urban rail line; S2. Construct a dynamic passenger flow model, and based on historical passenger flow arrival rates and current real-time passenger flow data, use the dynamic passenger flow model to predict the passenger flow distribution when each train arrives at each platform. S3. Based on the predicted passenger flow distribution, and combined with the operational status data of each train, a communication topology map between trains is constructed. Based on this map, and utilizing the preceding train following communication topology, each train interacts with only the train preceding it to obtain information about that train. i Only with the car in front i -1 Establish communication link, first vehicle i =1 No preceding vehicle connection; S4. Input the acquired real-time station passenger flow data and previous train information into the dynamic passenger flow model, and collaboratively execute the prediction of future passenger flow distribution and update the dynamic passenger flow model; at the same time, based on the updated passenger flow distribution prediction results, construct a multi-objective distributed optimization model for the train group, wherein the multi-objective distributed optimization model for the train group aims to minimize the weighted sum of passenger dissatisfaction and operating costs. S5. Design an autonomous decision-maker based on Sarsa reinforcement learning for each train, and input the train's operating status data, updated passenger flow distribution predictions, and constraints of the train group's multi-objective distributed optimization model into the autonomous decision-maker. Through interaction with the simulation environment, perform offline training to learn the optimal running interval and stop time adjustment strategy. S6. During the deployment phase, the autonomous decision-makers trained offline are used to execute online autonomous decisions, and the optimized control commands output by the autonomous decision-makers of each train are obtained.

[0006] Furthermore, the passenger flow distribution status includes: From the site s Head to the station k Number of arriving passengers :

[0007]

[0008] in, Indicates the passenger arrival rate. Indicates train i Arrival Station s At that moment, Indicates train i The moment of leaving the previous station, Indicates the total number of trains. Indicates the total number of stations on the line. Represents the smoothing coefficient. This indicates the current measured arrival rate. This indicates the arrival rate at the previous moment; platform s Waiting for the train i Total passenger flow :

[0009] in, Indicates platform s The number of stranded passengers; train i On the site s Number of passengers getting off the bus : ; in, Indicates on the site k Boarding and the destination is station s The number of passengers; platform s Total demand of passengers boarding :

[0010] in, Indicates the boarding speed. Indicates the disembarkation rate. S Indicates the total number of platforms. Indicates train i The moment of departure from the station; train i On the site s The actual number of passengers on board : ; train i Leave the station s Number of passengers on board at the time :

[0011] in, Indicates train i Leave the previous station s The number of passengers on the train at that time; platform s Update on the number of stranded passengers : .

[0012] Furthermore, regarding the first train i The expression for information exchange using =1 is as follows:

[0013] in, Indicates the arrival of the first train at the station. k The predicted total passenger flow at that time, This indicates that the first train is located at the station. s Time to site k Passenger flow forecast, Indicates site k Passenger arrival rate This indicates the expected arrival station of the first train. k At that moment, This indicates that the first train has departed the station. s At that moment, s Indicates a platform. S Indicates the total number of sites. Indicates the total number of stations on the line; For excluding the first car i The expression for information exchange between trains other than =1 is as follows:

[0014] in, Indicates train i Arrival Station k The predicted total passenger flow at that time, Indicates train i Previous train numbering, Indicates the preceding sequence of vehicles x On the site k The estimated number of passengers to be transported. Indicates the total number of trains. i Indicates the number of trains.

[0015] Furthermore, the objective function expression of the multi-objective distributed optimization model for the train group is as follows:

[0016]

[0017]

[0018] in, Represents a multi-objective function. f p The function representing passenger dissatisfaction. Represents the operating cost function. , , and All represent weighting coefficients. , , and All represent cost coefficients. Indicates site jAverage passenger wait time Indicates site j The degree to which the train's occupancy rate deviates from the comfort range. Indicates the total number of platforms. Representing an interval j Traction energy consumption, Indicates site j Stop time.

[0019] Furthermore, the constraints of the multi-objective distributed optimization model for the train group include: Punctuality constraints:

[0020] in, Indicates the time when the first train departs from the first station. and These represent the time window boundaries for the first and last stations, respectively. Indicates train i Total runtime Indicates train i The moment of leaving the last station, Indicates the total number of trains. i Indicates the number of trains; Site capacity constraints:

[0021] like ,but

[0022] in, Indicates site s At any moment t Total number of people waiting for the train Indicates the maximum capacity of the platform. Indicates site k Total number of people waiting for the train Indicates platform s The number of stranded passengers, s Indicates a platform. S Indicates the total number of sites; Line average load factor constraint:

[0023] in, Indicates the total number of stations on the line. Indicates train i Leave the station s The number of passengers on the train at that time This indicates the train's rated passenger capacity. This indicates the upper limit of the average full load rate of the line. Indicates the total number of stations on the line; Front and rear vehicle safety protection constraints:

[0024]

[0025] in, Indicates train i Arrival Station s At that moment, Indicates the vehicle in front i -1 Arrival Station s At that moment, Indicates the minimum running interval. Indicates train i Leave the station s At that moment, Indicates the vehicle in front i -1 Leave the station s The moment; Running time and downtime constraints:

[0026]

[0027] in, Indicates train i Arrival Station j At that moment, Indicates train i Leave the previous station j At the moment of -1, Indicates the maximum running interval. j Indicates the number of stations.

[0028] Furthermore, performing offline training in S5 includes: Construct a simulation environment and observe the current station's state within it. The state space expression is as follows:

[0029] in, Representing the state space, j Indicates the number of stations. Indicates the total number of stations on the line; Based on the current site status, combined with the current Q-value table and The strategy selects the action, where, for each site j Define running time adjustment actions and stop time adjustment actions:

[0030]

[0031]

[0032] in, Represents the action space, Indicates site j Runtime adjustment action, Indicates site j The stop time adjustment action, Indicates site j Maximum runtime Indicates site j Minimum runtime, Indicates site j Maximum stop time, Indicates site j The minimum stop time for running time adjustment actions and stop time adjustment actions Negative values ​​indicate a shorter time, positive values ​​indicate a longer time, and 0 indicates that the original plan will be maintained. Execute actions to adjust running time and stop time; Based on the adjustment result, an immediate reward is obtained, and the process transitions to the next state. The expression for the reward function is as follows:

[0033]

[0034]

[0035]

[0036] in, Indicates the state s Execute action a The instant reward function obtained afterward , , and This represents the weight coefficient of each item in the reward function. This represents the waiting time cost coefficient. Indicates train i On the site j The resulting passenger waiting time item This represents the cost coefficient for deviation from full load rate. Indicates train i On the site j The full load rate deviation item, This represents the energy cost coefficient. Indicates train i In the interval j Traction energy consumption item, This represents the cost coefficient for stopping at the station. Indicates train i On the site j Stop time, x Represents the site index variable. y Represents the site index variable. Indicates train i Arrival Station y At that moment, Indicates site y The number of passengers boarding, Indicates site y The boarding speed Indicates the train is at the station. x The number of passengers on the train This indicates the train's rated passenger capacity. and These indicate the train is at the station. x and sites x At the moment of +1, A function representing the relationship between train traction force and speed. v Indicates the train's speed. dt Represents the time derivative. This indicates the time when train i arrives at station x.

[0037] In the new state, select the next action; Based on the chosen next action, update the Q value using the following formula:

[0038] in, Indicates the state Execute action State-action value function Indicates site j The corresponding state, Indicates on the site j The chosen action Indicates the learning rate. Indicates an immediate reward. Indicates the discount factor. Indicates the next station j The state corresponding to +1 This indicates the action selected at the next stop. Indicates the next state Execute action Q-value, learning rate An adaptive decay strategy is adopted, gradually reducing the decay rate as the number of training epochs increases:

[0039] in, Indicates the first e Learning rate during training rounds This represents the initial learning rate. Indicates the attenuation coefficient. e Indicates the current training round number; Repeat the above process until the train reaches the final station, completing one training round; offline training is completed when the number of training rounds reaches the maximum value or the Q value converges.

[0040] Furthermore, the aforementioned The expression for the strategy selection action is as follows:

[0041] in, Indicates on the site j The optimal action to choose Represents a random number between 0 and 1. Indicates site j The corresponding state, This represents the learning rate.

[0042] Furthermore, the constraints satisfied by state selection and action space update are as follows:

[0043]

[0044]

[0045]

[0046] in, Indicates the train is at the station. j runtime Indicates site j Runtime adjustment action, Indicates the train is at the station. j Stop time, Showing the site j The stop time adjustment action, Indicates the minimum running interval. Indicates the maximum running interval. Indicates the minimum stopping time. Indicates the maximum dwell time. Indicates the train is at the station. j +1 running time, Indicates the train is at the station. j +1 stop time.

[0047] The beneficial effects of this invention are: This invention achieves adaptive adjustment and collaborative optimization of train intervals under dynamic passenger flow conditions by constructing a distributed collaborative framework based on vehicle-to-vehicle communication, establishing a dynamic passenger flow prediction model, and designing an autonomous decision-maker based on Sarsa reinforcement learning. This improves passenger service quality and reduces operating costs while ensuring operational safety, and has at least the following beneficial effects: Strong adaptability: Through distributed reinforcement learning, the train interval can be dynamically and adaptively adjusted, without the need to preset a fixed timetable. It can cope with irregular passenger flow fluctuations such as tidal passenger flow, and significantly improve the matching degree between transport capacity and demand. Multi-objective optimization: Simultaneously optimize passenger waiting time, train load factor balance, traction energy consumption and station stopping costs to achieve a comprehensive balance between operational service quality and economy. Compared with a fixed timetable, it can reduce the average passenger waiting time by about 5.56% and reduce traction energy consumption by about 1.32%. Distributed architecture: The distributed decision-making framework based on vehicle-to-vehicle communication avoids the single point of failure risk of centralized control, improves system robustness and scalability, and reduces the central computing load by enabling each train to make autonomous decisions. Improved collaboration mechanism: Through collaborative updates of passenger flow forecasts and the preceding train following topology, effective collaboration among multiple trains is achieved, and the subsequent train's decision-making fully considers the preceding train's carrying capacity, avoiding passenger flow backlog caused by blind optimization; Strong generalization ability: Reinforcement learning methods can learn the optimal strategy autonomously from operational data without manual parameter tuning, and can adapt to changes in passenger flow patterns through online learning mechanisms, and have cross-line transfer capabilities; Highly practical: The method can be embedded into existing automatic train control systems and seamlessly integrated with ATO systems. It is applicable to different urban rail lines and operating scenarios, providing technical support for the construction of intelligent urban rail systems. Attached Figure Description

[0048] Figure 1 This is a flowchart of the method of the present invention.

[0049] Figure 2 This is a schematic diagram of a distributed train group collaborative framework.

[0050] Figure 3 A schematic diagram for simulating and validating the dynamic passenger flow prediction model. Detailed Implementation

[0051] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.

[0052] Example like Figure 1 As shown, this invention provides an autonomous decision-making method for urban rail train groups, the implementation of which is as follows: S1. Initialize the system and acquire real-time operating status data of each train and passenger flow data of each station on the urban rail line.

[0053] In this embodiment, real-time operational status data of each train and passenger flow data of each platform are acquired along the urban rail line. Train operational status data includes location, speed, acceleration, arrival time, and departure time; platform passenger flow data includes the number of arriving passengers, waiting passengers, boarding passengers, and alighting passengers.

[0054] Initialize system parameters, including the total number of line stations N, the total number of trains T, and the rated passenger capacity of each train. C max Minimum running interval R min Maximum operating interval R max Minimum stop time D min Maximum stopping time D max Learning rate Discount Factor Exploration rate Reinforcement learning parameters, etc.

[0055] The collection of trains is denoted as: T = {1, 2, ..., NT} The collection of sites is denoted as: S = {1, 2, ..., NS}.

[0056] S2. Construct a dynamic passenger flow model, and based on historical passenger flow arrival rates and current real-time passenger flow data, use the dynamic passenger flow model to predict the passenger flow distribution when each train arrives at each platform. In this embodiment, the dynamic passenger flow model uses an exponential smoothing prediction method combined with a passenger flow accumulation model, as shown in the following expression: Passenger arrival rate prediction model:

[0057] in, express t From the site s Head to the station k Passenger arrival rate This indicates the current measured arrival rate. This represents the smoothing coefficient (range 0.1-0.3). This represents the predicted arrival rate at the previous moment. Passenger flow cumulative prediction model:

[0058] in, Indicates train i From the previous station s -1 Departure time to arrival station s Between moments, from the station s Head to the station k The cumulative number of arriving passengers. Indicates train i Arrive at the current station (site) s (The time frame is used as an upper limit.) Indicates train i Leaving the previous station (station) s The time at which -1 is taken as the lower limit of integration. Loss function for dynamic passenger flow model update:

[0059] in, This represents the total predicted loss. This indicates the predicted passenger flow. This indicates the actual observed passenger flow. This represents the regularization coefficient (usually between 0.01 and 0.1). The L2 regularization term, representing the change in arrival rate, is used to prevent excessive fluctuations in the forecast. This represents the passenger arrival rate vector at the current moment. This represents the passenger arrival rate vector at the previous moment.

[0060] This dynamic passenger flow model continuously updates the passenger arrival rate by monitoring passenger flow sensor data (such as AFC gate data and video passenger flow statistics) at each station in real time. Based on the current location and estimated arrival time of the trains, the system predicts the passenger flow distribution when each train arrives at each platform.

[0061] In this embodiment, a time-series-based dynamic passenger flow prediction model is constructed. This model can predict the passenger flow distribution when each train arrives at subsequent platforms based on historical passenger flow arrival rates and current real-time passenger flow data.

[0062] Define from site s Head to the station k Number of arriving passengers :

[0063] Among them, passenger arrival rate Based on historical passenger flow big data statistical analysis, we classify and model the data according to weekdays / holidays and peak / off-peak periods, and update the data in real time using exponential smoothing:

[0064] in, Indicates the passenger arrival rate. Indicates train i Arrival Station s At that moment, Indicates train i The moment of leaving the previous station, Indicates the total number of trains. Indicates the total number of stations on the line. Represents the smoothing coefficient. This indicates the current measured arrival rate. This indicates the arrival rate at the previous moment; Passenger arrival rates are categorized into peak-hour rates based on the characteristics of different time periods. v peak Off-peak rates v off_peak . Define platform s Waiting for the train i Total passenger flow :

[0065] in, Indicates platform s The number of stranded passengers. Define train i On the site s Number of passengers getting off the bus :

[0066] in, Indicates on the site k Boarding and the destination is station s The number of passengers.

[0067] Define platform s Total demand of passengers boarding :

[0068] in, Indicates the boarding speed. Indicates the disembarkation rate. S Indicates the total number of platforms. Indicates train i The moment of departure from the station; definition i On the site s The actual number of passengers on board (Constrained by vehicle capacity):

[0069] Define train i Leave the station s Number of passengers on board at the time :

[0070] in, Indicates train i Leave the previous station s The number of passengers on the train at that time.

[0071] Define platform s Update on the number of stranded passengers : .

[0072] S3. Based on the predicted passenger flow distribution, and combined with the operational status data of each train, a communication topology map between trains is constructed. Based on this map, and utilizing the preceding train following communication topology, each train interacts with only the train preceding it to obtain information about that train. i Only with the car in front i -1 Establish communication link, first vehicle i =1 No preceding vehicle connection; In this embodiment, as Figure 2 As shown, a distributed train group collaboration framework based on vehicle-to-vehicle communication is constructed. Figure 2 This is a schematic diagram of a distributed train group collaborative framework, illustrating the vehicle-to-vehicle communication network structure based on a leading-vehicle following topology. Figure 2 In this system, each train node interacts with other trains via onboard units (OBUs). Each train establishes a communication link only with the train preceding it, forming a chain-like communication topology. The diagram clearly illustrates the direction of information flow between trains, including status information (location, speed, passenger capacity) transmitted from the preceding train to the following train, as well as the estimated number of passengers to be transported from each platform, thus enabling collaborative updates to passenger flow forecasts. This distributed communication architecture avoids the single-point-of-failure risk of centralized control, improving the system's robustness and scalability. The train communication topology graph G(V, A) is defined, where V represents the set of train nodes, and A is an adjacency matrix representing the communication connections between trains. A preceding-train-following topology is adopted, meaning each train only interacts with the train preceding it.

[0073] The adjacency matrix is ​​defined as:

[0074] Based on this topology, each train obtains the status and passenger information of the preceding train through vehicle-to-vehicle communication, and updates its dynamic passenger flow model by combining this information with platform passenger flow data collected by its own sensors. For the first car (i=1):

[0075] in, Indicates the arrival of the first train at the station. k The predicted total passenger flow at that time, This indicates that the first train is located at the station. s Time to site k Passenger flow forecast, Indicates site k Passenger arrival rate This indicates the expected arrival station of the first train. k At that moment, This indicates that the first train has departed the station. s At that moment, s Indicates a platform. S Indicates the total number of sites. Indicates the total number of stations on the line; For subsequent trains (i∈{2,...,T}):

[0076] in, Indicates train i Arrival Station k The predicted total passenger flow at that time, Indicates train i Previous train numbering, Indicates the preceding sequence of vehicles x On the site k The estimated number of passengers to be transported. Indicates the total number of trains. i This represents the number of trains. In the formula, the second term subtracts the number of passengers already transported by the previous train, achieving a collaborative update of passenger flow forecasts.

[0077] In this embodiment, in the distributed collaborative framework, the train communication protocol adopts a periodic broadcast mechanism. When each train arrives at a station, it broadcasts its own status information to subsequent trains, including its current location, speed, passenger capacity, estimated departure time, and the estimated number of passengers to be transported at each platform. The communication cycle is synchronized with the train arrival interval to ensure real-time information.

[0078] S4. Input the acquired real-time station passenger flow data and previous train information into the dynamic passenger flow model, and collaboratively execute the prediction of future passenger flow distribution and update the dynamic passenger flow model; at the same time, based on the updated passenger flow distribution prediction results, construct a multi-objective distributed optimization model for the train group, wherein the multi-objective distributed optimization model for the train group aims to minimize the weighted sum of passenger dissatisfaction and operating costs. In this embodiment, a multi-objective distributed optimization model for train groups oriented towards tidal passenger flow is established, with the objective function being the weighted minimization of passenger dissatisfaction and total operating cost:

[0079] Passenger dissatisfaction function f p Defined as:

[0080] Operating cost function Defined as:

[0081] in, Represents a multi-objective function. f p The function representing passenger dissatisfaction. Represents the operating cost function. , , and All represent weighting coefficients. , , and All represent cost coefficients. Indicates site j Average passenger wait time Indicates site j The degree to which the train's occupancy rate deviates from the comfort range. Indicates the total number of platforms. Representing an interval j Traction energy consumption, Indicates site j Stop time.

[0082] The constraints include: Punctuality constraints:

[0083] in, Indicates the time when the first train departs from the first station. and These represent the time window boundaries for the first and last stations, respectively. This represents the total travel time of train i. This indicates the time when train i departs from the last station. Indicates the total number of trains. i Indicates the number of trains; Platform capacity constraints:

[0084] like ,but

[0085] in, Indicates site s At any moment t Total number of people waiting for the train Indicates the maximum capacity of the platform. Indicates site k Total number of people waiting for the train Indicates platform s The number of stranded passengers, s Indicates a platform. S Indicates the total number of sites; Line average load factor constraint:

[0086] in, Indicates the total number of stations on the line. Indicates train i Leave the station s The number of passengers on the train at that time This indicates the train's rated passenger capacity. This indicates the upper limit of the average full load rate of the line; Front and rear vehicle safety protection constraints:

[0087]

[0088] in, Indicates train i Arrival Station s At that moment, Indicates the vehicle in front i -1 Arrival Station s At that moment, Indicates the minimum running interval. Indicates train i Leave the station s At that moment, Indicates the vehicle in front i -1 Leave the station s The moment; Running time and downtime constraints:

[0089]

[0090] in, Indicates train i Arrival Station j At that moment, Indicates train i Leave the previous station j At the moment of -1, Indicates the maximum running interval. j Indicates the number of stations.

[0091] S5. Design an autonomous decision-maker based on Sarsa reinforcement learning for each train, and input the train's operating status data, updated passenger flow distribution predictions, and constraints of the train group's multi-objective distributed optimization model into the autonomous decision-maker. Through interaction with the simulation environment, perform offline training to learn the optimal running interval and stop time adjustment strategy. In this embodiment, each train independently performs Sarsa (State-Action-Reward-State-Action) training, continuously optimizing its decision-making strategy through interaction with the environment. During the training process, each train follows this procedure: Build a simulation environment and observe the current state within the simulation environment. (Current site); Based on the current site's state, combined with the current Q-value table and Select the action based on the strategy; execute the action, adjusting the running time and stop time; obtain immediate rewards based on the adjustment results. And move to the next state. ; Select the next action in the new state Update Q value: Repeat the above process until the train reaches its final destination.

[0092] Training termination conditions: The number of training rounds reaches the maximum value E; the rate of change of Q value is lower than the set tolerance. The average cumulative reward converges to a stable value.

[0093] In this embodiment, the manifestation of constraints during the training process is explained as follows: The constraints of the multi-objective distributed optimization model for train groups are reflected in the following ways: 1. The state space contains constraint-related variables: the preceding train interval reflects safety protection constraints, the platform passenger flow reflects station capacity constraints, and the train load factor reflects the line average load factor constraints. 2. Action space setting boundary constraints: The range of values ​​for the running time and stop time adjustment actions directly corresponds to the upper and lower limits of the constraint conditions, ensuring that the adjusted time meets the requirements.

[0094] 3. Constraint checking and penalty mechanism: If the selected action results in a violation of constraints such as punctuality, capacity or safety interval, the system rejects the action and gives a negative reward penalty, forcing the decision-maker to learn a strategy that satisfies the constraints.

[0095] The interaction with the simulation environment is manifested in the following way: The training process is implemented through interaction with the simulation environment, as detailed below: 1. Construct a simulation environment: Establish a simulation system that includes lines, trains, passenger flow, and platforms to simulate real-world operating scenarios. 2. Interactive training process: The decision-maker observes the state of the simulation environment → selects and adjusts actions → the simulation environment updates the state of train position, passenger flow distribution, etc. according to the actions → calculates indicators such as waiting time and energy consumption → returns to the new state and reward → the decision-maker updates the learning parameters, and the process is repeated. 3. Multi-scenario training: The simulation environment randomly generates different scenarios such as morning peak, evening peak, off-peak, and sudden passenger flow, enabling the decision-maker to learn strategies to adapt to various operational situations.

[0096] In this embodiment, an autonomous decision-maker based on the Sarsa (State-Action-Reward-State-Action) algorithm is designed for each train, and adaptive adjustment of running interval and stopping time is achieved through reinforcement learning.

[0097] State space definition: The state space consists of site indices.

[0098] in, Representing the state space, j Indicates the number of stations. This indicates the total number of stations on the line.

[0099] Action space definition: For each station j Define running time adjustment actions and stop time adjustment actions:

[0100]

[0101]

[0102] Among them, for runtime adjustment actions and stop time adjustment actions Negative values ​​indicate a reduction in time, positive values ​​indicate an extension of time, and 0 indicates maintaining the original plan. In the formula, Represents the action space, Indicates site j Runtime adjustment action, Indicates site j The stop time adjustment action, Indicates site j Maximum runtime Indicates site j Minimum runtime, Indicates site j Maximum stop time, Indicates site j The minimum stop time for running time adjustment actions and stop time adjustment actions Negative values ​​indicate a shorter time, positive values ​​indicate a longer time, and 0 indicates that the original plan will be maintained.

[0103] Reward Function Design: Define a comprehensive reward function, considering passenger waiting time, load balancing, energy costs, and station stopping costs.

[0104] Passenger waiting time item:

[0105] Load factor deviation (target load factor 60%):

[0106] Traction energy consumption items:

[0107] in, Indicates the state s Execute action a The instant reward function obtained afterward , , and This represents the weight coefficient of each item in the reward function. This represents the waiting time cost coefficient. Indicates train i On the site j The resulting passenger waiting time item This represents the cost coefficient for deviation from full load rate. Indicates train i On the site j The full load rate deviation item, This represents the energy cost coefficient. Indicates train i In the interval j Traction energy consumption item, This represents the cost coefficient for stopping at the station. Indicates train i On the site j Stop time,x Represents the site index variable. y Represents the site index variable. Indicates train i Arrival Station y At that moment, Indicates site y The number of passengers boarding, Indicates site y The boarding speed Indicates the train is at the station. x The number of passengers on the train This indicates the train's rated passenger capacity. and These indicate the train is at the station. x and sites x At the moment of +1, A function representing the relationship between train traction force and speed. v Indicates the train's speed. dt Represents the time derivative. This indicates the time when train i arrives at station x.

[0108] Action selection strategy: Adopt Strategy for action selection:

[0109] in, Indicates on the site j The optimal action to choose Represents a random number between 0 and 1. Indicates site j The corresponding state, Indicates the learning rate. This represents the action that maximizes the Q value.

[0110] All constraints must be met. If the selected action violates the constraints, a new action must be selected.

[0111] State-action value function update: Q-value is updated using the Sarsa algorithm.

[0112] in, Indicates the learning rate. This represents the discount factor.

[0113] Learning rate An adaptive decay strategy is adopted, gradually reducing the decay rate as the number of training epochs increases:

[0114] in, Indicates the first eLearning rate during training rounds This represents the initial learning rate. Indicates the attenuation coefficient. e This indicates the current training round number. This mechanism ensures rapid learning in the early stages of training and fine-tuning in the later stages.

[0115] State selection and action space update: Select the corresponding state based on the current station location of the train, and dynamically update the action space of the current station based on the decision results of the previous station to ensure that the safety interval and timetable constraints are met.

[0116] Constraints:

[0117]

[0118]

[0119] in, Indicates the train is at the station. j runtime Indicates site j Runtime adjustment action, Indicates the train is at the station. j Stop time, Showing the site j The stop time adjustment action, Indicates the minimum running interval. Indicates the maximum running interval. Indicates the minimum stopping time. Indicates the maximum dwell time. Indicates the train is at the station. j +1 running time, Indicates the train is at the station. j +1 stop time.

[0120] S6. During the deployment phase, the autonomous decision-makers trained offline are used to execute online autonomous decisions, and the optimized control commands output by the autonomous decision-makers of each train are obtained.

[0121] In this embodiment, during the system deployment phase, each train performs online autonomous decision-making based on the trained Q-value table. Before arriving at each station, each train updates its dynamic passenger flow model based on real-time collected passenger flow data and collaborative information obtained from communication with the preceding train, and selects the optimal action based on the current state.

[0122] When a significant change in passenger flow distribution is detected (such as passenger flow rate exceeding the historical statistical range), the system automatically triggers an adaptive update mechanism: Passenger flow model parameter update: Re-estimate arrival rate based on passenger flow data within the most recent time window. boarding speed With disembarkation rate ; Q-value table fine-tuning: Retain the learned Q-value table as the initial value, and only perform a small number of rounds of online learning updates on the Q-values ​​of the affected sites; Exploration rate dynamic adjustment: When environmental changes are detected, the exploration rate is temporarily increased. This enhances the ability to explore new strategies.

[0123] Control command output and execution: Each train's autonomous decision-maker outputs optimized control commands, including: Adjusted running time and station dwell time; optimized speed curve (solved based on running time constraints); train passenger status and platform passenger flow prediction information (for communication with subsequent trains). Control commands are sent to the Automatic Train Operation (ATO) system for execution after being securely verified by the central monitoring system.

[0124] In this embodiment, the method can be extended to random delay suppression scenarios. When a train is detected to be delayed, the delayed train, the preceding train, and the following train each execute differentiated decision objectives: the delayed train takes minimizing the delay time as its main objective and may adopt a cross-station strategy; the preceding train actively absorbs the passenger flow stranded on the platform; and the following train restricts its running interval to avoid the propagation of delays, thereby realizing a coordinated response of the train group to sudden disturbances.

[0125] The present invention will be further described below.

[0126] The autonomous decision-making method for adaptive operation intervals of urban rail train groups based on tidal passenger flow using distributed reinforcement learning is implemented as follows: A1. System Deployment and Parameter Setting: The system of this invention was first deployed on Line 1 of a certain urban rail transit system. This line has 14 stations (NS=14), 8 operating trains (NT=8), and a rated passenger capacity of Cmax=310 people per train. Real-time data on train position, speed, and the number of passengers waiting and boarding / alighting at each platform are collected using onboard sensors and platform passenger flow detection equipment. System parameters are set as follows: minimum safe operating interval Rmin=90 seconds, maximum operating interval Rmax=180 seconds, minimum stop time Dmin=20 seconds, maximum stop time Dmax=60 seconds. The Sarsa learning parameters are set as follows: initial learning rate... Discount factor Initial exploration rate The maximum number of training rounds is E=5000 rounds.

[0127] A2. Historical Passenger Flow Data Analysis and Modeling: Based on the 6 months of historical operational data collected in step A1, statistical analysis is performed on the passenger arrival rate at each station. The time periods are divided into four periods: morning peak (7:00-9:00), evening peak (17:00-19:00), off-peak period, and nighttime low peak period.

[0128] Arrival rate models for typical stations (such as transfer stations and commercial area stations) are obtained through data fitting. For example, the average arrival rate of a certain transfer station during the morning rush hour is... =8 people / second, off-peak period is =3 people / second; boarding speed =0.8 people / second, disembarkation rate =1.2 people / second. These parameters serve as the initial inputs to the passenger flow forecasting model.

[0129] A3. Train-to-Train Communication Network Configuration: Based on the train deployment in step A1, construct a leading-train following communication topology. Each train is equipped with an onboard unit (OBU) supporting DSRC or LTE-V communication protocols. i Only with the car in front i -1 Establish communication link, first vehicle ( i =1) No preceding train connection. Configure communication protocol: Each train broadcasts an information packet upon arrival at a station, including vehicle ID, current location, speed, passenger capacity, estimated departure time, and the estimated number of passengers to be transported at each station. The broadcast period is set to be triggered upon arrival, with a delay of less than 100 milliseconds to ensure that following trains receive the status of the preceding train in a timely manner.

[0130] A4. Distributed Passenger Flow Forecasting Collaborative Update: Input the real-time passenger flow data obtained in step A1 and the preceding vehicle information obtained in step A3 into the dynamic passenger flow model established in step A2, and perform collaborative forecasting updates. For example... Figure 3 As shown, the complete evolution of passenger flow from arrival at the station, waiting for the train, boarding the train, to being on the train is dynamically tracked in the model.

[0131] Taking train number 3 as an example, when it is at station 5 and about to depart for station 6: (1) The estimated number of passengers to be transported by the first train and the second train at the sixth station is 150 and 120 respectively, obtained through vehicle-to-vehicle communication; (2) The number of people waiting at station 6 is 380, as obtained from the platform sensors. (3) Based on the historical rate model, the predicted number of new arriving passengers from the current time until the third train arrives at the sixth station is approximately [number missing]. = 8 90 = 720 people; (4) Update the estimated total number of passengers waiting at station 6 when train 3 arrives: people. Similarly, recursive predictions are made for stations 7 to 14 to form a complete prediction of future passenger flow distribution.

[0132] A5. Sarsa Reinforcement Learning Training: Train operation status data, updated passenger flow distribution predictions, and constraints of the multi-objective distributed optimization model for train groups are input into the autonomous decision-maker for offline training. The training process uses a simulated environment, replaying historical passenger flow data to simulate real-world operational scenarios. Taking a specific training round as an example: (1) Train 5 departs from the first station, and its state is initialized to... ; (2) At station 1, according to the Q-value table and Strategy selection action: runtime adjustment Seconds, stop time adjustment Second; (3) After the action is executed, the train proceeds to the second station according to the adjusted operating schedule, during which the immediate reward is calculated: Although the extended travel time increases the waiting time for passengers at the second station, it avoids excessively high train occupancy rates, resulting in a comprehensive reward of [amount missing]. ; (4) Upon arrival at station 2, the status transitions to... Choose a new action Second, Second; (5) Update the Q value according to the Sarsa update formula:

[0133] (6) Repeat the above process until the 5th train arrives at the terminal station, completing one training round.

[0134] After 5,000 rounds of training, the Q-value table converged, and the decision-making strategies of each train tended to stabilize.

[0135] A6. Online Operation and Performance Verification: Deploy the autonomous decision-maker trained offline in step A5 to the actual operating environment and conduct online testing during the morning peak hours (7:30-8:30).

[0136] System operation process: (1) Each train obtains an initial passenger flow forecast and Q-value table from the central system before departure; (2) During train operation, real-time passenger flow data at the platform is collected, and information about the preceding train is obtained through train-to-train communication to dynamically update passenger flow forecasts; (3) Before arriving at each station, based on the current status Query the Q-value table and select the optimal action (exploration rate reduced to 0). ); (4) Execute actions, adjust the running speed curve and station dwell time, and automatically control the train operation by the ATO system; (5) After leaving the station, broadcast update information to the following train and record the running data for subsequent model updates.

[0137] A7. Adaptive Update Mechanism Verification: To verify the system's adaptive capability, a large-scale event was simulated near a station during the third week of operation, causing a sudden surge in passenger flow. The system detected a sudden increase in the measured arrival rate at station 10 from an average of 4 people / second to 15 people / second, triggering an adaptive update. (1) Automatic update of passenger flow rate parameters:

[0138] (2) Temporarily increase the exploration rate to This will enhance the exploration of new strategies; (3) Conduct 100 rounds of online fine-tuning training on the Q value involving the 10th station, and optimize the decision-making strategy using the latest passenger flow data; (4) After the fine-tuning is completed, the system automatically adjusts the running interval and stopping time of each train at stations 9 to 11, shortening the interval to increase the capacity and effectively alleviate passenger backlog.

[0139] A8. Comprehensive Performance Evaluation: By comparing the operational data of a fixed timetable under the same passenger flow scenario with that of the method of this invention, the system performance is evaluated: (1) Passenger waiting time: Under this invention, the total waiting time for all passengers is 10254 hours, with an average waiting time of 365.3 seconds; under a fixed timetable, the total waiting time is 10851 hours, with an average waiting time of 386.8 seconds. The method of this invention reduces the waiting time for each passenger by 21.5 seconds, saving approximately 5.56%. (2) Train load factor distribution: Under a fixed timetable, the load factor distribution of each train section is uneven, with some sections having a load factor of over 87% and others only 28%; Under the method of this invention, the load factor of each train is mostly concentrated in the comfortable range of 45%-75%, achieving a better load factor balance. (3) Traction energy consumption: Under a fixed timetable, the total traction energy consumption of 8 trains is 1577.2 kWh; under the method of this invention, the total energy consumption is 1556.4 kWh, resulting in energy savings of 20.8 kWh and an energy saving rate of approximately 1.32%. The reduction in energy consumption is mainly due to the optimization of the running interval, which reduces unnecessary acceleration and deceleration operations; (4) System responsiveness: The average delay of autonomous decision-making calculation for each train is less than 50 milliseconds, which meets the requirements of real-time control; the success rate of train-to-train communication is 99.8%, and the average information transmission delay is 85 milliseconds.

[0140] In summary, through the above implementation methods, the present invention enables urban rail train groups to make autonomous decisions and coordinate optimization under dynamic tidal passenger flow conditions, significantly improving the quality of operation services and economic benefits, and verifying the effectiveness and practicality of the method.

Claims

1. A method for autonomous decision-making of urban rail train groups, characterized in that, Includes the following steps: S1. Initialize the system and acquire real-time operating status data of each train and passenger flow data of each station on the urban rail line; S2. Construct a dynamic passenger flow model, and based on historical passenger flow arrival rates and current real-time passenger flow data, use the dynamic passenger flow model to predict the passenger flow distribution when each train arrives at each platform. S3. Based on the predicted passenger flow distribution, and combined with the operational status data of each train, a communication topology map between trains is constructed. Based on this map, and utilizing the preceding train following communication topology, each train interacts with only the train preceding it to obtain information about that train. i Only with the car in front i -1 Establish communication link, first vehicle i =1 No preceding vehicle connection; S4. Input the acquired real-time station passenger flow data and previous train information into the dynamic passenger flow model, and collaboratively execute the prediction of future passenger flow distribution and update the dynamic passenger flow model; at the same time, based on the updated passenger flow distribution prediction results, construct a multi-objective distributed optimization model for the train group, wherein the multi-objective distributed optimization model for the train group aims to minimize the weighted sum of passenger dissatisfaction and operating costs. S5. Design an autonomous decision-maker based on Sarsa reinforcement learning for each train, and input the train's operating status data, updated passenger flow distribution predictions, and constraints of the train group's multi-objective distributed optimization model into the autonomous decision-maker. Through interaction with the simulation environment, perform offline training to learn the optimal running interval and stop time adjustment strategy. S6. During the deployment phase, the autonomous decision-makers trained offline are used to execute online autonomous decisions, and the optimized control commands output by the autonomous decision-makers of each train are obtained.

2. The autonomous decision-making method for urban rail train groups according to claim 1, characterized in that, The passenger flow distribution status includes: From the site s Head to the station k Number of arriving passengers : in, Indicates the passenger arrival rate. Indicates train i Arrival Station s At that moment, Indicates train i The moment of leaving the previous station, Indicates the total number of trains. Indicates the total number of stations on the line. Represents the smoothing coefficient. This indicates the current measured arrival rate. This indicates the arrival rate at the previous moment; platform s Waiting for the train i Total passenger flow : in, Indicates platform s The number of stranded passengers; train i On the site s Number of passengers getting off the bus : ; in, Indicates on the site k Boarding and the destination is station s The number of passengers; platform s Total demand of passengers boarding : in, Indicates the boarding speed. Indicates the disembarkation rate. S Indicates the total number of platforms. Indicates train i The moment of departure from the station; train i On the site s The actual number of passengers on board : ; train i Leave the station s Number of passengers on board at the time : in, Indicates train i Leave the previous station s The number of passengers on the train at that time; platform s Update on the number of stranded passengers : 。 3. The autonomous decision-making method for urban rail train groups according to claim 1, characterized in that, For the first car i The expression for information exchange using =1 is as follows: in, Indicates the arrival of the first train at the station. k The predicted total passenger flow at that time, This indicates that the first train is located at the station. s Time to site k Passenger flow forecast, Indicates site k Passenger arrival rate This indicates the expected arrival station of the first train. k At that moment, This indicates that the first train has departed the station. s At that moment, s Indicates a platform. S Indicates the total number of sites. Indicates the total number of stations on the line; For excluding the first car i The expression for information exchange between trains other than =1 is as follows: in, Indicates train i Arrival Station k The predicted total passenger flow at that time, Indicates train i Previous train numbering, Indicates the preceding sequence of vehicles x On the site k The estimated number of passengers to be transported. Indicates the total number of trains. i Indicates the number of trains.

4. The autonomous decision-making method for urban rail train groups according to claim 1, characterized in that, The objective function expression of the multi-objective distributed optimization model for the train group is as follows: in, Represents a multi-objective function. f p The function representing passenger dissatisfaction. Represents the operating cost function. , , and All represent weighting coefficients. , , and All represent cost coefficients. Indicates site j Average passenger wait time Indicates site j The degree to which the train's occupancy rate deviates from the comfort range. Indicates the total number of platforms. Representing an interval j Traction energy consumption, Indicates site j Stop time.

5. The autonomous decision-making method for urban rail train groups according to claim 1, characterized in that, The constraints of the multi-objective distributed optimization model for the train group include: Punctuality constraints: in, Indicates the time when the first train departs from the first station. and These represent the time window boundaries for the first and last stations, respectively. Indicates train i Total runtime Indicates train i The moment of leaving the last station, Indicates the total number of trains. i Indicates the number of trains; Site capacity constraints: like ,but in, Indicates site s At any moment t Total number of people waiting for the train Indicates the maximum capacity of the platform. Indicates site k Total number of people waiting for the train Indicates platform s The number of stranded passengers, s Indicates a platform. S Indicates the total number of sites; Line average load factor constraint: in, Indicates the total number of stations on the line. Indicates train i Leave the station s The number of passengers on the train at that time This indicates the train's rated passenger capacity. This indicates the upper limit of the average full load rate of the line; Front and rear vehicle safety protection constraints: in, Indicates train i Arrival Station s At that moment, Indicates the vehicle in front i -1 Arrival Station s At that moment, Indicates the minimum running interval. Indicates train i Leave the station s At that moment, Indicates the vehicle in front i -1 Leave the station s The moment; Running time and downtime constraints: in, Indicates train i Arrival Station j At that moment, Indicates train i Leave the previous station j At the moment of -1, Indicates the maximum running interval. j Indicates the number of stations.

6. The autonomous decision-making method for urban rail train groups according to claim 1, characterized in that, The offline training performed in S5 includes: Construct a simulation environment and observe the current station's state within it. The state space expression is as follows: in, Representing the state space, j Indicates the number of stations. Indicates the total number of stations on the line; Based on the current site status, combined with the current Q-value table and The strategy selects the action, where, for each site j Define running time adjustment actions and stop time adjustment actions: in, Represents the action space, Indicates site j Runtime adjustment action, Indicates site j The stop time adjustment action, Indicates site j Maximum runtime Indicates site j Minimum runtime, Indicates site j Maximum stop time, Indicates site j The minimum stop time for running time adjustment actions and stop time adjustment actions Negative values ​​indicate a shorter time, positive values ​​indicate a longer time, and 0 indicates that the original plan will be maintained. Execute actions to adjust running time and stop time; Based on the adjustment result, an immediate reward is obtained, and the process transitions to the next state. The expression for the reward function is as follows: in, Indicates the state s Execute action a The instant reward function obtained afterward , , and This represents the weight coefficient of each item in the reward function. This represents the waiting time cost coefficient. Indicates train i On the site j The resulting passenger waiting time item This represents the cost coefficient for deviation from full load rate. Indicates train i On the site j The full load rate deviation item, This represents the energy cost coefficient. Indicates train i In the interval j Traction energy consumption item, This represents the cost coefficient for stopping at the station. Indicates train i On the site j Stop time, x Represents the site index variable. y Represents the site index variable. Indicates train i Arrival Station y At that moment, Indicates site y The number of passengers boarding, Indicates site y The boarding speed Indicates the train is at the station. x The number of passengers on the train This indicates the train's rated passenger capacity. and These indicate the train is at the station. x and sites x At the moment of +1, A function representing the relationship between train traction force and speed. v Indicates the train's speed. dt Represents the time derivative. Indicates train i Arrival Station x The moment; In the new state, select the next action; Based on the chosen next action, update the Q value using the following formula: in, Indicates the state Execute action State-action value function Indicates site j The corresponding state, Indicates on the site j The chosen action Indicates the learning rate. Indicates an immediate reward. Indicates the discount factor. Indicates the next station j The state corresponding to +1 This indicates the action selected at the next stop. Indicates the next state Execute action Q-value, learning rate An adaptive decay strategy is adopted, gradually reducing the decay rate as the number of training epochs increases: in, Indicates the first e Learning rate during training rounds This represents the initial learning rate. Indicates the attenuation coefficient. e Indicates the current training round number; Repeat the above process until the train reaches the final station, completing one training round; offline training is completed when the number of training rounds reaches the maximum value or the Q value converges.

7. The autonomous decision-making method for urban rail train groups according to claim 6, characterized in that, The The expression for the strategy selection action is as follows: in, Indicates on the site j The optimal action to choose Represents a random number between 0 and 1. Indicates site j The corresponding state, This represents the learning rate.

8. The autonomous decision-making method for urban rail train groups according to claim 6, characterized in that, The constraints that the selection of states and the updating of the action space must satisfy are as follows: in, Indicates the train is at the station. j runtime Indicates site j Runtime adjustment action, Indicates the train is at the station. j Stop time, Showing the site j The stop time adjustment action, Indicates the minimum running interval. Indicates the maximum running interval. Indicates the minimum stopping time. Indicates the maximum dwell time. Indicates the train is at the station. j +1 running time, Indicates the train is at the station. j +1 stop time.