A train operation adjustment method and device based on reinforcement learning and electronic equipment
By training the reinforcement learning model with the DDPG algorithm and optimizing the train operation adjustment strategy, the problems of high computational complexity and large differences in schemes in existing methods were solved, and efficient train operation adjustment was achieved.
Patent Information
- Application Number
- CN202411404990.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-09
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2044-10-09
AI Technical Summary
The existing reinforcement learning-based train operation adjustment method has high computational complexity, making it difficult to make overall adjustments to the train operation plan, and the adjusted plan is significantly different from the original plan.
The deep deterministic policy gradient (DDPG) algorithm is used to train the reinforcement learning model. By establishing a set of train operation states, a set of actions, and a target reward function, the train operation adjustment strategy is optimized to ensure that the adjusted plan is not much different from the original plan.
The calculation complexity is reduced, the overall train delay time is shortened, and the adjusted plan is not much different from the original plan, avoiding large-scale adjustments.
Smart Images

Figure CN119117052B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of rail transportation, and in particular to a train operation adjustment method, device and electronic equipment based on reinforcement learning. Background Art
[0002] With the rapid development of high-speed railways, the scale of the high-speed rail network has expanded, and the network structure has become more complex. This, coupled with the impact of unexpected events such as severe weather and equipment failures, can lead to train delays, which in severe cases can have significant social impacts and economic losses. Train operation adjustment is a key component of railway dispatching and command. By dynamically adjusting train operation plans, it is possible to coordinate the resumption of orderly operation of various trains as quickly as possible, reducing delays and narrowing the scope of impact. Therefore, research on train operation adjustment is of great practical significance. The complexity of the train operating environment means that train operation adjustment is a large-scale, complex combinatorial optimization problem. Reinforcement learning can learn optimal policies through interaction with the environment. It can autonomously learn and gradually optimize in complex environments. It is suitable for decision-making problems in complex environments and is therefore often used in train operation adjustment. However, existing reinforcement learning-based train operation adjustment methods have disadvantages such as high computational complexity and difficulty in making comprehensive adjustments to train operation plans. Summary of the Invention
[0003] In view of this, the present application proposes a train operation adjustment method, device and electronic equipment based on reinforcement learning, which can make overall adjustments to the operation plans of multiple trains on the same train operation line, reduce the overall train delay time, and ensure that the adjusted train operation plan will not be too different from the original train operation plan, thereby avoiding large-scale adjustments to the operation plans of the train group.
[0004] According to one aspect of the present application, a train operation adjustment method based on reinforcement learning is provided, comprising: obtaining a first train planned operation plan for multiple trains on the same train operation line; the train operation line includes multiple stations; the first train planned operation plan includes the first planned arrival time of each train in the multiple trains at each station in the train operation line, the first planned stop time of each train at each station, and the first planned operation time of each train in each operation section in the train operation line; the operation section represents the section between two adjacent stations in the train operation line; establishing a train operation state set and a train operation action set; the train operation state set includes the delay value of each train at each station; the train operation action set includes the stop time of each train at each station and the travel speed of each train in each operation section; the delay value represents the difference between the time when each train arrives at each station and the time when each train plans to arrive at the The difference between the first planned arrival times of each station; configuring a target reward function according to the delay value of each train at each station, the stop time of each train at each station and the first planned stop time, the running time of each train in each operating interval and the first planned running time; the running time of each train in each operating interval is determined according to the distance of each operating interval and the travel speed of each train in each operating interval; according to the train operating state set, the train operating action set and the target reward function, training a reinforcement learning model based on the deep deterministic policy gradient DDPG algorithm to obtain a trained reinforcement learning model; using the trained reinforcement learning model, obtaining a second train planned operation plan; the second train planned operation plan includes the second planned arrival time of each train at each station, the second planned stop time of each train at each station and the second planned running time of each train in each operating interval.
[0005] In one possible implementation, if the first planned stop duration of the first train at the first station is not 0, then the stop duration of the first train at the first station in the train operation action set is also not 0, and the stop duration of the first train at the first station is greater than or equal to the minimum preset stop duration; wherein, the first train is any train among the multiple trains; and the first station is any station among the multiple stations.
[0006] In one possible implementation, the running speed of the first train in each running interval in the train running action set is less than or equal to the maximum preset running speed of the first train, and the running speed of the first train in the first running interval is less than or equal to the maximum allowable running speed of the first running interval; the first running interval is any running interval among the running intervals.
[0007] In one possible implementation, the target reward function includes a first reward function, a second reward function, and a third reward function; the value of the target reward function is obtained by weighted summing the values of the first reward function, the second reward function, and the third reward function; wherein the value of the first reward function is calculated based on the difference between the delay value of each train at the second station and the delay value of each train at the third station; the value of the second reward function is calculated based on the difference between the delay value of each train at the departure station among the multiple stations and the delay value of each train at the terminal station among the multiple stations; the value of the third reward function is calculated based on the difference between the first planned stop time of each train at the first station and the stop time of each train at the first station, and the difference between the first planned running time of each train in the first operating section and the running time of each train in the first operating section; the first station is any station among the multiple stations; the second station is any station among the multiple stations except the terminal station; the third station is the next station adjacent to the second station among the multiple stations; and the first operating section is any operating section among the multiple operating sections.
[0008] In a possible implementation, the reinforcement learning model comprises a current actor network, a target actor network, a current critic network and a target critic network; the train operation line comprises M+1 stations; and the training of the reinforcement learning model based on a deep deterministic policy gradient (DDPG) algorithm according to the set of train operation states, the set of train operation actions and the target return function comprises: inputting a kth train operation state into the current actor network to obtain an action difference value corresponding to the kth train operation state, wherein the kth train operation state is a set of late departure values of the trains at a kth station in the train operation line in the set of train operation states, the action difference value corresponding to the kth train operation state comprises a difference between a planned stop duration of the trains at the kth station and a first planned stop duration of the trains at the kth station, and a difference between a planned running duration of the trains in a k+1th running interval and a first planned running duration of the trains in the k+1th running interval, the k+1th running interval is a running interval between the kth station and a k+1th station in the train operation line, 0≤k≤M, obtaining an action corresponding to the kth train operation state according to the action difference value corresponding to the kth train operation state, wherein the action corresponding to the kth train operation state comprises a stop duration of the trains at the kth station and a running speed of the trains in the k+1th running interval selected from the set of train operation actions, inputting the kth train operation state and the action corresponding to the kth train operation state into the current critic network to obtain a current return value, wherein the current return value is calculated according to the target return function, updating parameters of the current actor network according to the current return value, and stopping updating the parameters of the current actor network until a first preset training condition is met to obtain a trained current actor network, inputting a k+1th train operation state into the target actor network to obtain an action difference value corresponding to the k+1th train operation state, obtaining an action corresponding to the k+1th train operation state according to the action difference value corresponding to the k+1th train operation state, inputting the k+1th train operation state and the action corresponding to the k+1th train operation state into the target critic network to obtain a target return value, wherein the target return value is calculated according to the target return function, updating parameters of the current critic network according to the current return value and the target return value, and stopping updating the parameters of the current critic network until a second preset training condition is met to obtain a trained current critic network.Using the parameters of the trained current actor network as the parameters of the target actor network, and using the parameters of the trained current critic network as the parameters of the target critic network, a trained target actor network and a trained target critic network are obtained; and obtaining a trained reinforcement learning model based on the trained current actor network, the trained current critic network, the trained target actor network, and the trained target critic network.
[0009] In one possible implementation, the current Actor network and the target Actor network include a hidden layer and an output layer; the k-th train running state is input into the current Actor network to obtain the action difference corresponding to the k-th train running state, including: inputting the k-th train running state into the hidden layer of the current Actor network and applying the linear rectification ReLU activation function to obtain a first eigenvalue; inputting the first eigenvalue into the output layer of the current Actor network and applying the hyperbolic tangent tanh activation function and then multiplying by a preset coefficient to obtain the difference between the running time of each train in the k+1-th running interval and the first planned running time of each train in the k+1-th running interval; inputting the first eigenvalue into the output layer of the current Actor network and applying the tanh activation function to obtain a second eigenvalue; dividing the second eigenvalue by 2 and applying the ReLU activation function to obtain the difference between the stop time of each train at the k-th station and the first planned running time of each train at the k+1-th station. The difference between the first planned stop time of the train at the k-th station; the inputting the k+1-th train running state into the target Actor network to obtain the action difference corresponding to the k+1-th train running state, including: inputting the k+1-th train running state into the hidden layer of the target Actor network and applying the ReLU activation function to obtain a third eigenvalue; inputting the third eigenvalue into the output layer of the target Actor network and applying the tanh activation function and then multiplying by the preset coefficient to obtain the difference between the running time of each train in the k+2-th running interval and the first planned running time of each train in the k+2-th running interval; inputting the third eigenvalue into the output layer of the target Actor network and applying the tanh activation function to obtain a fourth eigenvalue; dividing the fourth eigenvalue by 2 and then applying the ReLU activation function to obtain the difference between the stop time of each train at the k+1-th station and the first planned stop time of each train at the k+1-th station.
[0010] In one possible implementation, the use of the trained reinforcement learning model to obtain a second train planned operation plan includes: obtaining an initial delay set; the initial delay set includes the delay values of the stations that each train has arrived at on the train operation route; determining the second planned stop duration of each train at each station and the second planned operation duration of each train in each operation interval based on the initial delay set, the train operation action set and the trained reinforcement learning model; determining the second planned arrival time of each train at each station based on the initial delay set, the second planned stop duration of each train at each station and the second planned operation duration of each train in each operation interval.
[0011] According to another aspect of the present application, a train operation adjustment device based on reinforcement learning is provided, including: an acquisition module for acquiring a first train planned operation plan for multiple trains on the same train operation line; the train operation line includes multiple stations; the first train planned operation plan includes the first planned arrival time of each train among the multiple trains at each station in the train operation line, the first planned stop time of each train at each station, and the first planned operation time of each train in each operation section in the train operation line; the operation section represents the section between two adjacent stations in the train operation line; a state set and action set establishment module for establishing a train operation state set and a train operation action set; the train operation state set includes the delay value of each train at each station; the train operation action set includes the stop time of each train at each station and the travel speed of each train in each operation section; the delay value represents the difference between the time when each train arrives at each station and the planned arrival time of each train at each station the difference between the first planned arrival time and the first planned arrival time; a reward function configuration module, for configuring a target reward function according to the delay value of each train at each station, the stop time of each train at each station and the first planned stop time, the running time of each train in each operating interval and the first planned running time; the running time of each train in each operating interval is determined according to the distance of each operating interval and the travel speed of each train in each operating interval; a training module, for training a reinforcement learning model based on a deep deterministic policy gradient DDPG algorithm according to the train operating state set, the train operating action set and the target reward function to obtain a trained reinforcement learning model; an operation plan adjustment module, for obtaining a second train planned operation plan using the trained reinforcement learning model; the second train planned operation plan includes the second planned arrival time of each train at each station, the second planned stop time of each train at each station and the second planned running time of each train in each operating interval.
[0012] According to another aspect of the present application, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to implement the above-mentioned reinforcement learning-based train operation adjustment method when executing the instructions stored in the memory.
[0013] According to another aspect of the present application, a non-volatile computer-readable storage medium is provided, on which computer program instructions are stored, wherein the computer program instructions, when executed by a processor, implement the above-mentioned reinforcement learning-based train operation adjustment method.
[0014] According to another aspect of the present application, a computer program product is provided, comprising a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above-mentioned reinforcement learning-based train operation adjustment method.
[0015] The train operation adjustment method based on reinforcement learning of the present application establishes a state set according to the delay value of the train at each station, establishes an action set according to the stop time of the train at each station and the speed of the train in each operation section, sets a reward function according to the delay value of the train and the gap between the original train operation plan and the action adopted by the train in the action set, and trains the reinforcement learning model based on the DDPG algorithm. The adjusted train operation plan is obtained through the trained reinforcement learning model, which can make overall adjustments to the operation plans of train groups on the same train operation line, reduce the overall delay time of the train, and ensure that the adjusted train operation plan will not be too different from the original train operation plan, thereby avoiding large-scale adjustments to the operation plans of train groups.
[0016] Other features and aspects of the present application will become apparent from the following detailed description of exemplary embodiments with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate exemplary embodiments, features, and aspects of the application and, together with the description, serve to explain the principles of the application.
[0018] Figure 1 A schematic diagram showing the process of reinforcement learning.
[0019] Figure 2 A flowchart of a train operation adjustment method based on reinforcement learning according to an embodiment of the present application is shown.
[0020] Figure 3 A schematic diagram showing the input and output process of the neural network in the traditional DDPG algorithm.
[0021] Figure 4 A structural diagram of a train operation adjustment device based on reinforcement learning according to an embodiment of the present application is shown.
[0022] Figure 5 A block diagram of an electronic device 1900 according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0023] Various exemplary embodiments, features, and aspects of the present application will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise indicated.
[0024] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.
[0025] In addition, numerous specific details are provided in the detailed description below to better illustrate the present application. Those skilled in the art will appreciate that the present application can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art are not described in detail in order to highlight the main purpose of the present application.
[0026] In reinforcement learning, an agent interacts with its environment to acquire a state and then uses this state and the environment to determine its next action. After taking an action, the agent updates its state by interacting with the environment and receives rewards based on the changes in state and the actions taken. During learning, the agent repeats this process repeatedly, learning the appropriate decision-making method.
[0027] The train operation adjustment method based on reinforcement learning mainly achieves the purpose of continuous learning to optimize the selection strategy through the interaction between the intelligent agent and the environment. Figure 1 A schematic diagram of the reinforcement learning process is shown, as Figure 1 As shown in the figure, the agent selects a train operation adjustment plan based on the initial selection policy, makes adjustments based on this plan, interacts with the environment, and then receives a reward. Based on the reward, the agent updates its strategy selection method and then returns to the strategy selection step. Through step-by-step iterative optimization, a convergent strategy selection method (i.e., the optimal train operation adjustment plan) is eventually obtained, resulting in a trained reinforcement learning model.
[0028] A very important step in the reinforcement learning process is to establish a reinforcement learning environment for the agent to interact. The reinforcement learning environment generally consists of the following aspects: state set, action set, state transition matrix, and reward function.
[0029] The state set is used to express the basic description of the problem. In the train operation adjustment method based on reinforcement learning, the continuous time is usually discretized to model the train operation process as a Markov process. The state set is used to describe the specific operation of the train in such a discrete time sequence. Different ways of looking at the problem will lead to significant differences in the description of the state set. For example, some related research mainly adjusts the train arrival and departure time sequence in a certain station, so the arrival and departure times of L trains in s station are used as the state set to describe the operation state of the train. This state set description is only limited to the consideration of a certain station. For train modeling, this state set description is relatively accurate and convenient for train operation planning, but when the same state set description is applied to more stations and multiple stations are considered as a whole, the description space of the state set will be very large, so such a state set description greatly limits the possibility of adjusting the entire train operation network, making it difficult to adjust the entire train operation plan on the train operation line.
[0030] The action set is the action selected to jump from one state to the next state. The most important thing in the action set is the rationality of the selected action. Some related research uses the update scheme to select the action, which is biased towards search algorithms. Basically, the adjustment strategy (i.e., action) is selected by permutation and combination, and then the feasibility of the selected adjustment strategy is judged and filtered when interacting with the environment. However, such an action set is factorial in size, which will generate a huge search space, greatly increasing the computational complexity.
[0031] The state transition matrix is used to describe the jump of the state. When the state of the train is not at the terminal station, the train will jump to the next state after taking action a in s station. By defining the probability of jumping to state s+1 after taking action a in state s, the Markov process of the train before reaching the terminal station can be completely described.
[0032] The reward function needs to be set according to the actual needs. In related research, the mapping of the total late time of the train to the terminal station is usually taken as the reward function, such as taking the negative value of the total late time as the reward function. However, this reward function setting method only considers the influence of the total late time of the train on the entire operation line, without considering the influence of the late time of the train at each station on each station, and without considering the gap between the adjusted train operation scheme and the original train operation scheme, which may lead to a large difference between the adjusted train operation scheme and the original train operation scheme.
[0033] Existing reinforcement learning-based train operation adjustment methods have shortcomings such as high computational complexity, difficulty in making overall adjustments to the train operation plan, and the possibility of significant differences between the adjusted train operation plan and the original train operation plan. To address the above issues, this application proposes a reinforcement learning-based train operation adjustment method that optimizes the establishment of a state set, an action set, and the setting of a reward function. The corresponding state transition is determined by the previously established state set, action set, and reward function.
[0034] Train operation adjustment mainly changes the operation of the train by giving different train operation plans, and finally achieves the recovery of train delays or approaches the original train operation plan to a certain extent. The research object of the train operation adjustment problem should not be limited to the adjustment of a certain train, but the operation of a group of trains on a certain operation line, and the overall adjustment of the train group is made to restore the overall delay of these trains to the greatest extent. Therefore, this application describes the train operation adjustment problem as follows: for N trains on a train operation line, the original train operation plan is {P1, P2,…, PN}, and the initial train delay situation is {Delay1, Delay2,…, DelayN}, where Pi represents the original operation plan of train i on the entire operation line, and Delayi represents the delay situation of train i at the first few stations, 1≤i≤N. A reinforcement learning model needs to be trained using reinforcement learning methods. The reinforcement learning model can give a new train operation plan {P1', P2', ..., PN'} based on the original train operation plan and the initial delay situation, where Pi' represents the new operation plan of train i on the entire operation line; the new operation plan can restore the overall delay situation of the train group as much as possible compared to the original operation plan, and the new operation plan will not differ too much from the original operation plan.
[0035] In establishing the state set, the most important part to be considered in adjusting the train operation plan is the train delay situation, so the train delay situation can be used as the train operation state.
[0036] In establishing the action set, constraints can be set for the train adjustment strategy (i.e., action) based on the actual train operation situation to ensure the rationality of the action. The search subspace (i.e., action set) of the adjustment strategy can be constructed based on the constraints to be selected, thereby optimizing the selection of the adjustment strategy. The main constraints to be considered may include: (1) station capacity: the number of trains stopping at the station cannot exceed the maximum number of trains that can be accommodated; (2) departure time interval: due to the number of arrival and departure lines, trains traveling in the same direction at the station generally cannot depart at the same time; (3) train stop time: for trains stopping, sufficient time is required for passengers to get on and off, and for trains passing through, they cannot occupy the passing track at the same time; (4) necessary travel time: the maximum travel speed of the train and the distance between stations will limit the minimum time required for the train to travel between stations. This time cannot be adjusted to dynamically make up for the delay time.
[0037] In setting the reward function, the ultimate goal of train operation adjustment is to recover from train delays. Therefore, the reward function needs to be designed based on the train's delays. Furthermore, the train operation adjustment method of this application also aims to ensure that the resulting train operation adjustment plan does not differ significantly from the original plan. Therefore, the reward function can be calculated based on two factors: the train's delays and the difference between the train operation adjustment plan and the original plan. When setting the reward function based on train delays, considering the presence of multiple stations along the train route, two approaches can be considered for the reward function: a temporary reward function, calculating a reward value each time the train arrives at a station, providing immediate feedback; and an overall reward function, calculating a reward value upon reaching the terminal station, providing feedback on the total train delay. By setting these two different reward functions and the proportion of each reward function, the reinforcement learning model can balance its choice of the optimal adjustment plan: one that minimizes the total train delay and the other that minimizes the impact of all delayed trains on a station. At the same time, setting a reward function based on the gap between the train operation adjustment plan and the original train operation plan can avoid large-scale adjustments to the train group's operation plan.
[0038] Based on the above ideas, this application proposes a train operation adjustment method based on reinforcement learning, which can make overall adjustments to the operation plans of multiple trains on the same train operation line, reduce the computational complexity, reduce the overall train delay time, and ensure that the adjusted train operation plan will not be too different from the original train operation plan, avoiding large-scale adjustments to the operation plans of the train group.
[0039] Figure 2A flow chart of a train operation adjustment method based on reinforcement learning according to an embodiment of the present application is shown, which is used to adjust the operation plans of multiple trains on the same train operation line, such as Figure 2 As shown, the method may include:
[0040] S201. Obtain the first train planned operation plan of multiple trains on the same train running line; the train running line includes multiple stations; the first train planned operation plan includes the first planned arrival time of each train among the multiple trains at each station in the train running line, the first planned stop time of each train at each station and the first planned operation time of each train in each operation section in the train running line; the operation section represents the section between two adjacent stations in the train running line.
[0041] The first train planned operation plan is the original train planned operation plan. For example, the first train planned operation plan may include information such as the time at which each train in the train operation route was originally planned to arrive at each station (i.e., the first planned arrival time), the length of time each train was originally planned to stop at each station (i.e., the first planned stop length), the time each train was originally planned to depart from each station (which can be referred to as the first planned departure time), the length of time each train was originally planned to run in each operation interval (i.e., the first planned operation length), and the speed at which each train was originally planned to travel in each operation interval (which can be referred to as the first planned interval travel speed).
[0042] S202. Establish a train operation status set and a train operation action set; the train operation status set includes the delay value of each train at each station; the train operation action set includes the stop time of each train at each station and the travel speed of each train in each operation interval; the delay value represents the difference between the arrival time of each train at each station and the first planned arrival time of each train at each station.
[0043] Assume that there are N trains and M+1 stations on the train route, where the kth station represents the kth station that the train arrives at after departing from the starting station, 0≤k≤M, the 0th station represents the starting station, and the Mth station represents the terminal station.
[0044] For the establishment of the train operation state set (corresponding to the state set in the reinforcement learning environment), the train operation adjustment method of the embodiments of the present application considers the late adjustment of the train group on the same operation line, and therefore the defined train operation state set should be the set of late conditions of all trains on the train operation line. In actual situations, the prediction of train late is made station by station according to the stations on the train operation line, and when the train operation state set is established, the late conditions of each train at each station are also used as the train operation state.
[0045] Exemplarily, the train operation state set can be denoted as S = {S0, S2, …, S M}, S k represents the kth state, S k = {d1, d2, …, d N} ; wherein d i represents the late value of the ith train at the kth station (i.e. the difference between the time of the train arriving at the kth station and the first planned arrival time of the train at the kth station, if the train arrives at the kth station later than the first planned arrival time, the late value is positive; if the train arrives at the kth station earlier than the first planned arrival time, the late value is negative), 1≤i≤N.
[0046] Since the train operation state defined in the embodiments of the present application needs to consider the train late value of all trains at a station, the train operation state in the train operation state set is not limited. Exemplarily, an upper limit value of late and a lower limit value of late can be defined, so that all late values in the train operation state set are within the range of the lower limit value of late to the upper limit value of late, and the late value can be discretized, for example, the late value can be an integer minute.
[0047] For the establishment of the train operation action set (corresponding to the action set in the reinforcement learning environment), the embodiments of the present application set two actions of the train from the kth station to the k+1th station under the state S k , including the stop duration of the train at the kth station and the running speed of the train between the kth station and the k+1th station.
[0048] Exemplarily, the train operation action set can be denoted as Act = {Act1, Act2, …, Act M}, Act k+1 represents the action taken by the train under the kth state (i.e. the action taken by the train between the kth station and the k+1th station), Act k+1 = {V1, P1; V2, P2; …; V N , P N} ; wherein V iP represents the running speed of the ith train between the kth station and the k+1th station. i T represents the stop duration of the ith train at the kth station.
[0049] To ensure the rationality of the actions in the train operation action set, constraints can be set for the train operation action set. For the stop duration of each train at each station in the train operation action set, the stop duration can be constrained according to the stop behavior and passing behavior of each train at each station in the original train planning operation scheme. For the running speed of each train in each running section in the train operation action set, the running speed can be constrained according to the maximum running speed of each train and the speed limit of each running section.
[0050] For example, if the first planned stop duration of the first train at the first station is not 0, the stop duration of the first train at the first station in the train operation action set is also not 0, and the stop duration of the first train at the first station is greater than or equal to the minimum preset stop duration. The first train is any train in the multiple trains, and the first station is any station in the multiple stations.
[0051] For example, if the first planned stop duration of the ith train at the kth station is not 0 (i.e., the ith train plans to stop at the kth station in the original train planning operation scheme), the stop duration of the ith train at the kth station in the train operation action set is also not 0, and the stop duration is not less than the minimum preset stop duration P min , i.e., Act k+1 P i is not 0 and P i ≥ P min . P min can be set according to actual needs. If the first planned stop duration of the ith train at the kth station is 0 (i.e., the ith train plans to pass through the kth station in the original train planning operation scheme), the stop duration of the ith train at the kth station in the train operation action set can be 0 or not. In this way, if a train originally plans to stop at a station, the behavior of the train at the station in the given train operation action set is also stopping. If a train originally plans to pass through a station, the behavior of the train at the station in the given train operation action set can be passing or stopping. This can ensure that the original planned stop behavior of the train is not affected, thereby ensuring the rationality of the stop duration of the train at each station in the train operation action set.
[0052] Exemplarily, the running speed of the first train in each running section in the train operation action set is less than or equal to the maximum preset running speed of the first train, and the running speed of the first train in the first running section is less than or equal to the maximum allowed running speed of the first running section; the first running section is any one of the running sections.
[0053] As an example, if the maximum running speed of the i th train is V max1 , the running speed of the i th train in each running section in the train operation action set cannot be higher than V max1 , and if the maximum allowed running speed of the running section between the k th station and the k+1 th station is V max2 , the running speed of the i th train in the running section between the k th station and the k+1 th station in the train operation action set cannot be higher than V max2 , that is, V k+1 in Act i ≤ V max1 and V i ≤ V max2 . In this way, the rationality and feasibility of the running speed of the train in each running section in the train operation action set can be ensured.
[0054] In this way, the train operation action is constrained in the process of establishing the train operation action set, so that the rationality and feasibility of the selectable train operation action can be ensured, and the calculation complexity in the subsequent process of selecting the train operation action can be reduced.
[0055] S203, configuring a target reward function according to the delay value of each train at each station, the stop duration of each train at each station and the first planned stop duration, and the running duration of each train in each running section and the first planned running duration; the running duration of each train in each running section is determined according to the distance of each running section and the running speed of each train in each running section.
[0056] The reward function can be configured according to the delay value of the train and the difference between the original train planned operation scheme and the action adopted in the train operation action set.
[0057] The reward function configured based on a train's delay value must be determined based on the specific delay value. Since the delay value changes continuously as the train moves, the reward function based on the train's delay value can be calculated from two perspectives. First, when the train arrives at each station, a corresponding delay value is generated. A reward value is calculated based on the delay value obtained for each action taken by the train. This reward value is provided to the reinforcement learning model for each action taken, serving as an immediate reward. Second, a total delay value is generated after the train completes its route. This total delay value is used to calculate another reward value, which is provided to the reinforcement learning model after the train completes its route. By weightedly combining these two reward values, the reinforcement learning model can ultimately learn relevant decisions.
[0058] The reward function, configured based on the difference between the original train plan and the actions taken in the set of train operations, is fed into the reinforcement learning model after the train has completed its route. After the train has completed its route, the set of actions taken by the train constitutes the new train plan. Based on the difference between the new plan and the original plan (i.e., the extent of the adjustment to the plan for the entire train group), a reward value is calculated and fed into the reinforcement learning model, guiding it to learn the extent of the adjustment to the plan.
[0059] In a possible implementation, the target reward function may include a first reward function, a second reward function, and a third reward function.
[0060] Exemplarily, the value of the first reward function can be calculated based on the difference between the delay value of each train at the second station and the delay value of each train at the third station; the second station is any station among the multiple stations except the terminal station; the third station is the next station adjacent to the second station among the multiple stations.
[0061] As an example, assuming that there are N trains on the train route, the calculation formula of the first reward function R1 can be as follows:
[0062]
[0063] Among them, Delay k-1,i Delay represents the delay value of the i-th train in the k-1-th state (i.e., the delay value of the i-th train at the k-1-th station). k,i Delay represents the delay value of the i-th train in the k-th state (i.e., the delay value of the i-th train at the k-th station). k-1,i - Delay k,i) represents the reduction of the delay value of the i-th train at the k-th station relative to the delay value at the k-1-th station. If (Delay k-1,i - Delay k,i ) is a positive value, which means that the delay value of the i-th train at the k-th station is reduced relative to the delay value at the previous station. k-1,i - Delay k,i ) is negative, indicating that the delay of the i-th train at the k-th station has increased relative to the delay at the previous station. In the train operation adjustment problem defined in this application, we expect the delay of the train at each station to gradually decrease, so we expect the value of R1 to be positive and the larger the better.
[0064] In this way, when the train arrives at each station, the first reward function value is calculated by adding the reduction in the train's delay value at that station relative to the delay value at the previous station. Instant feedback can be given to the reinforcement learning model every time the train takes an action, so that the reinforcement learning model learns in the direction of gradually reducing the train's delay value at each station in the process of selecting actions.
[0065] For example, the value of the second reward function may be calculated based on the difference between the delay value of each train at the departure station among the plurality of stations and the delay value of each train at the terminal station among the plurality of stations.
[0066] As an example, assuming that there are N trains and M+1 stations on the train route, the calculation formula of the second reward function R2 can be as follows:
[0067]
[0068] Among them, Delay 0,i Delay represents the initial delay value of the i-th train at the departure station (i.e., the 0th station). M,i Delay represents the delay value of the i-th train at the terminal station (i.e. the M-th station). 0,i - Delay M,i ) represents the total delay value of the i-th train after running the entire route. If (Delay 0,i - Delay M,i ) is a positive value, it means that the total delay value of the i-th train on the entire running line is reduced. 0,i - Delay M,i ) is a negative value, indicating that the total delay value of the i-th train on the entire route increases. In the train operation adjustment problem defined in this application, we expect the total delay value of the train to decrease, so we expect the value of R2 to be positive and the larger the better.
[0069] In this way, after the train has traveled the entire route, the total delay value of the train is calculated as the second reward function value, which can enable the reinforcement learning model to learn in the direction of reducing the total delay value of the train in the process of selecting actions.
[0070] Exemplarily, the value of the third reward function can be calculated based on the difference between the first planned stop time of each train at the first station and the stop time of each train at the first station, and the difference between the first planned running time of each train in the first operating interval and the running time of each train in the first operating interval; the first station is any station among the multiple stations; the first operating interval is any operating interval among the operating intervals.
[0071] As an example, assuming that there are N trains on the train route, the calculation formula of the third reward function R3 can be as follows:
[0072]
[0073] Among them, Δ k,i,s represents the difference between the first planned stop time of the i-th train at the k-th station and the stop time of the i-th train at the k-th station in the set of actions taken after the i-th train has completed the entire route, Δ k,i,r It represents the difference between the first planned running time of the i-th train between the k-th station and the k+1-th station and the running time between the k-th station and the k+1-th station in the set of actions taken by the i-th train after running the entire running route. The running time between the k-th station and the k+1-th station can be calculated by dividing the distance between the k-th station and the k+1-th station by the speed taken by the i-th train between the k-th station and the k+1-th station. In the train operation adjustment problem defined in this application, we expect that the adjusted train operation plan will not be too different from the original train operation plan, so we expect the absolute value of R3 to be as small as possible.
[0074] For example, the value of the target reward function can be obtained by weighted summing the values of the first reward function, the second reward function, and the third reward function. The value R of the target reward function can be calculated according to the following formula:
[0075] R=a1·R1+a2·R2+a3·R3 (4)
[0076] Wherein, a1, a2, and a3 are positive numbers. a1, a2, and a3 can be set by those skilled in the art according to actual needs. For the train operation adjustment problem defined in this application, we expect that the value of R is as large as possible.
[0077] As an example, a1, a2, and a3 may all be 1, that is, R=R1+R2+R3.
[0078] In the configuration of the reward function in the embodiment of the present application, the values of the first reward function and the second reward function are calculated according to the train delay value. The reward values of these two parts can ensure that the reinforcement learning model learns in the direction of reducing the train delay value in the process of selecting actions, so that the adjusted train operation plan can achieve train delay recovery; the value of the third reward function is calculated according to the gap between the original train operation plan and the actions adopted in the train operation action set. On the one hand, this part of the reward value can ensure that the action selection strategy learned by the reinforcement learning model will not be adjusted significantly. On the other hand, it can avoid the adjusted train operation plan causing most trains to arrive at the station early.
[0079] S204. According to the train operation state set, the train operation action set and the target reward function, the reinforcement learning model is trained based on the deep deterministic policy gradient (DDPG) algorithm to obtain a trained reinforcement learning model.
[0080] The Deep Deterministic Policy Gradient (DDPG) algorithm is a hybrid algorithm based on the Deep Q-Network (DQN) and Policy Gradient (PG). Based on the basic ideas of the Q-learning algorithm in reinforcement learning, the DDPG algorithm uses a neural network structure instead of the Q-value table in the Q-learning algorithm to implement the two processes of selecting the optimization policy and updating the Q-value table after receiving rewards. The neural network structure in the DDPG algorithm consists of the following four neural networks: the current actor network (Actor network), the target actor network, the current value network (Critic network), and the target critic network. The current actor network and the target actor network have the same structure, with the input being the state variable and the output being the action selected under the input state. The current critic network and the target critic network have the same structure, with the input being the state variable and the action taken and the output being the Q-value (i.e., the reward value) obtained by taking the action under the state.
[0081] Figure 3 A schematic diagram of the input and output process of the neural network in the traditional DDPG algorithm is shown, as Figure 3As shown, the input of the current Actor network is the current state S, and the output is the action a selected to be taken in the current state S; the input of the target Actor network is the next state S', and the output is the action a' selected to be taken in the next state S'; this process replaces the process of selecting an action according to the relevant values in the Q value table in the Q-learning algorithm. The input of the current Critic network is the current state S and the action a, and the output is the current reward value Q(S, a) obtained after the action a is taken in the current state S; the input of the target Critic network is the next state S' and the action a', and the output is the target reward value Q'(S', a') obtained after the action a' is taken in the next state S'; this process replaces the calculation process of the Q value in the Q-learning algorithm.
[0082] The reinforcement learning method based on the DDPG algorithm needs to first use the current Actor network and the reinforcement learning environment to interact to generate a certain amount of samples, and then update the neural network parameters. In the entire reinforcement learning process, the generation of samples and the updating of parameters are the two most important processes.
[0083] For the generation method of the samples, the current state S is input into the current Actor network to obtain the action a selected in the current state S, the current state S and the action a are input into the reinforcement learning environment, and after interacting with the reinforcement learning environment, the next state S' and the corresponding reward value r after the action a is taken in the current state S can be obtained. The four variables are combined into a four-tuple (S, a, r, S') as a sample and put into the experience pool. The above process is repeatedly performed, different actions are taken in different states to obtain the corresponding next state variable and reward value, thereby completing the sample collection process. Subsequently, the neural network parameters can be updated by using the four-tuple samples generated above.
[0084] The neural network parameter update method first samples and updates the current actor network and the current critic network. After a period of time, the current actor network parameters are used to update the target actor network parameters (i.e., the current actor network parameters are used as the target actor network parameters). The current critic network parameters are then updated using the current critic network parameters (i.e., the current critic network parameters are used as the target critic network parameters). The current critic network parameter update process is as follows: a sample (S, a, r, S') is taken from the experience pool. The current state S and action a are input into the current critic network to obtain the current reward value Q(S, a). The next state S' is input into the target actor network to obtain the next action a'. The next state S' and the next action a' are input into the target critic network to obtain the target reward value Q'(S', a'). Then, based on Q(S, a) and Q'(S', a'), the current critic network parameters are updated using gradient descent to make the current critic network output Q(S, a) as close as possible to the label Q'(S', a'). The update process of the current Actor network parameters is as follows: use the current Actor network to obtain the action a under the current state S, use the current Crtic network to obtain the current reward value Q(S,a), and continuously update the parameters of the current Actor network according to Q(S,a) so that the output Q(S,a) can be maximized.
[0085] For the train operation adjustment problem defined in the present application, the state set in the reinforcement learning environment needs to consider the delay values of all trains at a station, so the train operation states in the state set are not limited. In contrast, the traditional Q-learning reinforcement learning method requires a limited state set, each state corresponds to a value in the Q-value table, and the corresponding action under state S is selected accordingly. In actual application scenarios, although the train delay values are discretized, there are still an infinite number of states, so a neural network is used to replace the Q-value table to describe and process various train delay states. Similarly, the train operation adjustment scheme (i.e., the length of time a train stops at each station and the speed of a train traveling in each operating section) is used as the action of reinforcement learning, which also has an infinite number of problems, so a neural network structure is also used to make decisions in the action selection process. Although in the DQN reinforcement learning method, a neural network is also used to replace the Q-value table, the neural network structure outputs the Q-value corresponding to the limited action set according to the input state S, so it is not applicable in the application scenario considered in the present application. In summary, the reinforcement learning application scenario in the present application is complex, and it is difficult to use some traditional reinforcement learning algorithms, but the DDPG algorithm can effectively describe and process the reinforcement learning application scenario in the present application, solving the problem of a large state set and action set established in the present application.
[0086] The process of training the reinforcement learning model of the embodiment of the present application based on the DDPG algorithm is introduced below.
[0087] In a possible implementation, the reinforcement learning model includes a current Actor network, a target Actor network, a current Critic network, and a target Critic network; the train operation line includes M+1 stations; and the training of the reinforcement learning model based on the deep deterministic policy gradient (DDPG) algorithm according to the train operation state set, the train operation action set, and the target return function to obtain the trained reinforcement learning model can include:
[0088] (1) Inputting the k-th train running state into the current Actor network, obtaining the action difference corresponding to the k-th train running state; the k-th train running state is the delay value of each train in the train running state set at the k-th station in the train running line; the action difference corresponding to the k-th train running state includes the difference between the stop time of each train at the k-th station and the first planned stop time of each train at the k-th station, and the difference between the running time of each train in the k+1-th running interval and the first planned running time of each train in the k+1-th running interval; the k+1-th running interval is the running interval between the k-th station and the k+1-th station in the train running line; 0≤k≤M.
[0089] (2) According to the action difference corresponding to the k-th train running state, the action corresponding to the k-th train running state is obtained; the action corresponding to the k-th train running state includes the stop time of each train selected from the train running action set at the k-th station and the running speed of each train in the k+1-th running section.
[0090] Assuming that the first planned stop duration of the i-th train at the k-th station obtained from the first train planned operation plan is p1, and the difference between the stop duration of the i-th train at the k-th station output by the current Actor network and the first planned stop duration of the i-th train at the k-th station is Δ1, then the stop duration of the i-th train at the k-th station selected from the train operation action set can be calculated as p2=p1+Δ1.
[0091] Assume that the running time of the i-th train in the k+1 running interval obtained from the first train plan is t1, the difference between the running time of the i-th train in the k+1 running interval output by the current Actor network and the first planned running time of the i-th train in the k+1 running interval is Δ2, and the distance of the k+1 running interval is s, then the running time of the i-th train in the k+1 running interval can be calculated as t2=t1+Δ2, and the speed v of the i-th train in the k+1 running interval selected from the train running action set is i =s / t2.
[0092] (3) Inputting the k-th train running state and the action corresponding to the k-th train running state into the current critic network to obtain a current reward value; the current reward value is calculated according to the target reward function.
[0093] The current reward value r can be calculated according to the above formula (4).
[0094] (4) updating the parameters of the current Actor network according to the current reward value until a first preset training condition is met, stopping updating the parameters of the current Actor network, and obtaining a trained current Actor network.
[0095] You can stop updating when the current reward value reaches the maximum value to obtain the trained current Actor network.
[0096] (5) Inputting the k+1th train running state into the target Actor network to obtain the action difference corresponding to the k+1th train running state.
[0097] The k+1th train operation status is the delay value of each train in the train operation status set at the k+1th station in the train operation line; the action difference corresponding to the k+1th train operation status includes the difference between the stop time of each train at the k+1th station and the first planned stop time of each train at the k+1th station, and the difference between the running time of each train in the k+2th operation section (that is, the running section between the k+1th station and the k+2th station in the train operation line) and the first planned running time of each train in the k+2th operation section.
[0098] (6) Obtaining the action corresponding to the k+1th train running state according to the action difference corresponding to the k+1th train running state.
[0099] The specific process of obtaining the action corresponding to the k+1th train running state according to the action difference corresponding to the k+1th train running state can refer to the above step (2) and will not be repeated here.
[0100] (7) Inputting the k+1th train operation state and the action corresponding to the k+1th train operation state into the target critic network to obtain a target reward value; the target reward value is calculated according to the target reward function.
[0101] The target return value r' can be calculated according to the above formula (4).
[0102] It should be noted that, since the values of R2 and R3 need to be calculated after the train has completed the entire route in the process of calculating the value of the target reward function, when calculating the reward value corresponding to the action taken under each train operating state, it is necessary to simulate the process of the train traveling the entire route starting from the initial operating state of the train at the departure station to obtain the corresponding reward value.
[0103] (8) Updating the parameters of the current critic network according to the current reward value and the target reward value until a second preset training condition is met, stopping updating the parameters of the current critic network, and obtaining a trained current critic network.
[0104] The gradient descent method can be used to update the parameters of the current Critic network. When the current reward value is closest to the target reward value, the update is stopped to obtain the trained current Critic network.
[0105] (9) The parameters of the trained current actor network are used as the parameters of the target actor network, and the parameters of the trained current critic network are used as the parameters of the target critic network, thereby obtaining a trained target actor network and a trained target critic network.
[0106] (10) Obtaining a trained reinforcement learning model based on the trained current actor network, the trained current critic network, the trained target actor network, and the trained target critic network.
[0107] The setting of the neural network in the reinforcement learning model of the embodiment of the present application is adjusted relative to the traditional DDPG algorithm. The output of the Actor network (including the current Actor network and the target Actor network) in the traditional DDPG algorithm is the action taken under state s, while in the reinforcement learning model of the embodiment of the present application, the output of the Actor network is adjusted to the difference between the action taken under state s and the action originally planned under state s, and the actual action taken under state s is subsequently calculated based on the action difference. The output of the Actor network can reflect the gap between the adjusted train plan operation plan and the original train plan operation plan, thereby making it easier to calculate the reward value corresponding to the adjustment range of the train operation plan.
[0108] For example, the current actor network and the target actor network can include a hidden layer and an output layer. There can be two hidden layers, each consisting of a linear layer and a linear rectification function (ReLU). The output layer can be one layer, using a hyperbolic tangent (tanh) activation function and a ReLU activation function.
[0109] Exemplarily, after inputting the running status of the kth train into the hidden layer of the current Actor network and applying the ReLU activation function, the first eigenvalue can be obtained; the first eigenvalue is input into the output layer of the current Actor network, and the tanh activation function is applied in the output layer and then multiplied by the preset coefficient, the running time of each train in the k+1th running interval and the first planned running time of each train in the k+1th running interval can be obtained; the first eigenvalue is input into the output layer of the current Actor network and the tanh activation function is applied in the output layer, the second eigenvalue can be obtained; the second eigenvalue is divided by 2 and then the ReLU activation function is applied (if there are N second eigenvalues, each second eigenvalue is divided by 2 and then the ReLU activation function is applied respectively), the difference between the stop time of each train at the kth station and the first planned stop time of each train at the kth station can be obtained.
[0110] Exemplarily, the operating status of the k+1th train is input into the hidden layer of the target Actor network and the ReLU activation function is applied to obtain the third eigenvalue; the third eigenvalue is input into the output layer of the target Actor network and the tanh activation function is applied and then multiplied by the preset coefficient to obtain the difference between the running time of each train in the k+2th running interval and the first planned running time of each train in the k+2th running interval; the third eigenvalue is input into the output layer of the target Actor network and the tanh activation function is applied to obtain the fourth eigenvalue; the fourth eigenvalue is divided by 2 and the ReLU activation function is applied (if there are N fourth eigenvalues, each fourth eigenvalue is divided by 2 and the ReLU activation function is applied respectively), and the difference between the stop time of each train at the k+1th station and the first planned stop time of each train at the k+1th station can be obtained.
[0111] In actual train operation, in order to allow passengers to get on and off the train, the train needs to have a basic stop time at the corresponding station. The originally planned stop time can be considered as the shortest stop time. In the reinforcement learning model of the embodiment of the present application, after the output layer of the Actor network applies the tanh activation function to the eigenvalues output by the hidden layer, half of the output eigenvalue is input into the ReLU activation function to obtain the difference between the stop time taken and the originally planned stop time. This part will not produce a negative output after passing through the ReLU activation function, which can ensure that the output action will not have an action in which the stop time taken is less than the originally planned stop time, thereby ensuring the rationality of the stop time of the train at each station in the adjusted train plan operation plan. In the reinforcement learning model of the embodiment of the present application, after the output layer of the Actor network applies the tanh activation function to the eigenvalues output by the hidden layer, it is multiplied by a preset coefficient as the difference between the interval operation time taken and the originally planned interval operation time. Through such processing, the adjusted train plan operation plan can be generated based on the original train plan operation plan, and the existing knowledge can be effectively used to generate the adjusted train plan operation plan.
[0112] For example, the current critic network and the target critic network can be composed of two linear layers and a ReLU activation function. The input of the current critic network and the target critic network can pass through two linear layers and then use a ReLU activation function to obtain the output result.
[0113] S205. Utilize the trained reinforcement learning model to obtain a second train planned operation plan; the second train planned operation plan includes the second planned arrival time of each train at each station, the second planned stop time of each train at each station, and the second planned operation time of each train in each operation section.
[0114] In one possible implementation, obtaining the second train operation plan by using the trained reinforcement learning model may include:
[0115] (1) Obtaining an initial delay set; the initial delay set includes the delay values of the stations that each train has arrived at on the train route.
[0116] If a train has arrived at n stations along the train route, the initial delay set may include the delay values of the train at the first n stations, that is, the initial train set may represent the running status of the first n trains.
[0117] Exemplarily, the initial delay set may include the delay value of each train at the departure station. In this case, the initial delay set may represent the initial train operation status.
[0118] (2) Determine the second planned stop duration of each train at each station and the second planned running duration of each train in each running section based on the initial delay set, the train running action set, and the trained reinforcement learning model.
[0119] The initial delay set and the train operation action set can be input into a trained reinforcement learning model. The trained reinforcement learning model has already learned the optimal selection strategy. Through the trained reinforcement learning model, the optimal stop duration of each train at subsequent stations other than the station it has already departed (i.e., the second planned stop duration) and the optimal speed of each train in each operating section it has not traveled (which can be called the second planned section speed) can be selected from the train operation action set. Based on the second planned section speed of the train in the operating section it has not traveled and the distance of the operating section, the second planned operating time of the train in the operating section can be calculated.
[0120] (3) Determine the second planned arrival time of each train at each station based on the initial delay set, the second planned stop time of each train at each station, and the second planned running time of each train in each running section.
[0121] Based on the delay value of each train at the arrived stations and the first planned arrival time, the actual arrival time of each train at these stations can be obtained; based on the actual arrival time of each train at these stations, the second planned stop time of each train at subsequent stations except the station it has departed, and the second planned running time of each subsequent running section, the departure time of each train from subsequent stations except the station it has departed (which can be called the second planned departure time) and the arrival time at subsequent stations (i.e., the second planned arrival time) can be obtained, so that the second train planned operation plan can be obtained.
[0122] The reinforcement learning-based train operation adjustment method of the embodiment of the present application establishes a state set according to the delay value of the train at each station, establishes an action set according to the stop time of the train at each station and the speed of the train in each operation section, sets a reward function according to the delay value of the train and the gap between the original train operation plan and the action adopted by the train in the action set, and trains the reinforcement learning model based on the DDPG algorithm. The adjusted train operation plan is obtained through the trained reinforcement learning model, and the operation plan of the train group on the same train operation line can be adjusted as a whole to reduce the overall delay time of the train, and can ensure that the adjusted train operation plan will not be too different from the original train operation plan, thereby avoiding large-scale adjustment of the operation plan of the train group.
[0123] In one embodiment, the parameters of the reinforcement learning model can be set as follows: the number of Actor network layers (i.e., network depth) can be 3 layers (2 hidden layers and 1 output layer), the number of neurons in the two hidden layers of the Actor network (i.e., network width) can be set to 256, the hidden layer of the Actor network uses the ReLU activation function, the output layer uses the tanh activation function and the ReLU activation function, and the output of the Actor network is a vector of the size of the number of trains * 2; the number of Critic network layers can be 2 layers (2 linear layers), the number of neurons in the two linear layers of the Critic network can be set to 256, the Critic network uses the ReLU activation function, and the output of the Critic network is the reward value; the learning rate can be set to 0.001, the future reward decay rate can be set to 0.99, the number of training samples per batch can be set to 64, and the maximum capacity of the experience pool samples can be set to 1000. The reinforcement learning model can be trained using train operation data from the high-speed railway line operating section from Guangzhou South Station to Chibi North Station on the Wuhan-Guangzhou High-Speed Railway. The train route includes 15 stations (Guangzhou South, Guangzhou North, Qingyuan, Yingde West, Shaoguan, Lechang East, Binzhou West, Leiyang West, Hengyang East, Hengshan West, Zhuzhou West, Changsha South Wuhan-Guangzhou, Miluo East, Yueyang East, and Chibi North) and 14 train operating sections (the distance traveled by each train operating section can be obtained by querying). Train operation data can include the train number, train departure date, the name of the station the train arrives at, the time the train arrives at each station, the time the train is scheduled to arrive at each station, the time the train departs from each station, the scheduled departure time of the train at each station, the track number when the train enters the station, and the delay time (i.e., the difference between the actual time the train arrives at the station and the scheduled arrival time). Table 1 shows some examples of train operation data.
[0124] Table 1
[0125]
[0126] Except for some larger stations on this route, it is generally less affected by cross-line trains. The simulation is a single line, so only trains departing from Guangzhou South Station to Chibi North Station are considered. Trains entering this line in the middle and trains leaving through a station in the middle are not considered. The operation plans of 30 trains at the first 7 stations are used as adjustment objects, and the training rounds are set to 50 rounds. The operation of the train is sampled 100 times in each round. The reinforcement learning model is trained and finally a trained reinforcement learning model and a train operation adjustment schedule are obtained. Table 2 shows some examples of the generated train operation adjustment schedule. Among them, Stop i Indicates the planned stop time of the train at the i-th station after adjustment, Run iIt represents the planned running time of the adjusted train in the i-th operating section (i.e. the operating section between the i-th station and the i+1-th station), in minutes.
[0127] Table 2
[0128]
[0129]
[0130] On the whole, the new train operation plan can play a role in delay recovery for the train's delay value. Some of the obviously unreasonable adjustment plans that appear therein (for example, the planned running time of train 6 in the seventh running interval in Table 2 is 3 minutes, and the running time is less than the minimum running time of the running interval) can be removed. After removing the unreasonable train operation plan, the total train operation recovery time is 109 minutes, and each train recovers an average of 4 minutes in the running line of 7 stations. In order to verify the adjustment effect of the new train operation plan on the actual operation of the train, the embodiment of the present application calculates and compares the actual running time of the new train operation plan with the actual running time according to the original train operation plan. Table 3 shows the delay recovery of the first 10 trains, in minutes.
[0131] Table 3
[0132] Train number Actual running time of the new plan Originally planned actual running time Late recovery time 1 42 67.5 25.5 2 46 67 21 3 56.5 72.5 16 4 72.5 71 -1.5 5 48.5 68.5 20 6 70.5 67 -3.5 7 73 76 3 8 78 73.5 -4.5 9 56 73.5 17.5 10 49.5 69 19.5
[0133] According to the comparison results of the actual running time of the new plan and the actual running time of the original plan, it can be seen that the reinforcement learning model trained in the embodiment of the present application has basic train delay adjustment capabilities, can reduce the overall train delay time, and can play a role in delay recovery for most trains.
[0134] The embodiment of the present application also proposes a train operation adjustment device based on reinforcement learning.
[0135] Figure 4 A structural diagram of a train operation adjustment device based on reinforcement learning according to an embodiment of the present application is shown as follows: Figure 4As shown, the apparatus can comprise: an acquisition module 401 configured to acquire a first train plan operation scheme of a plurality of trains on a same train operation line; the train operation line comprising a plurality of stations; the first train plan operation scheme comprising a first planned arrival time of each train of the plurality of trains at each station of the train operation line, a first planned stop duration of each train at each station, and a first planned operation duration of each train in each operation section of the train operation line; the operation section representing a section between two adjacent stations of the train operation line; a state set and action set establishment module 402 configured to establish a train operation state set and a train operation action set; the train operation state set comprising a late value of each train at each station; the train operation action set comprising a stop duration of each train at each station and a travel speed of each train in each operation section; the late value representing a difference between a time of each train arriving at each station and the first planned arrival time of each train at each station; a reward function configuration module 403 configured to configure a target reward function according to the late value of each train at each station, the stop duration of each train at each station and the first planned stop duration, and the operation duration of each train in each operation section and the first planned operation duration; the operation duration of each train in each operation section being determined according to a distance of each operation section and the travel speed of each train in each operation section; a training module 404 configured to train a reinforcement learning model based on a deep deterministic policy gradient (DDPG) algorithm according to the train operation state set, the train operation action set, and the target reward function, to obtain a trained reinforcement learning model; and an operation scheme adjustment module 405 configured to obtain a second train plan operation scheme by using the trained reinforcement learning model; the second train plan operation scheme comprising a second planned arrival time of each train at each station, a second planned stop duration of each train at each station, and a second planned operation duration of each train in each operation section.
[0136] In a possible implementation, if the first planned stop duration of the first train at the first station is not 0, the stop duration of the first train at the first station in the train operation action set is also not 0, and the stop duration of the first train at the first station is greater than or equal to a minimum preset stop duration; wherein the first train is any train of the plurality of trains, and the first station is any station of the plurality of stations.
[0137] In one possible implementation, the running speed of the first train in each running interval in the train running action set is less than or equal to the maximum preset running speed of the first train, and the running speed of the first train in the first running interval is less than or equal to the maximum allowable running speed of the first running interval; the first running interval is any running interval among the running intervals.
[0138] In one possible implementation, the target reward function includes a first reward function, a second reward function, and a third reward function; the value of the target reward function is obtained by weighted summing the values of the first reward function, the second reward function, and the third reward function; wherein the value of the first reward function is calculated based on the difference between the delay value of each train at the second station and the delay value of each train at the third station; the value of the second reward function is calculated based on the difference between the delay value of each train at the departure station among the multiple stations and the delay value of each train at the terminal station among the multiple stations; the value of the third reward function is calculated based on the difference between the first planned stop time of each train at the first station and the stop time of each train at the first station, and the difference between the first planned running time of each train in the first operating section and the running time of each train in the first operating section; the first station is any station among the multiple stations; the second station is any station among the multiple stations except the terminal station; the third station is the next station adjacent to the second station among the multiple stations; and the first operating section is any operating section among the multiple operating sections.
[0139] In a possible implementation, the reinforcement learning model comprises a current Actor network, a target Actor network, a current Critic network and a target Critic network; the train operation route comprises M+1 stations; the training module 404 is further configured to: input a kth train operation state into the current Actor network to obtain an action difference value corresponding to the kth train operation state; the kth train operation state is a late point value of the trains at a kth station in the train operation route in the set of train operation states; the action difference value corresponding to the kth train operation state comprises a difference between a stop duration of the trains at the kth station and a first planned stop duration of the trains at the kth station, and a difference between a running duration of the trains in a k+1th running interval and a first planned running duration of the trains in the k+1th running interval; the k+1th running interval is a running interval between the kth station and a k+1th station in the train operation route; 0≤k≤M; obtain an action corresponding to the kth train operation state according to the action difference value corresponding to the kth train operation state; the action corresponding to the kth train operation state comprises a stop duration of the trains at the kth station and a running speed of the trains in the k+1th running interval selected from the set of train operation actions; input the kth train operation state and the action corresponding to the kth train operation state into the current Critic network to obtain a current reward value; the current reward value is calculated according to the target reward function; update parameters of the current Actor network according to the current reward value, and stop updating the parameters of the current Actor network when a first preset training condition is met, to obtain a trained current Actor network; input a k+1th train operation state into the target Actor network to obtain an action difference value corresponding to the k+1th train operation state; obtain an action corresponding to the k+1th train operation state according to the action difference value corresponding to the k+1th train operation state; input the k+1th train operation state and the action corresponding to the k+1th train operation state into the target Critic network to obtain a target reward value; the target reward value is calculated according to the target reward function; update parameters of the current Critic network according to the current reward value and the target reward value, and stop updating the parameters of the current Critic network when a second preset training condition is met, to obtain a trained current Critic network; use the parameters of the trained current Actor network as the parameters of the target Actor network, and use the parameters of the trained current Critic network as the parameters of the target Critic network, to obtain a trained target Actor network and a trained target Critic network.Obtain a trained reinforcement learning model based on the trained current actor network, the trained current critic network, the trained target actor network, and the trained target critic network.
[0140] In one possible implementation, the current Actor network and the target Actor network include a hidden layer and an output layer; the training module 404 is further used to: input the running state of the k-th train into the hidden layer of the current Actor network and apply the linear rectification ReLU activation function to obtain a first eigenvalue; input the first eigenvalue into the output layer of the current Actor network and apply the tanh activation function and then multiply by a preset coefficient to obtain the difference between the running time of each train in the k+1-th running interval and the first planned running time of each train in the k+1-th running interval; input the first eigenvalue into the output layer of the current Actor network and apply the hyperbolic tangent tanh activation function to obtain a second eigenvalue; divide the second eigenvalue by 2 and apply the ReLU activation function to obtain the stop time of each train at the k-th station The operation status of the k+1 train is input into the hidden layer of the target Actor network and the ReLU activation function is applied to obtain a third eigenvalue; the third eigenvalue is input into the output layer of the target Actor network and the tanh activation function is applied and then multiplied by the preset coefficient to obtain the difference between the operation duration of each train in the k+2 operation interval and the first planned operation duration of each train in the k+2 operation interval; the third eigenvalue is input into the output layer of the target Actor network and the tanh activation function is applied to obtain a fourth eigenvalue; the fourth eigenvalue is divided by 2 and the ReLU activation function is applied to obtain the difference between the stop duration of each train at the k+1 station and the first planned stop duration of each train at the k+1 station.
[0141] In one possible implementation, the operation plan adjustment module 405 is further used to: obtain an initial delay set; the initial delay set includes the delay values of the stations that the trains have arrived at on the train operation route; determine the second planned stop duration of the trains at the stations and the second planned running duration of the trains in the operation intervals according to the initial delay set, the train operation action set and the trained reinforcement learning model; determine the second planned arrival time of the trains at the stations according to the initial delay set, the second planned stop duration of the trains at the stations and the second planned running duration of the trains in the operation intervals.
[0142] In some embodiments, the functions or modules included in the device provided in the embodiments of the present application can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.
[0143] The reinforcement learning-based train operation adjustment device of the embodiment of the present application can make overall adjustments to the operation plans of the train groups on the same train operation line, reduce the overall train delay time, and ensure that the adjusted train operation plan will not be too different from the original train operation plan, thereby avoiding large-scale adjustments to the operation plans of the train groups.
[0144] The present application also provides a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, implement the aforementioned reinforcement learning-based train operation adjustment method. The computer-readable storage medium may be a volatile or non-volatile computer-readable storage medium.
[0145] An embodiment of the present application also proposes an electronic device, comprising: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to implement the above-mentioned reinforcement learning-based train operation adjustment method when executing the instructions stored in the memory.
[0146] An embodiment of the present application also provides a computer program product, including a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above-mentioned reinforcement learning-based train operation adjustment method.
[0147] Figure 5 FIG1 shows a block diagram of an electronic device 1900 according to an embodiment of the present application. For example, the electronic device 1900 can be provided as a server or a terminal device. Figure 5 Electronic device 1900 includes a processing component 1922, which further includes one or more processors and memory resources represented by memory 1932 for storing instructions executable by processing component 1922, such as applications. The applications stored in memory 1932 may include one or more modules, each corresponding to a set of instructions. In addition, processing component 1922 is configured to execute instructions to perform the above-mentioned reinforcement learning-based train operation adjustment method.
[0148] The electronic device 1900 may further include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output interface 1958 (I / O interface). The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server 2003. TM , Mac OS X TM , Unix TM ,Linux TM , FreeBSD TM or similar.
[0149] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions, which can be executed by the processing component 1922 of the electronic device 1900 to complete the above-mentioned reinforcement learning-based train operation adjustment method.
[0150] The present application may be a system, method and / or computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present application.
[0151] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination thereof. As used herein, a computer-readable storage medium is not to be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through an electrical wire.
[0152] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.
[0153] The computer program instructions for performing the operation of the present application can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data or source code or object code written in any combination of one or more programming languages, wherein the programming language includes object-oriented programming languages such as Smalltalk, C++, and conventional procedural programming languages such as "C" language or similar programming languages. Computer-readable program instructions can be executed completely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or executed completely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer by any type of network including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (such as by using an Internet service provider to connect to the Internet). In certain embodiments, by utilizing the state information of computer-readable program instructions to personalize electronic circuits, such as programmable logic circuits, field programmable gate arrays (FPGAs) or programmable logic arrays (PLAs), the electronic circuits can execute computer-readable program instructions, thereby realizing various aspects of the present application.
[0154] Various aspects of the present application are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0155] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0156] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0157] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the system, method and computer program product according to multiple embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and the part of the module, program segment or instruction includes one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a special hardware-based system that performs the function or action of the specification, or can be implemented by a combination of special hardware and computer instructions.
[0158] While various embodiments of the present application have been described above, the above descriptions are illustrative, non-exhaustive, and not intended to be limiting of the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or technological improvements in the marketplace, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A train operation adjustment method based on reinforcement learning, characterized in that: include: Obtaining the first train planned operation plan of multiple trains on the same train operation line; The train operation route includes a plurality of stations; The first train operation plan includes a first planned arrival time of each train among the multiple trains at each station on the train operation route, a first planned stop time of each train at each station, and a first planned operation time of each train in each operation section on the train operation route; the operation section represents a section between two adjacent stations on the train operation route; Establishing a train operation state set and a train operation action set; the train operation state set includes the delay value of each train at each station; the train operation action set includes the stop time of each train at each station and the travel speed of each train in each operation section; the delay value represents the difference between the arrival time of each train at each station and the first planned arrival time of each train at each station; configuring a target reward function based on the delay value of each train at each station, the stop time of each train at each station and the first planned stop time, the running time of each train in each operating section and the first planned running time; the running time of each train in each operating section is determined based on the distance of each operating section and the travel speed of each train in each operating section; According to the train operation state set, the train operation action set and the target reward function, the reinforcement learning model is trained based on the deep deterministic policy gradient DDPG algorithm to obtain a trained reinforcement learning model; Using the trained reinforcement learning model, a second train planned operation plan is obtained; the second train planned operation plan includes the second planned arrival time of each train at each station, the second planned stop time of each train at each station, and the second planned operation time of each train in each operation section.
2. The method according to claim 1, characterized in that If the first planned stop duration of the first train at the first station is not 0, then the stop duration of the first train at the first station in the train operation action set is also not 0, and the stop duration of the first train at the first station is greater than or equal to the minimum preset stop duration; wherein, the first train is any train among the multiple trains; and the first station is any station among the multiple stations.
3. The method according to claim 2, characterized in that In the train operation action set, the running speed of the first train in each operation section is less than or equal to the maximum preset running speed of the first train, and the running speed of the first train in the first operation section is less than or equal to the maximum allowed running speed of the first operation section; The first operating interval is any one of the operating intervals.
4. The method according to claim 1, wherein The target reward function includes a first reward function, a second reward function, and a third reward function; the value of the target reward function is obtained by weighted summing the values of the first reward function, the second reward function, and the third reward function; The value of the first reward function is calculated based on the difference between the delay value of each train at the second station and the delay value of each train at the third station; The value of the second reward function is calculated based on the difference between the delay value of the originating station of each train among the multiple stations and the delay value of the terminal station of each train among the multiple stations; The value of the third reward function is calculated based on the difference between the first planned stop time of each train at the first station and the stop time of each train at the first station, and the difference between the first planned running time of each train in the first running section and the running time of each train in the first running section; The first station is any station among the multiple stations; the second station is any station among the multiple stations except the terminal station; the third station is the next station adjacent to the second station among the multiple stations; and the first operating interval is any operating interval among the operating intervals.
5. The method according to claim 4, characterized in that The reinforcement learning model includes a current actor network, a target actor network, a current value critic network, and a target critic network; the train route includes M+1 stations; The method further comprises training a reinforcement learning model based on a deep deterministic policy gradient (DDPG) algorithm according to the train operation state set, the train operation action set, and the target reward function to obtain a trained reinforcement learning model, including: Input the k-th train running state into the current Actor network to obtain the action difference corresponding to the k-th train running state; the k-th train running state is the delay value of each train in the train running state set at the k-th station in the train running line; the action difference corresponding to the k-th train running state includes the difference between the stop time of each train at the k-th station and the first planned stop time of each train at the k-th station, and the difference between the running time of each train in the k+1-th running section and the first planned running time of each train in the k+1-th running section; the k+1-th running section is the running section between the k-th station and the k+1-th station in the train running line; 0≤k≤M; Obtaining an action corresponding to the k-th train running state according to the action difference corresponding to the k-th train running state; the action corresponding to the k-th train running state includes the stop time of each train at the k-th station and the travel speed of each train in the k+1-th running section selected from the train running action set; Inputting the k-th train running state and the action corresponding to the k-th train running state into the current critic network to obtain a current reward value; the current reward value is calculated according to the target reward function; Updating the parameters of the current Actor network according to the current reward value until a first preset training condition is met, stopping updating the parameters of the current Actor network, and obtaining a trained current Actor network; Inputting the k+1th train running state into the target Actor network to obtain the action difference corresponding to the k+1th train running state; Obtaining an action corresponding to the k+1th train running state according to the action difference corresponding to the k+1th train running state; Inputting the k+1th train running state and the action corresponding to the k+1th train running state into the target critic network to obtain a target reward value; the target reward value is calculated according to the target reward function; Updating the parameters of the current critic network according to the current reward value and the target reward value until a second preset training condition is met, stopping updating the parameters of the current critic network, and obtaining a trained current critic network; Using the parameters of the trained current Actor network as the parameters of the target Actor network, and using the parameters of the trained current Critic network as the parameters of the target Critic network, to obtain a trained target Actor network and a trained target Critic network; A trained reinforcement learning model is obtained according to the trained current Actor network, the trained current Critic network, the trained target Actor network, and the trained target Critic network.
6. The method according to claim 5, characterized in that The current Actor network and the target Actor network include a hidden layer and an output layer; Inputting the k-th train running state into the current Actor network to obtain the action difference corresponding to the k-th train running state includes: Inputting the k-th train running state into the hidden layer of the current Actor network and applying a linear rectification ReLU activation function to obtain a first eigenvalue; Inputting the first eigenvalue into the output layer of the current Actor network and applying a hyperbolic tangent tanh activation function and then multiplying by a preset coefficient, obtains the difference between the running time of each train in the k+1th running section and the first planned running time of each train in the k+1th running section; Inputting the first eigenvalue into the output layer of the current Actor network and applying a tanh activation function to obtain a second eigenvalue; dividing the second eigenvalue by 2 and applying a ReLU activation function to obtain the difference between the stop duration of each train at the k-th station and the first planned stop duration of each train at the k-th station; Inputting the k+1th train running state into the target Actor network to obtain the action difference corresponding to the k+1th train running state includes: Inputting the k+1th train running state into the hidden layer of the target Actor network and applying a ReLU activation function to obtain a third eigenvalue; Inputting the third eigenvalue into the output layer of the target Actor network and applying a tanh activation function and then multiplying by the preset coefficient, obtains the difference between the running time of each train in the k+2th running section and the first planned running time of each train in the k+2th running section; The third eigenvalue is input into the output layer of the target Actor network and the tanh activation function is applied to obtain the fourth eigenvalue; the fourth eigenvalue is divided by 2 and the ReLU activation function is applied to obtain the difference between the stop time of each train at the k+1th station and the first planned stop time of each train at the k+1th station.
7. The method according to claim 5 or 6, characterized in that The method of obtaining a second train operation plan by using the trained reinforcement learning model includes: Obtaining an initial delay set; the initial delay set includes delay values of stations that each train has arrived at on the train route; Determining, based on the initial delay set, the train operation action set, and the trained reinforcement learning model, a second planned stop duration of each train at each station and a second planned operation duration of each train at each operation section; The second planned arrival time of each train at each station is determined based on the initial delay set, the second planned stop time of each train at each station and the second planned running time of each train in each running section.
8. A train operation adjustment device based on reinforcement learning, characterized in that: include: An acquisition module, configured to acquire a first train operation plan of multiple trains on the same train operation line; The train operation route includes a plurality of stations; The first train operation plan includes a first planned arrival time of each train in the plurality of trains at each station in the train operation route, a first planned stop time of each train at each station, and a first planned operation time of each train in each operation section in the train operation route; The operating section refers to the section between two adjacent stations in the train operating route; A state set and action set establishment module is configured to establish a train operation state set and a train operation action set; the train operation state set includes the delay value of each train at each station; the train operation action set includes the stop time of each train at each station and the travel speed of each train in each operation interval; the delay value represents the difference between the arrival time of each train at each station and the first planned arrival time of each train at each station; a reward function configuration module, configured to configure a target reward function based on the delay value of each train at each station, the stop duration of each train at each station and the first planned stop duration, the running duration of each train in each operating interval and the first planned running duration; the running duration of each train in each operating interval is determined based on the distance of each operating interval and the travel speed of each train in each operating interval; A training module is used to train the reinforcement learning model based on the deep deterministic policy gradient (DDPG) algorithm according to the train operation state set, the train operation action set, and the target reward function to obtain a trained reinforcement learning model; An operation plan adjustment module is used to use the trained reinforcement learning model to obtain a second train planned operation plan; the second train planned operation plan includes the second planned arrival time of each train at each station, the second planned stop time of each train at each station, and the second planned operation time of each train in each operation section.
9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to implement the method according to any one of claims 1 to 7 when executing the instructions stored in the memory.
10. A non-volatile computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Railway traffic system scheduling optimization method based on automaton and reinforcement learning
CN116001864A
Train adjustment method and system based on AI Agent, medium and equipment
CN118358626A