A real-time adjustment method for train operation based on cooperative competition game

By applying the deep reinforcement learning method of cooperative competition game in train operation scheduling, the real-time and optimization performance problems of train operation diagram adjustment under emergencies are solved, and efficient and real-time train operation diagram adjustment is achieved.

CN118722789BActive Publication Date: 2025-05-16BEIJING JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410629200.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-21
Publication Date
2025-05-16
Estimated Expiration
2044-05-21

AI Technical Summary

Technical Problem

It is difficult for the existing technology to adjust the train operation chart in real time in the event of emergencies, which makes it difficult for dispatchers to make precise optimization decisions quickly, and the existing methods have problems such as poor real-time performance and difficult to ensure overall optimization performance.

Method used

Deep reinforcement learning method based on cooperative competition game is adopted, by obtaining line static data and train dynamic data, real-time reward function and delay reward function are established, multi-agent game reinforcement learning is used to conduct competitive games when the train has not reached the terminal, and cooperative games are conducted after arrival, and strategy network is trained to achieve real-time adjustment.

Benefits of technology

Real-time adjustment of train operation charts under emergencies is achieved, overall optimal performance is improved, modeling problems caused by complex and changing train operation environments are solved, and resource utilization efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118722789B_ABST
    Figure CN118722789B_ABST
Patent Text Reader

Abstract

The present invention relates to a real-time adjustment method for train operation based on cooperative competitive game, the method comprising: based on the static data of the designated line, the dynamic data at the initial moment; the constraint conditions, establishing the immediate reward function and the delayed reward function in the process of training the target network; the immediate reward function is obtained based on the competitive game strategy and the delay time of the train itself when each train has not arrived at the terminal; the delayed reward function is obtained based on the cooperative game strategy and the total delay time of all trains after each train arrives at the terminal; training the updated policy network, and obtaining the loss function of the policy network in real time; when the loss function and the reward function meet the convergence conditions, obtaining the trained policy network, which is used to realize the real-time adjustment of the train operation diagram. The above method uses multi-train game deep reinforcement learning to continuously learn and interact with the environment, solving the problem of difficult modeling caused by the complex and changeable train operation environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to railway transportation dispatching technology, and in particular to a real-time train operation adjustment method based on cooperative competition game. Background Art

[0002] High-speed railways are characterized by high speed, large transport capacity, high safety, comfort and convenience. However, the strong coupling of train operation under networked operation conditions, the complex and changeable operating environment, and long-distance cross-line operation have brought great challenges to the safe and efficient operation of high-speed railways.

[0003] When an interference event occurs, the dispatcher handles the interference event based on manual experience, which makes it difficult to grasp the global train information in real time and quickly make accurate and optimized adjustment decisions. Using technologies such as big data, artificial intelligence, and large language models to intelligently optimize train operation adjustment plans has become a major demand for the intelligent development of high-speed rail. The train operation adjustment problem is NP hard, and the commonly used solution methods can be divided into three categories: operations research methods, simulation methods, and artificial intelligence methods. Since operations research methods require the establishment of accurate models, for optimization problems of complex systems such as high-speed rail, more assumptions and simplifications are usually made. In the solution process, there will be problems of poor feasibility and poor real-time performance, and simulation methods have problems such as large simulation scale and poor real-time performance.

[0004] In recent years, with the development of artificial intelligence technology, heuristic and deep reinforcement learning algorithms have shown significant advantages in complex information processing and control decision-making. This type of method integrates domain knowledge and constraints, and can quickly obtain an approximate global satisfactory solution, which is suitable for optimization problems with many constraints and imprecise models. However, the basic idea of ​​the heuristic algorithm is to regard the problem as a static process of single-stage decision-making, while the train and timetable adjustment problem is a multi-stage decision-making dynamic process. When solving such problems, there are problems such as poor real-time performance and difficulty in ensuring overall optimization performance. Reinforcement learning algorithms can solve multi-stage decision-making problems. By constantly interacting with the environment, corresponding strategies are obtained after training. They have strong real-time characteristics and provide a way to solve the adjustment of train timetables after emergencies. Therefore, it is of great significance to study the adjustment of train timetables based on deep reinforcement learning under emergencies. Summary of the invention

[0005] 1. Technical issues to be resolved

[0006] In view of the above-mentioned shortcomings and deficiencies of the prior art, the present invention provides a real-time adjustment method for train operation based on cooperative competition game, which solves the problem of adjusting the train operation diagram in the event of an emergency.

[0007] (II) Technical solution

[0008] In order to achieve the above object, the main technical solutions adopted by the present invention include:

[0009] In a first aspect, an embodiment of the present invention provides a method for real-time adjustment of train operation based on cooperative competition game, which includes:

[0010] S10, obtaining static data of a designated line and dynamic data of initial times of all trains associated with the designated line;

[0011] S20, based on predefined constraints and the initial state of each train in the specified line, establish an immediate reward function and a delayed reward function in the process of training the strategy network;

[0012] The instant reward function is obtained based on the competitive game strategy and the delay time of the train itself when each train fails to arrive at the terminal station;

[0013] The delay reward function is obtained after each train arrives at the terminal station based on the cooperative game strategy and the total delay time of all trains;

[0014] S30, when all trains have not arrived at the terminal station, based on the randomness and Nash equilibrium strategy of the interaction between the train and the reinforcement learning environment, it is determined whether the pre-defined constraints and the first condition satisfied by the immediate reward function are met, to select train actions, update the train status, and obtain an updated policy network;

[0015] S40, after all trains arrive at the terminal station, the updated policy network is trained according to the delayed reward function of the cooperative game strategy, and the loss function of the policy network is obtained in real time;

[0016] When the loss function and the reward function meet the convergence conditions, the trained policy network is obtained to realize real-time adjustment of the train operation diagram.

[0017] Optionally, the line static data includes one or more of the following: station capacity, planned operation schedule, minimum departure and arrival time interval, minimum stop time, and minimum interval operation time;

[0018] The train operation dynamic data includes one or more of the following: the current running position of the train, the remaining capacity of the preceding station, the arrival and departure delay time at the current time, and the train operation status;

[0019] The train operation status includes: stop, departure, arrival, and interval operation;

[0020] The dynamic data at the initial moment include: the train's initial running position is the departure station, the remaining capacity of the next station is the capacity of the next station, the initial arrival and departure delay time is 0, and the train's initial running status is stop.

[0021] Optionally, the predefined constraints include:

[0022] Departure time limit:

[0023] Minimum running time limit for interval:

[0024] Adjacent train interval limit: |a i,n -a i',n |≥ah n ,i'∈S M (1.3)

[0025] |d i,n -d i',n |≥d_h n , (1.4)

[0026] |d i,n -a i',n |≥da_h n , (1.5)

[0027] Station capacity restrictions:

[0028]

[0029] Crossing restrictions: (d i,n >d i′,n )⊙(a i,n+1 >a i′,n+1 )=1, (1.8)

[0030] Stop time limit:

[0031] Interval interruption time limit:

[0032]

[0033]

[0034] Among them, S N represents the set of stations, S M Represents a collection of trains; the last train departing from the originating station is numbered M, and the terminal station is numbered N; d i,n ,d i',n denote the departure time of train i and adjacent train i′ at station n, represents the scheduled departure time of train i at station n, a i,n+1 represents the arrival time of train i at the next station n+1, is the given minimum running time; a i,n 、a i',ndenote the arrival time of train i and adjacent train i' at station n, a_h n and d_h n are the minimum time intervals between the arrival and departure of train i and its adjacent train i' at station n, da_h n represents the departure and arrival interval constraint between train i and adjacent train i at the same station n; x t,i,n represents the train i at station n at time t; C n The capacity of each station; Indicates the minimum stop time; S NK Indicates the interrupt segment; H_dis_start and H_dis_end respectively indicate the start and end time of the interrupt.

[0035] Optionally, the status of each train in the specified line includes:

[0036] The state S of the train at time t t,i It is expressed as:

[0037] The state at the initial time t0 It is expressed as:

[0038] The action of train i at time t t,i Set to:

[0039]

[0040] Among them, the arrival time of the current train i is and departure time is n t,i Indicates the station where train i is located at the current time t or the last station passed or arrived at.

[0041] Optionally, the immediate reward function R t,i for:

[0042]

[0043] Delay reward function R t1,i for:

[0044]

[0045] Among them, F is the multiplication factor, H_dis_end-H_dis_start is the interval interruption time, r t,i It means that at time t, train i is at the departure station n t,i =0 is set to departure time Subtract the scheduled departure time a i,n , They represent the arrival time and planned arrival time of train i at station n respectively; the reward function including the immediate reward function and the delayed reward function is expressed as:

[0046]

[0047] T is the independent variable, and t1 represents the final time of the train operation diagram adjustment time interval.

[0048] Optionally, the first condition is: R t,i (action * t,i ,action * t,-i )≥R t,i (action t,i ,action * t,-i );

[0049] action * t,i represents the Nash equilibrium strategy of train i, action * t,-i It means that all trains except train i adopt Nash equilibrium strategy.

[0050] Optionally, in the first condition,

[0051] Assume action t-1,i =1 and the arrival condition and departure condition are met, when the random number rand≤ò, the action action is randomly selected t,i The value of makes train i continue to explore in the reinforcement learning environment; if the action selected t,i =1, indicating that train i passes station n at time t, n t,i =n t-1,i +1, when action t,i =0, it means that train i arrives at station n at time t;

[0052] Assume action t-1,i = 1 and the arrival condition is met but the departure condition is not met, let action t,i =0, indicating that train i arrives at station n at time t;

[0053] Assume action t-1,i = 1 and the arrival condition is not met, let action t,i =1, indicating that train i runs between station n-1 and station n at time t;

[0054] Assume action t-1,i=0 and meets the starting condition, then when the random number rand>ò, according to the strategy network Q(S t,i ,action t,i ,θ)Select action action t,i When the random number rand≤ò, randomly select action action t,i Value; selected action t,i = 1, indicating that train i departs from station n-1 at time t. t,i =0, it means that train i stops at station n-1 at time t;

[0055] Assume action t-1,i = 0 and the starting condition is not met, let action t,i =0, indicating that train i stops at station n-1 at time t;

[0056] action t,i After the selection is made, the train takes the selected action t,i , the immediate reward function and the delayed reward function will be obtained, and the next state will be inferred, and the policy network will be updated according to the immediate reward function and the delayed reward function.

[0057] Optionally, the S40 includes:

[0058] The updated policy network Q(S t,i ,action t,i ,θ) is:

[0059]

[0060] The loss function loss is:

[0061]

[0062] It means finding the average value, and rand is a random number;

[0063] Represents the target network that helps the policy network update the settings in the reinforcement learning environment. The architecture of the target network is the same as that of the policy network. The target network uses soft updates to update parameters. θ and θ′ represent the network parameters of the policy network to be trained and the network parameters of the target network; γ represents the discount factor, α represents the learning rate, and D represents the experience pool. Represents all the data in the experience pool (S t,i , action t,i , R t,i , S t+1,i ) collection; each time the train interacts with the reinforcement learning environment, it will get a set of data (St,i , action t,i , R t,i , S t+1,i ), put it into the experience pool to train the policy network.

[0064] In a second aspect, an embodiment of the present invention further provides a real-time train operation adjustment device based on cooperative competition game, comprising:

[0065] An acquisition unit, used for acquiring static line data of a designated line and dynamic data of initial time of all trains associated with the designated line;

[0066] A function building unit, used to build an immediate reward function and a delayed reward function in the process of training the strategy network based on predefined constraints and the initial state of each train in the specified line;

[0067] The instant reward function is obtained based on the competitive game strategy and the delay time of the train itself when each train fails to arrive at the terminal station;

[0068] The delay reward function is obtained after each train arrives at the terminal station based on the cooperative game strategy and the total delay time of all trains;

[0069] The policy network training unit is used to select train actions, update train states, and obtain updated policy networks based on the randomness and Nash equilibrium strategy of the interaction between the train and the reinforcement learning environment to determine whether the pre-defined constraints and the first condition satisfied by the immediate reward function are met when all trains have not arrived at the terminal station. After all trains arrive at the terminal station, the updated policy network is trained according to the delayed reward function of the cooperative game strategy, and the loss function of the policy network is obtained in real time.

[0070] When the loss function and the reward function meet the convergence conditions, the trained policy network is obtained to realize real-time adjustment of the train operation diagram.

[0071] In a third aspect, an embodiment of the present invention further provides a computing device, comprising: a processor, a memory, and a computer program stored in the memory, wherein the processor executes the computer program to implement a real-time adjustment method for train operation based on cooperative competitive game as described in any one of the first aspects.

[0072] (III) Beneficial effects

[0073] In the present invention, the advantages of multi-agent game deep reinforcement learning, such as simulating real environment, handling coordination and competition, distributed decision-making, learning complex strategies, flexibility and resource sharing, are integrated. Each train adjusts the train operation diagram according to the train plan operation diagram information and the location and time of temporary failure of the section's passing capacity, train operation rules, and multi-train system sharing information and experience to improve the overall optimal performance.

[0074] Using multi-train game deep reinforcement learning to continuously learn and interact with the environment makes the adjustment of the train timetable independent of the specific train model, solving the problem of difficult modeling caused by the complex and changeable train operation environment.

[0075] Use offline training to cope with complex and changing environmental changes and improve overall performance.

[0076] In the process of multi-train training, the competitive cooperation game strategy is used to enable each train to maximize its own interests while improving the overall optimal performance and making more effective use of line and station resources.

[0077] The learned train timetable adjustment strategy, i.e., the nonlinear relationship between the arrival and departure order of trains at stations, is saved in the strategy network of multi-train game deep reinforcement learning, generating a strategy for adjusting the train timetable to deal with emergencies. Subsequently, multiple trains can use the trained strategy to adjust the train timetable, and the quality of the strategy and the real-time solution can be guaranteed. BRIEF DESCRIPTION OF THE DRAWINGS

[0078] Figure 1 A flow chart of an integrated method for train speed curve optimization and operation diagram adjustment based on cooperative competition game provided in an embodiment of the present invention;

[0079] Figure 2 A train operation mode diagram based on cooperative competition game provided by an embodiment of the present invention;

[0080] Figure 3 An action selection diagram based on cooperative competition game provided by an embodiment of the present invention;

[0081] Figure 4 A structural diagram of a train operation diagram adjustment based on cooperative competition game provided by an embodiment of the present invention;

[0082] Figure 5 A train operation diagram generated based on a cooperative competition game method provided in an embodiment of the present invention;

[0083] Figure 6 A schematic diagram of a reward function provided by an embodiment of the present invention;

[0084] Figure 7 A schematic diagram of a loss function provided by an embodiment of the present invention;

[0085] Figure 8 A flowchart of a method for real-time adjustment of train operation based on cooperative competition game is provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0086] In order to better explain the present invention and facilitate understanding, the present invention is described in detail below through specific implementation modes in conjunction with the accompanying drawings.

[0087] When solving the problem of adjusting the high-speed train operation diagram with increasing traffic density and operating mileage, the action space and state space of the single-agent reinforcement learning algorithm increase exponentially, while multi-agent (train) reinforcement learning can integrate other agent information into the environment and only consider its own optimal action, thereby reducing the action and state space. Considering that the environment of multi-agent reinforcement learning is unstable, by using game theory to consider the competitive and cooperative relationship between trains, multi-agent game reinforcement learning is used to solve the problem of adjusting the train operation diagram.

[0088] In the process of using multi-agent game reinforcement learning to interact with the environment, since it takes a certain amount of time for a train to arrive at the terminal, the overall delay time cannot be obtained before all trains arrive at the terminal. Therefore, each agent is allowed to compete with each other, and each train strives to minimize its own delay time. After all trains arrive at the terminal, a cooperative game strategy is used to minimize the delay time of all trains, ultimately maximizing its own interests while improving the overall optimization performance. In this process, not only the delay time of each step and the total delay time reward function must be considered, but also the delay time of all agents. This patent invents a method of using multi-agent game reinforcement learning to comprehensively consider the agent's own reward function and overall optimization performance to achieve real-time adjustment of train operation.

[0089] In this embodiment, a train is used as an intelligent agent. Some embodiments use intelligent agents, and some embodiments use trains, which have the same meaning.

[0090] Embodiment 1

[0091] This embodiment provides a method for real-time adjustment of train operation based on cooperative competition game, and its execution subject can be any computing device, such as Figure 8 As shown, the method of this embodiment may include the following steps:

[0092] S10. Obtain static line data of a designated line and dynamic data of initial times of all trains associated with the designated line.

[0093] In this embodiment, the line static data includes one or more of the following: station capacity, planned operation schedule, minimum departure and arrival time interval, minimum stop time, and minimum interval operation time;

[0094] The train operation dynamic data includes the train operation dynamic data at the initial moment and the dynamic data obtained by interacting with the reinforcement learning environment at other moments. The dynamic data at the initial moment includes one or more of the following: the current running position of the train, the remaining capacity of the next station, the arrival and departure delay time at the current moment, and the train operation status; the train running position at the initial moment is the departure station, the remaining capacity of the next station is the capacity of the next station, the arrival and departure delay time at the initial moment is 0, and the train operation status at the initial moment is stop, etc.

[0095] S20, based on predefined constraints and the initial status of each train in the specified line, establish an immediate reward function and a delayed reward function in the process of training the strategy network.

[0096] The status of each train at the initial moment may include: the line static data in the above S10, the train operation dynamic data at the initial moment, the dynamic data at the initial moment, etc.

[0097] The instant reward function is obtained based on the competitive game strategy and the delay time of the train itself when each train fails to arrive at the terminal station;

[0098] The delay reward function is obtained after each train arrives at the terminal station based on the cooperative game strategy and the total delay time of all trains.

[0099] S30. When all trains have not arrived at the terminal, based on the randomness and Nash equilibrium strategy of the interaction between the train and the reinforcement learning environment, it is determined whether the pre-defined constraints and the first condition satisfied by the immediate reward function are met to select train actions, update the train status, and obtain an updated policy network.

[0100] It should be noted that, in this embodiment, only the states at the initial moment are read / read, and the states at other time periods except the initial moment are interactively updated with the reinforcement learning environment. In addition, the constraints and the first condition in the above S20 are used in the process of selecting actions, updating states, etc. of the train to obtain the immediate reward function, thereby obtaining an updated policy network.

[0101] S40. After all trains arrive at the terminal station, the updated policy network is trained according to the delayed reward function of the cooperative game strategy, and the loss function of the policy network is obtained in real time; when the loss function and the reward function (including the immediate reward function and the delayed reward function) meet the convergence conditions, the trained policy network is obtained to realize real-time adjustment of the train operation diagram.

[0102] The use of multi-agent game deep reinforcement learning to continuously learn and interact with the environment makes the adjustment of the train timetable independent of the specific train model, solving the problem of difficult modeling caused by the complex and changeable train operation environment. In the multi-train training process, the competitive and cooperative game strategy is used to enable each agent to maximize its own interests while improving the overall optimal performance, thereby making more effective use of line and station resources.

[0103] Embodiment 2

[0104] When the actual operation time of a high-speed train deviates from the planned operation diagram due to emergencies such as climate, disasters, and equipment failures, the train operation needs to be adjusted. Train operation adjustment optimization means dynamically adjusting the train operation diagram according to the actual operation status of the train and the network when the train operation is delayed, so that the delayed train can resume operation as soon as possible or reduce the delay time as much as possible, thereby reducing the number of delayed trains and the affected scope.

[0105] With the increase of traffic density and mileage, single train reinforcement learning can no longer meet the requirements. Since the state and action space increase exponentially with the number of trains and stations, it is necessary to use multi-train game reinforcement learning methods to solve the train schedule adjustment problem. In this embodiment, each train is an intelligent agent.

[0106] To better understand the method of this embodiment, Figures 1 to 7 , and the formulas involved are explained in detail.

[0107] The time range for train operation diagram adjustment is t∈(t0,t1), and the time is discretized into Time={t0,t0+1,t0+2,...,t1}, where t0 and t1 represent the start time and the end time of the operation diagram adjustment time interval, respectively, and the current time is t∈Time.

[0108] The method of this embodiment may specifically include the following steps:

[0109] Step 1: Obtain the static data of the designated line and the dynamic data of the initial time of all trains associated with the designated line.

[0110] Line static data includes: station capacity, planned operation schedule, minimum departure and arrival time interval, minimum stop time, minimum interval operation time, etc.

[0111] The dynamic data of train operation at the initial moment include: the current running position of the train, the remaining capacity of the next station, the arrival and departure delay time at the current moment, the running status of the train (stop, departure, arrival, section operation), the initial running position of the train is the departure station, the remaining capacity of the next station is the capacity of the next station, the arrival and departure delay time at the initial moment is 0, the running status of the train at the initial moment is stop, etc.

[0112] For example, manual surveys can be used to obtain static route data and dynamic data at the initial moment of the reinforcement learning environment.

[0113] Step 2: Based on the pre-defined constraints and the initial state of each train in the specified line, establish the immediate reward function and delayed reward function in the process of training the strategy network.

[0114] The instant reward function is obtained based on the competitive game strategy and the delay time of the train itself when each train fails to arrive at the terminal station;

[0115] The delay reward function is obtained after each train arrives at the terminal station based on the cooperative game strategy and the total delay time of the train.

[0116] Step 2-1: Establish the arrival and departure conditions of the constraints.

[0117] S N represents the set of stations, S M Represents a set of trains. The last train departing from the originating station is numbered M, and the last train departing from the destination station is numbered N.

[0118] ① Departure time limit: the departure time d of train i at station n i,n No earlier than the scheduled departure time

[0119]

[0120] ②Minimum running time limit for the interval: the arrival time a of train i at the next station n+1 i,n+1 Must be greater than or equal to the departure time d at the current departure station n i,n Plus minimum run time (Minimum running time for a given parameter):

[0121]

[0122] ③ Adjacent train interval limit: a_h n and d_h n are the minimum time intervals between the arrival and departure of train i and its adjacent train i' at station n, da_h n Represents the departure and arrival interval constraint of two trains i and i' at the same station n:

[0123] |a i,n -a i',n |≥a_h n ,i∈S M ,i'∈S M ,n∈S N, (1.3)

[0124] |d i,n -d i',n |≥d_h n ,i∈S M ,i'∈S M ,n∈S N , (1.4)

[0125] |d i,n -a i',n |≥da_h n ,i∈S M ,i'∈S M ,n∈S N , (1.5)

[0126] ④Station capacity limit: x t,i , n represents the train i at station n at time t. The number of trains at each station at the same time should be less than the capacity C of each station. n :

[0127]

[0128]

[0129] ⑤ Overtaking restriction: The symbol ⊙ represents the same or, the departure time sequence of all trains at station n is consistent with the arrival time sequence of the next station n+1:

[0130] (d i,n >d i′,n )⊙(a i,n+1 >a i′,n+1 )=1 (1.8)

[0131] a i,n+1 represents the arrival time of train i at the next station n+1, and i′ represents the trains other than i.

[0132] ⑥ Stop time limit: the departure time d of train i at station n i,n Should be greater than or equal to the arrival time a i,n Plus minimum stop time

[0133]

[0134] ⑦Interval interruption time limit: The train runs according to the original schedule before the interruption, which means that the train cannot enter the interruption section during the interruption period. NK , H_dis_start, H_dis_end represent the start and end time of the interruption respectively:

[0135]

[0136]

[0137] Arrival conditions: meet the minimum running time limit of the section (Formula 1.2); meet the overtaking limit (Formula 1.8); meet the arrival interval limit between trains (Formula 1.3); meet the section interruption time limit (Formula 1.10, Formula 1.11).

[0138] Departure conditions: not at the terminal station; meet the departure time limit (Formula 1.1); meet the stop time limit (Formula 1.9); meet the train departure interval constraint (Formula 1.4); meet the departure-arrival interval constraint between trains (Formula 1.5); meet the station capacity constraint (Formula 1.6, Formula 1.7); meet the interval interruption time limit (Formula 1.10, Formula 1.11).

[0139] Step 2-2: Set the state of each agent (i.e., the train state) to the arrival time and departure time of the current train at the next station at the current moment.

[0140] The train state is set to the arrival time of the current train i and departure time Therefore, it does not grow exponentially with the number of trains. t,i It indicates the station where train i is located at the current time t or the last station it passed or arrived at, such as Figure 2 The train operation mode diagram shown in the figure shows that the state S of train i at time t t,i It is expressed as:

[0141] The state at the initial time t0 It is expressed as:

[0142] Step 2-3: The action of each train is whether the current train is running at the current time.

[0143] The action of train i at time t t,i Set to:

[0144]

[0145] Step 2-4: Establish a reward function for training the policy network iterations.

[0146] Taking the train arrival delay time and departure delay time as the optimization target, a reward function is established. When the train does not arrive at the terminal station, the arrival and departure delay time of the train at the current station is used as the immediate reward function, so that each train i can compete in the game, and each train in each game stage adopts the Nash equilibrium strategy, which improves the convergence speed of multi-train reinforcement learning training and ensures that the strategy network can converge;

[0147] When all trains arrive at the terminal station, the cooperative game strategy is used, that is, the overall delay time is used as the delay reward function to train the reinforcement learning strategy network to improve the overall timetable adjustment performance; the arrival time and departure time of the current train at the previous station at the current moment are used as the state; and whether the current train is running at the current moment is used as the action.

[0148] When the train does not reach the terminal, a competitive game strategy is used to set the immediate reward function R according to the delay time of each train i. t,i for:

[0149]

[0150] Action * t,i represents the Nash equilibrium strategy of train i, action * t,-i Indicates that all trains except train i adopt the Nash equilibrium strategy, positivereward = 1, in order to prevent the immediate reward function R t,i The range is too large to converge, so normalization is performed. F represents the multiplication factor, and H_dis_end-H_dis_start represents the interval interruption time. t,i At time t, train i is at the departure station n. t,i =0 is set to departure time Subtract the scheduled departure time At the middle station t,i =1,...,N-1 is set as the departure delay time (departure time Subtract the scheduled departure time ) plus the arrival delay time (departure time Subtract the scheduled departure time );At the final destination n t,i =N is set as departure time Subtract the scheduled departure time

[0151]

[0152] After all trains arrive at the terminal, a cooperative game strategy is used to improve the overall train coordination performance. The delay reward function is set to the total delay time after all trains arrive at the last station, to avoid each agent only requiring its own delay time to be small and ignoring the overall performance:

[0153]

[0154] The reward function R mentioned below T,i Both are R including immediate reward function and delayed reward function T,i , R T,i It is expressed as: T is an intermediate variable.

[0155] Step 3: When all trains have not arrived at the terminal, based on the randomness and Nash equilibrium strategy of the interaction between the train and the reinforcement learning environment, it is determined whether the pre-defined constraints and the first condition satisfied by the immediate reward function are met to select the train action, update the train status, and obtain the updated policy network.

[0156] Each train i and the reinforcement learning environment continuously interact and learn, maximizing its own rewards while improving the overall coordination performance of the train operation system and minimizing delay time.

[0157] Each agent continuously interacts with the environment. For train i at the current time t, the train action selection is the same as the action action at the previous time t-1 t-1,i The value of is related to the arrival and departure constraints. Generate a random number rand for random exploration of train activity selection, set Assume that the last station reached or passed at time t-1 is n t-1,i =n-1; here ∈ is used to determine whether to select actions randomly or using a policy network

[0158] Assume action t-1,i =1 and meets the arrival condition and departure condition, then when the random number rand>∈, according to the strategy network Q(S t,i ,action t,i ,θ)Select action action t,i = argmax A Q(S t,i ,A,θ) values ​​are selected so that the strategy selected by each train in each stage is based on the Nash equilibrium strategy, so that the immediate reward function of each train i satisfies:

[0159] R t,i (action * t,i ,action * t,-i )≥Rt,i (action t,i ,action * t,-i ) (1.18)

[0160] When the random number rand≤∈, randomly select action action t,i The value of , to make the train continue to explore. If the action selected t,i =1, indicating that train i passes station n at time t, n t,i =n t-1,i +1, when action t,i =0, it means that train i arrives at station n at time t;

[0161] Assume action t-1,i = 1 and the arrival condition is met but the departure condition is not met, let action t,i =0, indicating that train i arrives at station n at time t;

[0162] Assume action t-1,i = 1 and the arrival condition is not met, let action t,i =1, indicating that train i runs between station n-1 and station n at time t.

[0163] Assume action t-1,i =0 and meets the starting condition, then when the random number rand>ò, according to the strategy network Q(S t,i ,action t,i ,θ)Select action action t,i When the random number rand≤ò, randomly select action action t,i Value. Selected action t,i = 1, indicating that train i departs from station n-1 at time t. t,i When =0, it means that train i stops at station n-1 at time t.

[0164] Assume action t-1,i = 0 and the starting condition is not met, let action t,i =0, indicating that train i stops at station n-1 at time t. Figure 3 .

[0165] Step 4: After all trains arrive at the terminal, the updated policy network is trained according to the delayed reward function of the cooperative game strategy, and the loss function of the policy network is obtained in real time; when the loss function and the reward function meet the convergence conditions, the trained policy network is obtained to realize real-time adjustment of the train operation diagram.

[0166] Whether the reward function and loss function in the multi-agent game reinforcement learning environment converge. If they converge, it means that the strategy network training is completed. If they do not converge, repeat steps 3 and 4 to continue training until convergence.

[0167] action t,i After the selection is completed, train i takes the selected action action t,i , we will get the reward function (1.15, 1.17) and infer the next state, and put (S t,i ,action t,i ,R t,i ,S t+1,i ) is stored in the experience pool, and the policy network (i.e. the pre-given neural network) is updated according to formula (1.19):

[0168]

[0169] α represents the learning rate, γ represents the discount factor, represents the target network, θ′←τθ+(1-τ)θ′, τ represents the update rate of the target network, and θ and θ' represent the network parameters of the policy network and the target network respectively. In the reinforcement learning environment, in order to make the update process of the policy network more stable, the target network is usually used. The structure of the policy network and the target network is the same, and the network parameters of the target network are obtained by soft updating the network parameters trained in the policy network.

[0170] The reward function of each round is the reward function R obtained by all interactions with the environment T,i The average value of ; and the loss function represents the error of the policy network, the formula is:

[0171]

[0172] Indicates the average value. The structure diagram of multi-agent game reinforcement learning is shown in Figure 4 .

[0173] The convergence of the reward function and the loss function indicates that the reinforcement learning training is completed, which means that it can be used to realize real-time train timetable adjustment. Otherwise, training needs to continue.

[0174] In the process of multi-train training, the competitive cooperation game strategy is used to enable each train to maximize its own interests while improving the overall optimal performance and making more effective use of line and station resources.

[0175] The learned train timetable adjustment strategy, i.e., the nonlinear relationship between the arrival and departure order of trains at stations, is saved in the strategy network of multi-train game deep reinforcement learning, generating a strategy for adjusting the train timetable to deal with emergencies. Subsequently, multiple trains can use the trained strategy to adjust the train timetable, and the quality of the strategy and the real-time solution can be guaranteed.

[0176] Experimental example

[0177] 24 trains running from Beijing to Dezhou East on the Beijing-Shanghai line are used. The five stations are represented as station 0 to station 4. The interruption interval is set to station 2 to station 3. The interruption event duration is 60 minutes. H_dis_start = 400min, H_dis_end = 460min, C n =[12,4,6,6,7]. The parameters in reinforcement learning are set as follows: learning rate, discount factor, update rate, batch size and training rounds are set to 0.0001, 0.9, 0.005, 512 and 600 respectively, the experience pool capacity is 50000, F=24, end =0.05,ò start =0.9, ò_decay=1000, the initial value of step_done is set to 1, and the policy network is used to select an action step_done=step_done+1. The adjusted train operation diagram is shown in Figure 5 In the example, the reward function and loss function are Figure 6 and Figure 7 In the middle, they are all convergent.

[0178] Embodiment 3

[0179] The embodiment of the present invention further provides a real-time train operation adjustment device based on cooperative competition game, comprising:

[0180] An acquisition unit, used for acquiring static line data of a designated line and dynamic data of initial time of all trains associated with the designated line;

[0181] A function building unit, used to build an immediate reward function and a delayed reward function in the process of training the strategy network based on predefined constraints and the initial state of each train in the specified line;

[0182] The instant reward function is obtained based on the competitive game strategy and the delay time of the train itself when each train fails to arrive at the terminal station;

[0183] The delay reward function is obtained after each train arrives at the terminal station based on the cooperative game strategy and the total delay time of all trains;

[0184] The policy network training unit is used to select train actions, update train states, and obtain updated policy networks based on the randomness and Nash equilibrium strategy of the interaction between the train and the reinforcement learning environment to determine whether the pre-defined constraints and the first condition satisfied by the immediate reward function are met when all trains have not arrived at the terminal station. After all trains arrive at the terminal station, the updated policy network is trained according to the delayed reward function of the cooperative game strategy, and the loss function of the policy network is obtained in real time.

[0185] When the loss function and the reward function meet the convergence conditions, the trained policy network is obtained to realize real-time adjustment of the train operation diagram.

[0186] Each step of the above method may correspond to the specific function of each unit in the device, please refer to the above records, and will not be described in detail here.

[0187] In addition, an embodiment of the present invention also provides a computing device, which includes: a processor, a memory, and a computer program stored in the memory, wherein the processor executes the computer program to implement the real-time adjustment method of train operation based on cooperative competition game as described in the above embodiments.

[0188] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0189] The present invention is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present invention. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions.

[0190] It should be noted that in the claims, any reference numerals placed between brackets shall not be construed as limiting the claims. The word "comprising" does not exclude the presence of components or steps not listed in the claims. The word "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The invention may be implemented by means of hardware comprising several different components and by means of a suitably programmed computer. In the claims enumerating several means, several of these means may be embodied by the same hardware. The use of the words first, second, third, etc., is for convenience of expression only and does not indicate any order. These words may be understood as part of the component name.

[0191] In addition, it should be noted that, in the description of this specification, the description of the terms "one embodiment", "some embodiments", "embodiment", "example", "specific example" or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, unless they are contradictory.

[0192] Although the preferred embodiments of the present invention have been described, those skilled in the art may make other changes and modifications to these embodiments after knowing the basic creative concept. Therefore, the claims should be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present invention.

[0193] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention should also include these modifications and variations.

Claims

1. A method for real-time adjustment of train operation based on cooperative competition game, characterized in that: include: S10, obtaining static data of a designated line and dynamic data of initial times of all trains associated with the designated line; S20, based on predefined constraints and the initial state of each train in the specified line, establish an immediate reward function and a delayed reward function in the process of training the strategy network; The instant reward function is obtained based on the competitive game strategy and the delay time of the train itself when each train fails to arrive at the terminal station; The delay reward function is obtained after each train arrives at the terminal station based on the cooperative game strategy and the total delay time of all trains; S30, when all trains have not arrived at the terminal station, based on the randomness and Nash equilibrium strategy of the interaction between the train and the reinforcement learning environment, it is determined whether the pre-defined constraints and the first condition satisfied by the immediate reward function are met, to select train actions, update the train status, and obtain an updated policy network; S40, after all trains arrive at the terminal station, the updated policy network is trained according to the delayed reward function of the cooperative game strategy, and the loss function of the policy network is obtained in real time; When the loss function and the reward function meet the convergence conditions, the trained policy network is obtained to realize real-time adjustment of the train operation diagram; Immediate reward function R t,i for: Delay reward function R t1,i for: Among them, F is the multiplication factor, H_dis_end-H_dis_start is the interval interruption time, r t,i It means that at time t, train i is at the departure station n t,i =0 is set to departure time Subtract the scheduled departure time a i,n , They represent the arrival time and planned arrival time of train i at station n respectively; the reward function including the immediate reward function and the delayed reward function is expressed as: T is the independent variable, t1 represents the final time of the train operation diagram adjustment time interval; The predefined constraints include: Departure time limit: Minimum running time limit for interval: Adjacent train interval limit: |a i,n -a i',n |≥a_h n ,i'∈S M (1.3) |d i,n -d i',n |≥d_h n , (1.4) |d i,n -a i',n |≥da_h n , (1.5) Station capacity restrictions: Crossing restrictions: (d i,n >d i′,n )☉(a i,n+1 >a i’,n+1 )=1, (1.8) Stop time limit: Interval interruption time limit: Among them, S N represents the set of stations, S M Represents a collection of trains; the last train departing from the originating station is numbered M, and the terminal station is numbered N; d i,n ,d i',n denote the departure time of train i and adjacent train i' at station n, represents the scheduled departure time of train i at station n, a i,n+1 represents the arrival time of train i at the next station n+1, is the given minimum running time; a i,n 、a i ' ,n denote the arrival time of train i and adjacent train i' at station n, a_h n and d_h n are the minimum time intervals between the arrival and departure of train i and its adjacent train i' at station n, da_h n represents the departure and arrival interval constraint between train i and adjacent train i' at the same station n; x t,i,n represents the train i at station n at time t; C n The capacity of each station; Indicates the minimum stop time; S NK Indicates the interruption section; H_dis_start and H_dis_end respectively indicate the start and end time of the interruption; The state S of the train at time t t,i It is expressed as: The state at the initial time t0 It is expressed as: The action of train i at time t t,i Set to: Among them, the arrival time of the current train i is and departure time is n t,i Indicates the station where train i is located at the current time t or the last station passed or arrived at; The first condition is: t,i (action * t,i ,action * t,-i )≥R t,i (action t,i ,action * t,-i ); action * t,i represents the Nash equilibrium strategy of train i, action * t,-i Indicates that all trains except train i adopt Nash equilibrium strategies; In the first condition, assume that action t-1,i =1 and the arrival condition and departure condition are met, when the random number rand≤ò, the action action is randomly selected t,i The value of makes train i continue to explore in the reinforcement learning environment; if the action selected t,i =1, indicating that train i passes station n at time t, n t,i =n t-1,i +1, when action t,i =0, it means that train i arrives at station n at time t; Assume action t-1,i = 1 and the arrival condition is met but the departure condition is not met, let action t,i =0, indicating that train i arrives at station n at time t; Assume action t-1,i = 1 and the arrival condition is not met, let action t,i =1, indicating that train i runs between station n-1 and station n at time t; Assume action t-1,i =0 and meets the starting condition, then when the random number rand>ò, according to the strategy network Q(S t,i ,action t,i ,θ)Select action action t,i When the random number rand≤ò, randomly select action action t,i Value; selected action t,i = 1, indicating that train i departs from station n-1 at time t. t,i =0, it means that train i stops at station n-1 at time t; Assume action t-1,i = 0 and the starting condition is not met, let action t,i =0, indicating that train i stops at station n-1 at time t; action t,i After the selection is made, the train takes the selected action t,i , the immediate reward function and the delayed reward function will be obtained, and the next state will be inferred, and the policy network will be updated according to the immediate reward function and the delayed reward function; The S40 includes: updating the strategy network Q(S t,i ,action t,i ,θ) is: The loss function loss is: It means finding the average value, and rand is a random number; represents the target network that helps the policy network update the settings in the reinforcement learning environment. The architecture of the target network is the same as that of the policy network. The target network uses soft updates to update parameters. θ and θ′ represent the network parameters of the policy network to be trained and the network parameters of the target network; γ represents the discount factor, α represents the learning rate, D represents the experience pool, and U(D) represents all the data in the experience pool (S t,i , action t,i , R t,i , S t+1,i ) collection; each time the train interacts with the reinforcement learning environment, it will get a set of data (S t,i , action t,i , R t,i , S t+1,i ), put it into the experience pool to train the policy network.

2. The method according to claim 1, characterized in that: Line static data includes one or more of the following: station capacity, planned operation schedule, minimum departure and arrival time interval, minimum stop time, and minimum interval operation time; The train operation dynamic data includes one or more of the following: the current running position of the train, the remaining capacity of the preceding station, the arrival and departure delay time at the current time, and the train operation status; The train operation status includes: stop, departure, arrival, and interval operation; The dynamic data at the initial moment include: the train's initial running position is the departure station, the remaining capacity of the next station is the capacity of the next station, the initial arrival and departure delay time is 0, and the train's initial running status is stop.

3. A real-time train operation adjustment device based on cooperative competition game, characterized in that: include: An acquisition unit, used for acquiring static line data of a designated line and dynamic data of initial time of all trains associated with the designated line; A function building unit, used to build an immediate reward function and a delayed reward function in the process of training the strategy network based on predefined constraints and the initial state of each train in the specified line; The instant reward function is obtained based on the competitive game strategy and the delay time of the train itself when each train fails to arrive at the terminal station; The delay reward function is obtained after each train arrives at the terminal station based on the cooperative game strategy and the total delay time of all trains; The policy network training unit is used to select train actions, update train states, and obtain updated policy networks based on the randomness and Nash equilibrium strategy of the interaction between the train and the reinforcement learning environment to determine whether the pre-defined constraints and the first condition satisfied by the immediate reward function are met when all trains have not arrived at the terminal station. After all trains arrive at the terminal station, the updated policy network is trained according to the delayed reward function of the cooperative game strategy, and the loss function of the policy network is obtained in real time. When the loss function and the reward function meet the convergence conditions, the trained policy network is obtained to realize real-time adjustment of the train operation diagram; Immediate reward function R t,i for: Delay reward function R t1,i for: Among them, F is the multiplication factor, H_dis_end-H_dis_start is the interval interruption time, r t,i It means that at time t, train i is at the departure station n t,i =0 is set to departure time Subtract the scheduled departure time a i,n , They represent the arrival time and planned arrival time of train i at station n respectively; the reward function including the immediate reward function and the delayed reward function is expressed as: T is the independent variable, t1 represents the final time of the train operation diagram adjustment time interval; The predefined constraints include: Departure time limit: Minimum running time limit for interval: Adjacent train interval limit: |a i,n -a i',n |≥a_h n ,i'∈S M (1.3) |d i,n -d i',n |≥d_h n , (1.4) |d i,n -a i',n |≥da_h n , (1.5) Station capacity restrictions: Crossing restrictions: (d i,n >d i′,n )☉(a i,n+1 >a i′n+1 )=1, (1.8) Stop time limit: Interval interruption time limit: Among them, S N represents the set of stations, S M Represents a collection of trains; the last train departing from the originating station is numbered M, and the terminal station is numbered N; d i,n ,d i',n denote the departure time of train i and adjacent train i' at station n, represents the scheduled departure time of train i at station n, a i,n+1 represents the arrival time of train i at the next station n+1, is the given minimum running time; a i,n 、a i',n denote the arrival time of train i and adjacent train i' at station n, a_h n and d_h n are the minimum time intervals between the arrival and departure of train i and its adjacent train i' at station n, da_h n represents the departure and arrival interval constraint between train i and adjacent train i' at the same station n; x t,i,n represents the train i at station n at time t; C n The capacity of each station; Indicates the minimum stop time; S NK Indicates the interruption section; H_dis_start and H_dis_end respectively indicate the start and end time of the interruption; The state St,i of the train at time t is expressed as: The state at the initial time t0 It is expressed as: The action of train i at time t t,i Set to: Among them, the arrival time of the current train i is and departure time is n t,i Indicates the station where train i is located at the current time t or the last station passed or arrived at; The first condition is: R t,i (action * t,i ,action * t,-i )≥R t,i (action t,i ,action * t,-i ); action * t,i represents the Nash equilibrium strategy of train i, action * t,-i Indicates that all trains except train i adopt Nash equilibrium strategies; In the first condition, assume that action t-1,i =1 and the arrival condition and departure condition are met, when the random number rand≤ò, the action action is randomly selected t,i The value of makes train i continue to explore in the reinforcement learning environment; if the action selected t,i =1, indicating that train i passes station n at time t, n t,i =n t-1,i +1, when action t,i =0, it means that train i arrives at station n at time t; Assume action t-1,i = 1 and the arrival condition is met but the departure condition is not met, let action t,i =0, indicating that train i arrives at station n at time t; Assume action t-1,i = 1 and the arrival condition is not met, let action t,i =1, indicating that train i runs between station n-1 and station n at time t; Assume action t-1,i =0 and meets the starting condition, then when the random number rand>ò, according to the strategy network Q(S t,i ,action t,i ,θ)Select action action t,i When the random number rand≤ò, randomly select action action t,i Value; selected action t,i = 1, indicating that train i departs from station n-1 at time t. t,i =0, it means that train i stops at station n-1 at time t; Assume action t-1,i = 0 and the starting condition is not met, let action t,i =0, indicating that train i stops at station n-1 at time t; action t,i After the selection is made, the train takes the selected action t,i , the immediate reward function and the delayed reward function will be obtained, and the next state will be inferred, and the policy network will be updated according to the immediate reward function and the delayed reward function; The S40 includes: updating the strategy network Q(S t,i ,action t,i ,θ) is: The loss function loss is: It means finding the average value, and rand is a random number; represents the target network that helps the policy network update the settings in the reinforcement learning environment. The architecture of the target network is the same as that of the policy network. The target network uses soft updates to update parameters. θ and θ′ represent the network parameters of the policy network to be trained and the network parameters of the target network; γ represents the discount factor, α represents the learning rate, D represents the experience pool, and U(D) represents all the data in the experience pool (S t,i , action t,i , R t,i , S t+1,i ) collection; each time the train interacts with the reinforcement learning environment, it will get a set of data (S t,i , action t,i , R t,i , S t+1,i ), put it into the experience pool to train the policy network.

4. A computing device, characterized in that include: It includes a processor, a memory and a computer program stored in the memory, and the processor executes the computer program to implement the real-time adjustment method of train operation based on cooperative competition game as described in any one of claims 1-2.

Citation Information

Patent Citations

  • High-speed train operation adjustment method and system based on Q learning

    CN113415322A

  • Timetable adjusting method and device under delay condition, and electronic equipment

    CN113525462A