Multi-train dynamic scheduling method and system based on MAPPO algorithm
By using a multi-train dynamic scheduling method based on the MAPPO algorithm, a distributed decision-making architecture and a multi-objective compound reward mechanism are constructed, which solves the problem of lack of coordinated global scheduling in train scheduling in existing technologies and realizes efficient train coordinated scheduling and safe operation.
Patent Information
- Application Number
- CN202511212223.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-08-28
AI Technical Summary
The train scheduling methods in the existing technology are still limited to single-train decision-making or local competition paradigms, lack the ability to coordinate global scheduling of trains, and find it difficult to achieve efficient coordinated scheduling of network resources under sudden interference events.
A multi-train dynamic scheduling method based on the MAPPO algorithm is adopted. By constructing a distributed decision-making architecture, each train is built into an independent intelligent entity, relying on local observation to achieve autonomous decision-making, integrating a multi-objective compound reward mechanism, establishing a centralized training and distributed execution collaborative paradigm, and optimizing group behavior coordination.
It significantly improves the efficiency of road network resource utilization and operational safety, can achieve millisecond-level response in large-scale road networks, effectively suppress the propagation of delays, and improve the punctuality and safety of train arrivals.
Smart Images

Figure CN120716796A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multi-train dynamic scheduling, and in particular to a multi-train dynamic scheduling method and system based on a MAPPO algorithm. Background Art
[0002] As the scale of the road network continues to expand and traffic density increases, the propagation effect of train delays caused by sudden interference events such as equipment failure, natural disasters, and transportation safety is becoming increasingly prominent.
[0003] like Figure 1 As shown in Figure 1, under normal circumstances, high-speed railway trains run efficiently and orderly on sections 1 to 4 between stations A to E according to the established timetable. However, in reality, the operation of the railway system will inevitably encounter various sudden interference events. Figure 1 As shown in Figure 2, a sudden interference event occurred in section 3 near station C, where H start Indicates the start time of the sudden interference event, H end represents the end time of the sudden disruption. Precisely because train operations within the railway system are closely interconnected in time and space, this localized failure quickly triggered a chain reaction. Specifically, when train k2 arrived at station C, its actual departure time was forced to adjust, resulting in the scheduled departure time interval between it and the trains preceding and following it falling short of the required safety interval. This directly prevented the train from departing as originally planned, inevitably leading to a departure delay. This initial delay was not isolated; it spread like a domino effect throughout the chain of subsequent operations. Under the impact of the sudden disruption, the actual and planned trajectories of subsequent trains exhibited significant and continuously increasing deviations. The original schedule was disrupted, and the arrival and departure times of subsequent trains were forced to be postponed across the board, severely disrupting operational order and significantly reducing overall transportation efficiency.
[0004] To address these issues, existing technologies primarily focus on two categories of approaches. One approach employs optimization models based on mathematical programming and intelligent algorithms. These approaches employ mathematical models (e.g., mixed integer programming, multi-objective particle swarm optimization, and the fireworks algorithm) that address multiple objectives (e.g., minimizing delay time, resource adjustments, energy consumption, and cost) and develop two-stage or multi-stage optimization strategies tailored to specific scenarios (e.g., snowstorms and single-line faults). These approaches have demonstrated success in achieving multi-objective coordination and adaptability, but they generally rely on predefined rules and struggle to meet the real-time response requirements of highly dynamic environments. The other approach utilizes reinforcement learning techniques, leveraging their interactive learning capabilities (e.g., Q-learning and its modifications, proximal policy optimization, and deep deterministic policy gradient algorithms). Through innovative approaches such as reward mechanisms, lightweight state representations, model-free methods, and multi-agent competitive learning, these approaches demonstrate significant potential for improving learning efficiency, effectively suppressing delay propagation, and collaboratively optimizing objectives (e.g., delay and energy consumption). However, current research on reinforcement learning methods is mostly still focused on single-train decision-making or simple competition and cooperation paradigms within a local scope. There is no effective solution for how to achieve coordinated scheduling of global network resources (such as station arrival and departure lines, and interval capacity), which has become a key bottleneck that urgently needs to be broken through. Summary of the Invention
[0005] Based on this, the purpose of the present invention is to provide a multi-train dynamic scheduling method and system based on the MAPPO algorithm, which is used to solve the technical problem that the train scheduling method in the existing technology is still limited to single-train decision-making or local competition paradigm under sudden interference events, and lacks the ability to coordinate global scheduling of trains.
[0006] In one aspect, the present invention provides a multi-train dynamic scheduling method based on the MAPPO algorithm, comprising: Acquiring multiple train operating status information, and identifying the operating status information based on a pre-trained neural network strategy model to generate discrete action decisions; Generate adjustment instructions based on discrete action decisions, including speed adjustment and start / stop; calculate new operating state information and compound reward feedback for the current decision step based on the adjustment instructions, and update the train environment state information to enter the next decision step until the end of the current round; wherein each round includes multiple decision steps, each decision step corresponds to generating an action decision, and each round generates a scheduling strategy, each scheduling strategy includes multiple action decisions; Determine whether the scheduling strategy corresponding to the current round complies with the preset scheduling strategy based on all rewards of the current round; If not, return to the step of obtaining the running status information of multiple trains until the scheduling strategy corresponding to the current round meets the preset scheduling strategy to dynamically coordinate and optimize the scheduling of multiple trains.
[0007] This multi-train dynamic scheduling method, based on the MAPPO algorithm, overcomes the scheduling conflict bottleneck caused by the lack of a collaborative communication mechanism in traditional single-agent systems through a dynamic scheduling framework based on multi-agent proximal policy optimization (MAPPO). Specifically, with the goal of minimizing network-wide train delays and delays of critical trains, a distributed decision-making framework implements a dynamic avoidance strategy for high-priority trains. Furthermore, this distributed decision-making architecture enables autonomous and coordinated train scheduling by dynamically constructing each train into an agent with local perception and real-time decision-making capabilities. Specifically, it leverages the 11-dimensional state characteristics of operating status information to achieve autonomous decision-making under local observation. It innovatively integrates a multi-objective compound reward mechanism: safe operation rewards and penalties, speeding penalties, delay improvement rewards, on-time arrival rewards, and destination arrival rewards. It also establishes a collaborative paradigm of centralized training and distributed execution to optimize group behavior coordination, significantly improving network resource utilization efficiency and operational safety. This solves the technical problem that existing train scheduling methods are still limited to single-train decision-making or local competition paradigms, lacking collaborative global scheduling capabilities.
[0008] In addition, the multi-train dynamic scheduling method based on the MAPPO algorithm according to the present invention may also have the following additional technical features: Furthermore, in the step of calculating the new operating state information and the compound reward feedback of the current decision step according to the adjustment instruction and updating the train environment state information, the total reward integration mechanism calculation formula of the compound reward feedback is: ; Where: k represents whether the train is at the station, where 1 represents the train is at the station, and 0 represents the train is not at the station or has just arrived at the station; r in represents the total reward returned by the environment when the train is in the station, where r in = r dwell , r dwell represents the stop efficiency bonus; r out represents the total reward returned by the environment when the train is outside the station or has just arrived at the station, where ; Where, r safe Indicates rewards and penalties for safe operations; r speed Indicates speeding penalty items; r delay Indicates delayed improvement reward; r punct Indicates an on-time arrival reward; rterm Indicates the reward for reaching the end point; ∩ indicates the intersection.
[0009] Furthermore, the calculation formula for delay improvement bonus is: ; Where, r delay Indicates delayed improvement reward; d t Indicates current delay; d t-1 represents the delay at the previous moment; d0 represents the difference between the current delay and the delay at the previous moment; ∩ represents the intersection; l delay Represents the delay improvement reward coefficient.
[0010] Furthermore, the calculation formula for the on-time arrival bonus is: ; Where, r punct Indicates an on-time arrival reward; represents the calculation of the punctual linear attenuation function, where Indicates the actual time when the train arrives at station s, It indicates the time when the train is scheduled to arrive at station s; l punct represents the on-time arrival bonus coefficient; s represents the station number.
[0011] Furthermore, the calculation formula for the destination reward is: ; Where, r term Indicates that the end point has been reached and the reward; l term Indicates the reward coefficient for reaching the destination; f Indicates the gain coefficient of train stopping at large stations; d t Indicates the current delay; 60 means 60 minutes, which is the maximum acceptable delay limit for the terminal.
[0012] Furthermore, the calculation formula for safety operation rewards and penalties is: ; Where, r safe Indicates rewards and penalties for safe operations; d front Indicates the distance between the current train and the preceding train; l safe Represents the safety operation reward coefficient.
[0013] Furthermore, the calculation formula for the stop efficiency bonus is: ; Where, r dwell represents the stop efficiency bonus; l dwell represents the stop efficiency bonus coefficient; t stop Indicates the actual stop time; t min Indicates the minimum stop time; t target Indicates the target stop time.
[0014] Furthermore, the calculation formula for speeding penalty items is: ; Where, r speed Indicates speeding penalty items; l speed represents the speeding penalty reward coefficient; v Indicates the current speed; v lim Indicates the maximum speed allowed.
[0015] Furthermore, the step of identifying the operating state information according to the pre-trained neural network strategy model to generate a discrete action decision includes: Obtaining a model initialization instruction, and initializing a pre-trained neural network strategy model according to the model initialization instruction; The operating state information is identified according to the initialized pre-trained neural network strategy model to generate discrete action decisions.
[0016] In one aspect, the present invention further provides a multi-train dynamic dispatching system based on the MAPPO algorithm, the system comprising: an acquisition module, configured to acquire operating status information of multiple trains and identify the operating status information according to a pre-trained neural network strategy model to generate discrete action decisions; A training module is configured to generate adjustment instructions based on discrete action decisions, including speed adjustment and start / stop instructions, calculate new operating state information and compound reward feedback for the current decision step based on the adjustment instructions, and update the train environment state information to proceed to the next decision step until the end of the current round; wherein each round includes multiple decision steps, each decision step corresponds to an action decision, and each round generates a scheduling strategy, each scheduling strategy including multiple action decisions; A judgment module is used to judge whether the scheduling strategy corresponding to the current round conforms to the preset scheduling strategy based on all rewards of the current round; The first execution module is used to return to the step of obtaining the operation status information of multiple trains if the scheduling strategy corresponding to the current round does not conform to the preset scheduling strategy, until the scheduling strategy corresponding to the current round conforms to the preset scheduling strategy to dynamically coordinate and optimize the scheduling of multiple trains. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 Schematic diagram for adjusting trains in the event of an unexpected interference event; Figure 2 This is a decision flow chart of an intelligent agent for adjusting the operation of a high-speed railway train in an embodiment of the present invention; Figure 3 Flowchart of a multi-train dynamic scheduling method based on the MAPPO algorithm in an embodiment of the present invention; Figure 4 This is an optimized timetable for dynamic train scheduling under sudden interference events in an embodiment of the present invention.
[0018] The following specific embodiments will further illustrate the present invention in conjunction with the above-mentioned drawings. DETAILED DESCRIPTION
[0019] To facilitate understanding of the present invention, the present invention will be described more fully below with reference to the accompanying drawings. The drawings illustrate several embodiments of the present invention. However, the present invention may be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and comprehensive understanding of the present invention.
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one skilled in the art to which this invention pertains. The terms used in this specification of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0021] To address the technical issues in sudden interference events, where existing train scheduling methods are limited to single-train decision-making or local competition, lacking the ability to coordinate global train scheduling, this application provides a multi-train dynamic scheduling method and system based on the MAPPO algorithm. Addressing the limitations of existing technologies in coordination and real-time decision-making, the architectural characteristics of Multi-Agent Proximal Policy Optimization (MAPPO) are deeply coupled with the requirements of high-speed rail scheduling. High-speed rail dynamic scheduling requires groups of trains to collaborate to minimize network delays, a task naturally facilitated by MAPPO's centralized training with decentralized execution (CTDE) mechanism. The high efficiency of neural network forward computation directly supports millisecond-level response in large-scale networks with more than 200 trains. Together, these two methods constitute a systematic solution that breaks through the bottlenecks of existing technologies.
[0022] Based on this, this application adopts the MAPPO method. This method constructs a distributed decision-making architecture, making each train an independent intelligent agent and relying on local observations to make autonomous decisions. It innovatively designs a composite reward mechanism that integrates punctuality, safe spacing, speeding rules, and minimum safe stop time rules. It establishes a collaborative paradigm of centralized training with decentralized execution (CTDE) to optimize group behavior coordination. An empirical study using the Shanghai-Kunming High-Speed Railway in China shows that this method demonstrates significant advantages under typical interference scenarios, such as speed restrictions due to equipment failures and line occupancy due to inclement weather. It can also be expanded to large-scale networks with more than 200 trains to enhance real-time response capabilities, providing a new theoretical paradigm and engineering practice basis for the intelligent scheduling of complex transportation systems. Specifically, the MAPPO dynamic scheduling framework is as follows Figure 2As shown in Figure 1, its core is a closed-loop adaptive real-time interactive system composed of train agents. The system's operating mechanism is as follows: the train agent continuously collects 11-dimensional state feature vectors from the high-speed rail operating environment, covering information on interference events, train status, line resources, and scheduling objectives, achieving a holistic perception of the operating situation. Interference event information includes the event impact location, event duration, and expected recovery time; train status information includes train section location, train speed, stop status, and stop time; line resource information includes the number of station-to-departure lines and section capacity; and scheduling objectives include key train identifiers and timetable offsets. Based on current local observations, the agent invokes a frozen neural network policy model trained online using the MAPPO algorithm. In a distributed execution mode, it generates five discrete action decisions, covering acceleration, deceleration, and speed maintenance for actions outside the station; and continued parking and departure for actions within the station. During interval operation, the train executes precise speed regulation of ±10 km / h (constrained within the 200 km / h-352 km / h safety zone) or maintains speed. At station stops, trains can choose to force departure or extend their stops. Safety constraints are rigidly integrated through a state-action linkage mechanism. The generated speed regulation or start-stop commands are applied to the high-speed rail simulation environment, which updates operating status information based on the train dynamics model, tracking interval rules, and resource utilization logic, while simultaneously calculating compound reward feedback. The intelligent agent then observes the updated environmental state information, including delay time and safety distance, forming a millisecond-level closed-loop control flow of "perception → decision → execution → evaluation", with a maximum number of decision steps exceeding 200, until all trains arrive at their final destination.
[0023] To facilitate understanding of the present invention, several embodiments of the present invention are provided below. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and comprehensive disclosure of the present invention.
[0024] Example 1 See also Figure 3 , which shows a multi-train dynamic scheduling method based on the MAPPO algorithm in a first embodiment of the present invention, the method includes steps S101 to S104: S101. Acquire the running status information of multiple trains, and identify the running status information according to a pre-trained neural network strategy model to generate discrete action decisions.
[0025] In this embodiment, the train is an intelligent agent, and its action decisions include actions outside the station and actions inside the station. Actions outside the station include acceleration, deceleration, and maintaining speed; actions inside the station include continuing to stop and starting. Figure 2As shown, the operating status information of multiple trains under sudden interference events is obtained. The operating status information includes four major status types: interference event information, train status information, line resource information, and dispatch target information. Furthermore, the four major status types include 11-dimensional state feature vectors. Specifically, interference event information includes the event impact location, event duration, and expected recovery time; train status information includes train section location, train speed, stop status, and stop time; line resource information includes the number of station-to-departure lines and section capacity; dispatch target information includes key train identifiers and timetable offsets. The 11-dimensional feature vector is used to accurately encode the train operating status. The core elements include train identification features, spatiotemporal state features, operating status features, station relationship features, safety status features, and departure control features.
[0026] In this embodiment, a reinforcement learning simulation environment is first constructed and the neural network strategy model is trained until convergence using train history or simulation data. Then, the scheduling execution phase is entered to obtain real-time operating status information including interference events, train status, line resources and scheduling goals, and input it into the trained strategy model to generate discrete action decisions for speed regulation or start and stop.
[0027] To ensure strict compatibility between the pre-trained neural network policy model and the pre-trained architecture, the pre-trained neural network policy model must first be initialized according to the model initialization instructions. The initialized pre-trained neural network policy model then identifies the operating state information to generate discrete action decisions. The input layer of the pre-trained architecture matches the 11-dimensional state feature vector, and the output layer corresponds to the joint action space.
[0028] In this embodiment, the initialized, pre-trained neural network policy model is capable of outputting action decisions based on real-time information collected from multiple train operating states. Its output is designed as follows: the first three neurons map actions outside the station (acceleration, deceleration, and maintaining speed); the last two neurons map actions within the station (continued stop and departure). After initialization, the model can be used directly for real-time inference, generating action probability distributions by processing state vectors containing core elements such as train identity, spatiotemporal state, and safety features.
[0029] S102. Generate adjustment instructions based on discrete action decisions, calculate new operating status information and the compound reward feedback of the current decision step based on the adjustment instructions, and update the train environment status information to enter the next decision step until the end of the current round.
[0030] In this embodiment, a neural network policy model acquires real-time operational status information, including interference events, train status, line resources, and scheduling objectives, to generate discrete action decisions for speed adjustment or start / stop. Adjustment instructions include speed adjustment (i.e., adjusting the operating speed within a certain interval) and start / stop (i.e., whether to continue stopping or starting at a station). After the generated adjustment instructions are applied to the high-speed rail simulation environment, the environment updates the train environment status information based on the train dynamics model, tracking interval rules, and resource utilization logic, and simultaneously calculates compound reward feedback. The intelligent agent then observes the updated train environment status information, forming a millisecond-level closed-loop control flow of "perception → decision → execution → evaluation". The maximum number of decision steps is greater than 200, until all trains arrive at the terminal. The environmental status information includes stop time, delay time, and safe distance.
[0031] In this embodiment, each round includes multiple decision steps, each of which generates an action decision. Each round generates a scheduling strategy, and each scheduling strategy includes multiple action decisions. Furthermore, as a specific example, the reward function is primarily designed around safe operation and delay control. To ensure on-time train arrival, the reward function should penalize late arrivals. The longer the delay, the greater the penalty, especially at key hubs, to minimize the impact on overall scheduling. At the same time, premature arrivals can disrupt operational order, so appropriate penalties should also be imposed to ensure that trains operate on schedule. Safe operation is a fundamental prerequisite for scheduling optimization. Train spacing must be maintained above a safe threshold. If the spacing is too small, a larger penalty is imposed to avoid the risk of rear-end collisions. Furthermore, rapid train speed changes can affect driving stability, so penalties should be imposed for drastic acceleration and deceleration to maintain safe and stable operation. In summary, the reward function should be designed to minimize train delays while ensuring safe distances and driving stability, enabling the agent to learn to optimize scheduling strategies and achieve reliable high-speed rail operation. Specifically, the total reward integration mechanism (i.e., reward function) for compound reward feedback is calculated as follows: ; Where: k represents whether the train is at the station, where 1 represents the train is at the station and 0 represents the train is not at the station or has just arrived at the station; r in represents the total reward returned by the environment when the train is in the station, where r in = r dwell , r dwell represents the stop efficiency bonus; r out represents the total reward returned by the environment when the train is outside the station or has just arrived at the station, where ; Where, r safe Indicates rewards and penalties for safe operations; r speed Indicates speeding penalty items; r delay Indicates delayed improvement reward; r punct Indicates an on-time arrival reward; r term Indicates the reward for reaching the end point; ∩ indicates the intersection.
[0032] Furthermore, the calculation formula for delay improvement bonus is: ; Where, r delay Indicates delayed improvement reward; d t Indicates current delay; d t-1 represents the delay at the previous moment; d0 represents the difference between the current delay and the delay at the previous moment, that is, d t-1 - d t The difference term of ; ∩ represents the intersection; l delay represents the delay improvement reward coefficient, l delay =300; 1 / 15 represents the attenuation coefficient.
[0033] The decay mechanism of delay improvement reward and the upper limit constraint of speeding penalty items jointly suppress the policy oscillation. In order to solve the problem that the traditional method is not strong enough in the case of continuous delay, such as cascading delay spread, a hysteresis differential mechanism is designed. d t-1 - d t Dynamically perceive delay trends, and the attenuation coefficient of 1 / 15 alleviates strategy shocks caused by short-term deterioration.
[0034] The internal logic of the delay improvement reward is: when the delay is improved (ie: t ≤d t-1 ), full weight excitation (λ delay =300), strengthen the positive feedback. t >d t-1 ) and does not reach the severe delay threshold (ie: d t <15min), the penalty intensity is attenuated to prevent the agent from giving up long-term optimization due to temporary deterioration. t <15min is set according to high-speed rail dispatching specifications to prevent small fluctuations from interfering with decision-making.
[0035] Furthermore, in order to drive station timing optimization, this embodiment provides an on-time arrival reward. Specifically, the calculation formula for the on-time arrival reward is: ; Where, r punct Indicates an on-time arrival reward; represents the calculation of the punctual linear attenuation function, where Indicates the actual time when the train arrives at station s, It indicates the time when the train is scheduled to arrive at station s; l punct Indicates the on-time arrival reward coefficient, balancing the weights of local and global goals ,l punct =50; s represents the station number.
[0036] Traditional binary punctuality judgments (on time / late) cannot guide micro-scheduling strategies. The on-time arrival reward in this embodiment innovatively introduces a flexible time window (±8 minutes) to achieve fine timing control. Among them, "±8 minutes" is the passenger tolerance threshold during the train operation, that is, the maximum acceptable deviation for high-speed rail passengers at intermediate stations.
[0037] The internal logic of the calculation formula for on-time arrival rewards is: a linear decay function maps time deviation to a continuous reward value (where the reward is maximum when the deviation = 0).
[0038] Furthermore, the calculation formula for the destination reward is: ; Where, r term Indicates that the end point has been reached and the reward; l term Indicates the reward coefficient for reaching the end point, l term =1500; f It represents the gain coefficient of train stopping at large stations, f= 1.8; d t Indicates the current delay; 60 means 60 minutes, which is the maximum acceptable delay limit for the terminal.
[0039] In order to resolve the contradiction between "priority protection for key trains" and "efficiency of ordinary trains", this application innovatively adopts the gain coefficient f Quantify the injection scheduling priority and break through the static limitations of traditional rule-based methods.
[0040] Specifically, the internal logic of the calculation formula for the destination arrival reward is: max(0.2,•) ensures that the lower limit of the basic reward is 0.2 to prevent strategy collapse (such as the reward returning to zero in the event of extreme delay).end / 60 compresses the terminal delay to the interval [0,1] to implement proportional penalty. f =1.8 Dynamically amplify the key train rewards to drive resource allocation; among them, when the train is an ordinary train, f =1.
[0041] In order to curb speeding and avoid chain delays, this application provides safe operation rewards and penalties. Furthermore, the calculation formula for safe operation rewards and penalties is: ; Where, r safe Indicates rewards and penalties for safe operations; d front Indicates the distance between the current train and the preceding train; l safe represents the safety operation reward coefficient, l safe =200.
[0042] In this embodiment, the physical safety rule (i.e., tracking interval ≥ 10 km) is converted into a differentiable optimization objective, avoiding the infeasible decisions caused by traditional hard constraints. The safety operation reward and penalty calculation formula converts the discrete safety rule (i.e., minimum interval 5 km) into a continuously differentiable function, supporting gradient optimization.
[0043] The internal logic of the safety operation reward and punishment calculation formula is: Danger zone (ie: d front <8km): Penalty value -(8-d front ) / 8•λ safe It increases exponentially as the distance decreases, simulating the need for emergency braking; Buffer (ie: 8≤d front <16km: Positive reward min(1-d) front / 16)•λ safe / 2 encourages active risk avoidance. The boundary value (i.e. 8 km) is aligned with the minimum safe distance D min =5km, with safety redundancy reserved.
[0044] In order to reduce resource conflicts, improve resource utilization, and support efficient execution of the basic target layer, this application provides a stop efficiency reward to balance passenger boarding and alighting needs (minimum stop time τ min ) and scheduling efficiency (target stop time τ target ), to resolve resource occupation conflicts. Furthermore, the calculation formula for the stop efficiency reward is: ; Where, r dwell represents the stop efficiency bonus; ldwell represents the stop efficiency bonus coefficient, l dwell =30, dynamically adjusted based on station level; t stop Indicates the actual stop time; t min Indicates the minimum stop time; t target Indicates the target stop time.
[0045] Specifically, the internal logic of the calculation formula for the stop efficiency reward is: when the target is not met (i.e., when τ stop <τ min When ), compensation reward (τ stop / τ min )•λ dwell •0.5, weight 0.5 to prevent overcompensation; after reaching the target (ie: when τ min ≤τ stop ≤τ target When the stop time is linearly decayed (1-(τ stop -τ min ) / (τ target -τ min ))•λ dwell , approaching the efficiency ceiling.
[0046] In order to address the dynamic risks (e.g., acceleration mutations) caused by aggressive speed regulation, a progressive penalty mechanism is designed to overcome the limitations of traditional binary constraints. Specifically, the speeding penalty calculation formula is: ; Where, r speed Indicates speeding penalty items; l speed represents the speeding penalty reward coefficient, l speed =-20; v Indicates the current speed; v lim Indicates the maximum speed allowed.
[0047] Specifically, the internal logic of the speeding penalty calculation formula is: proportional term ((vv lim ) / v lim )•λ speed For overspeed (i.e. v >1.08 v lim ) is penalized according to the degree of deviation (i.e., λ speed =-20). max(-100,•) sets the upper limit of the single-step penalty to avoid policy shock (e.g., a sharp fluctuation in reward when slightly exceeding the speed limit).v lim Dynamically bind section speed limits to adapt to changing environments.
[0048] S103. Determine whether the scheduling strategy corresponding to the current round complies with the preset scheduling strategy based on all rewards of the current round.
[0049] If the scheduling strategy corresponding to the current round meets the preset scheduling strategy, step S104 is executed; if the scheduling strategy corresponding to the current round does not meet the preset scheduling strategy, the process returns to step S101 until the scheduling strategy corresponding to the current round meets the preset scheduling strategy to dynamically coordinate and optimize multi-train scheduling.
[0050] S104. Dynamically and collaboratively optimize multi-train scheduling based on the scheduling strategy corresponding to the current round.
[0051] like Figure 4 The figure below shows an optimized train schedule for dynamic scheduling under sudden disruptions. Combined with Table 1, the MAPPO algorithm demonstrates significant comprehensive advantages over the traditional PPO algorithm, using the traditional PPO algorithm as a benchmark. In terms of on-time performance, the multi-agent collaborative decision-making mechanism employed in this application improves average terminal on-time performance by 20% (77% vs. 97%) and average key train on-time performance by 15% (85% vs. 100%). This is due to the centralized critic network's early identification of delay propagation paths and the effective implementation of intelligent adjustment strategies such as overtaking and route replanning, which significantly suppresses the spread of cascading delays. In terms of dynamic adaptation, the algorithm's convergence rounds are shortened by 33%, demonstrating the rapid policy optimization capabilities of a centralized training and distributed execution architecture in non-steady-state environments. The PPO algorithm in Table 1 is the traditional Proximal Policy Optimization (PPO) algorithm.
[0052] Table 1:
[0053] In summary, the multi-train dynamic scheduling method based on the MAPPO algorithm in the above-mentioned embodiments of the present invention, through a dynamic scheduling framework based on multi-agent proximal policy optimization (MAPPO), overcomes the scheduling conflict bottleneck caused by the lack of a collaborative communication mechanism in traditional single-agent systems. Specifically, with the goal of minimizing network-wide train delays and delays of critical trains, a distributed decision-making framework implements a dynamic avoidance strategy for high-priority trains. Furthermore, this distributed decision-making architecture enables autonomous and coordinated train scheduling by dynamically constructing each train into an agent with local perception and real-time decision-making capabilities. Specifically, it leverages the 11-dimensional state features covered by operating status information to achieve autonomous decision-making under local observation. It innovatively integrates a multi-objective composite reward mechanism: safe operation rewards and penalties, speeding penalties, delay improvement rewards, on-time arrival rewards, and destination arrival rewards. It also establishes a collaborative paradigm of centralized training and distributed execution to optimize group behavior coordination, significantly improving network resource utilization efficiency and operational safety. This solves the technical problem that existing train scheduling methods are limited to single-train decision-making or local competition paradigms, lacking the ability to coordinate global scheduling.
[0054] Example 2 The second embodiment of the present invention provides a multi-train dynamic dispatching system based on the MAPPO algorithm, including: an acquisition module, configured to acquire operating status information of multiple trains and identify the operating status information according to a pre-trained neural network strategy model to generate discrete action decisions; A training module is configured to generate adjustment instructions based on discrete action decisions, including speed adjustment and start / stop instructions, calculate new operating state information and compound reward feedback for the current decision step based on the adjustment instructions, and update the train environment state information to proceed to the next decision step until the end of the current round; wherein each round includes multiple decision steps, each decision step corresponds to an action decision, and each round generates a scheduling strategy, each scheduling strategy including multiple action decisions; A judgment module is used to judge whether the scheduling strategy corresponding to the current round conforms to the preset scheduling strategy based on all rewards of the current round; The first execution module is used to return to the step of obtaining the operation status information of multiple trains if the scheduling strategy corresponding to the current round does not conform to the preset scheduling strategy, until the scheduling strategy corresponding to the current round conforms to the preset scheduling strategy to dynamically coordinate and optimize the scheduling of multiple trains.
[0055] In summary, the multi-train dynamic scheduling system based on the MAPPO algorithm in the above-mentioned embodiments of the present invention, through a dynamic scheduling framework based on multi-agent proximal policy optimization (MAPPO), overcomes the scheduling conflict bottleneck caused by the lack of a collaborative communication mechanism in traditional single-agent systems. Specifically, with the goal of minimizing network-wide train delays and delays of critical trains, a distributed decision-making framework implements a dynamic avoidance strategy for high-priority trains. Furthermore, this distributed decision-making architecture enables autonomous and coordinated train scheduling by dynamically constructing each train into an agent with local perception and real-time decision-making capabilities. Specifically, it leverages the 11-dimensional state characteristics covered by operating status information to achieve autonomous decision-making under local observation. It innovatively integrates a multi-objective composite reward mechanism, including safe operation rewards and penalties, speeding penalties, delay improvement rewards, on-time arrival rewards, and destination arrival rewards. It also establishes a collaborative paradigm of centralized training and distributed execution to optimize group behavior coordination, significantly improving network resource utilization efficiency and operational safety. This solves the technical problem that existing train scheduling methods are limited to single-train decision-making or local competition paradigms, lacking the ability to coordinate global scheduling.
[0056] In addition, an embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the method in the above embodiment when the program is executed by a processor.
[0057] In addition, an embodiment of the present invention further provides a data processing device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method in the above embodiment when executing the program.
[0058] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0059] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting, or processing it in another suitable manner as necessary, and then storing it in a computer memory.
[0060] It should be understood that various components of the present invention may be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods may be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof may be used: a discrete logic circuit having logic gate circuits for implementing logic functions on data signals, an application-specific integrated circuit having suitable combinational logic gate circuits, a programmable gate array (PGA), a field-programmable gate array (FPGA), etc.
[0061] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0062] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the claims and their equivalents.
Claims
1. A multi-train dynamic scheduling method based on the MAPPO algorithm, characterized in that: include: Acquiring multiple train operating status information, and identifying the operating status information based on a pre-trained neural network strategy model to generate discrete action decisions; Generate adjustment instructions based on discrete action decisions, including speed adjustment and start / stop; calculate new operating state information and compound reward feedback for the current decision step based on the adjustment instructions, and update the train environment state information to enter the next decision step until the end of the current round; wherein each round includes multiple decision steps, each decision step corresponds to generating an action decision, and each round generates a scheduling strategy, each scheduling strategy includes multiple action decisions; Determine whether the scheduling strategy corresponding to the current round complies with the preset scheduling strategy based on all rewards of the current round; If not, return to the step of obtaining the running status information of multiple trains until the scheduling strategy corresponding to the current round meets the preset scheduling strategy to dynamically coordinate and optimize the scheduling of multiple trains.
2. The multi-train dynamic scheduling method based on the MAPPO algorithm according to claim 1 is characterized in that: In the step of calculating the new operating state information and the compound reward feedback of the current decision step according to the adjustment instruction and updating the train environment state information, the total reward integration mechanism calculation formula of the compound reward feedback is: ; Where: k represents whether the train is at the station, where 1 represents the train is at the station, and 0 represents the train is not at the station or has just arrived at the station; r in represents the total reward returned by the environment when the train is in the station, where r in = r dwell , r dwell represents the stop efficiency bonus; r out represents the total reward returned by the environment when the train is outside the station or has just arrived at the station, where ; Where, r safe Indicates rewards and penalties for safe operations; r speed Indicates speeding penalty items; r delay Indicates delayed improvement reward; r punct Indicates an on-time arrival reward; r term Indicates the reward for reaching the end point; ∩ indicates the intersection.
3. The multi-train dynamic scheduling method based on the MAPPO algorithm according to claim 2 is characterized in that: The calculation formula for delay improvement bonus is: ; Where, r delay Indicates delayed improvement reward; d t Indicates current delay; d t-1 represents the delay at the previous moment; d0 represents the difference between the current delay and the delay at the previous moment; ∩ represents the intersection; λ delay Represents the delay improvement reward coefficient.
4. The multi-train dynamic scheduling method based on the MAPPO algorithm according to claim 2, characterized in that: The calculation formula for the on-time arrival bonus is: ; Where, r punct Indicates an on-time arrival reward; represents the calculation of the punctual linear attenuation function, where Indicates the actual time when the train arrives at station s, It indicates the time when the train is scheduled to arrive at station s; λ punct represents the on-time arrival bonus coefficient; s represents the station number.
5. The multi-train dynamic scheduling method based on the MAPPO algorithm according to claim 2, characterized in that: The calculation formula for the reward for reaching the end point is: ; Where, r term Indicates that the end point has been reached and the reward; λ term Indicates the reward coefficient for reaching the destination; f Indicates the gain coefficient of train stopping at large stations; d t Indicates the current delay; 60 means 60 minutes, which is the maximum acceptable delay limit for the terminal.
6. The multi-train dynamic scheduling method based on the MAPPO algorithm according to claim 2, characterized in that: The calculation formula for safety operation rewards and penalties is: ; Where, r safe Indicates rewards and penalties for safe operations; d front Indicates the distance between the current train and the preceding train; λ safe Represents the safety operation reward coefficient.
7. The multi-train dynamic scheduling method based on the MAPPO algorithm according to claim 2, characterized in that: The calculation formula for the stop efficiency bonus is: ; Where, r dwell represents the stop efficiency bonus; λ dwell represents the stop efficiency bonus coefficient; τ stop Indicates the actual stop time; τ min Indicates the minimum stop time; τ target Indicates the target stop time.
8. The multi-train dynamic scheduling method based on the MAPPO algorithm according to claim 2, characterized in that: The calculation formula for speeding penalty items is: ; Where, r speed Indicates speeding penalty items; λ speed represents the speeding penalty reward coefficient; v Indicates the current speed; v lim Indicates the maximum speed allowed.
9. The multi-train dynamic scheduling method based on the MAPPO algorithm according to claim 1, characterized in that: The step of identifying the operating state information according to the pre-trained neural network strategy model to generate a discrete action decision includes: Obtaining a model initialization instruction, and initializing a pre-trained neural network strategy model according to the model initialization instruction; The operating state information is identified according to the initialized pre-trained neural network strategy model to generate discrete action decisions.
10. A multi-train dynamic dispatching system based on MAPPO algorithm, characterized in that: The system comprises: an acquisition module, configured to acquire operating status information of multiple trains and identify the operating status information according to a pre-trained neural network strategy model to generate discrete action decisions; A training module is configured to generate adjustment instructions based on discrete action decisions, including speed adjustment and start / stop instructions, calculate new operating state information and compound reward feedback for the current decision step based on the adjustment instructions, and update the train environment state information to proceed to the next decision step until the end of the current round; wherein each round includes multiple decision steps, each decision step corresponds to an action decision, and each round generates a scheduling strategy, each scheduling strategy including multiple action decisions; A judgment module is used to judge whether the scheduling strategy corresponding to the current round conforms to the preset scheduling strategy based on all rewards of the current round; The first execution module is used to return to the step of obtaining the operation status information of multiple trains if the scheduling strategy corresponding to the current round does not conform to the preset scheduling strategy, until the scheduling strategy corresponding to the current round conforms to the preset scheduling strategy to dynamically coordinate and optimize the scheduling of multiple trains.
Citation Information
Patent Citations
WiFi network resource scheduling method and system based on MAPPO algorithm
CN117412323A
Train operation real-time adjustment method based on cooperative competition game
CN118722789A
Bus real-time scheduling method and system based on MAPPO reinforcement learning, and storage medium
CN119886768A
Microgrid group low-carbon optimization operation method based on improved MAPPO algorithm
CN120073853A
Centralised and decentralised multi-agent systems
WO2023213403A1