Method for adjusting running diagram of virtual marshalling heavy-load train under autonomous operation
By constructing a virtual train formation heavy-haul train timetable adjustment model and using reinforcement learning algorithms, the train formation and operation plan are dynamically adjusted, which solves the problem of operation plan disorder caused by sudden events such as temporary speed limits under autonomous operation, and improves the responsiveness and transportation efficiency of the railway transportation system.
Patent Information
- Application Number
- CN202610025226.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-09
- Publication Date
- 2026-02-06
AI Technical Summary
Existing technologies are insufficient to effectively handle emergencies such as temporary speed restrictions in autonomous heavy-haul train systems, leading to disruptions in train operation plans and impacting the operational efficiency of the railway network.
A virtual train formation heavy-haul train timetable adjustment model is constructed and transformed into a reinforcement learning framework. The model is then solved using a knowledge-guided reinforcement learning algorithm to dynamically adjust train formations and operation plans, thereby optimizing the train timetable.
It has improved the railway transportation system's ability to respond to emergencies, alleviated the pressure caused by temporary speed restrictions, met transportation demand with flexibility and reliability, and improved transportation efficiency.
Smart Images

Figure CN121469686A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of heavy-haul train operation optimization technology, and in particular to a method for adjusting the timetable of heavy-haul trains with virtual formations under autonomous operation. Background Technology
[0002] With the continuous growth of freight demand, heavy-haul railway transportation faces the challenge of capacity saturation. Autonomous virtual train formation technology offers an effective solution to this problem. Through autonomous decision-making and wireless communication, virtual train formation enables dynamic train formation, allowing it to flexibly respond to real-time freight demand without human intervention. It automatically adjusts formation and operation plans, optimizes transportation resource allocation, reduces unnecessary stops and waiting times, and significantly improves freight transportation efficiency. This autonomous virtual train formation technology not only enhances the adaptability of the railway transportation system but also provides a more efficient solution for coping with complex and ever-changing transportation environments. Under normal circumstances, trains operate autonomously according to the planned timetable. However, external factors such as severe weather, equipment failure, and track maintenance can still disrupt train operations, with temporary speed restrictions being a common disruptive factor. When trains need to slow down on certain sections due to temporary speed restrictions, the original timetable and train sequence are disrupted, leading to delays and potentially triggering a chain reaction that affects the operational efficiency of the entire railway network.
[0003] To address this challenge, a virtual train formation method for adjusting heavy-haul train timetables is needed to shorten solution time and improve adaptability in multiple scenarios, thereby achieving autonomous, efficient, and flexible train operation management. Summary of the Invention
[0004] The purpose of this application is to provide a method for adjusting the timetable of heavy-haul trains with virtual formations under autonomous operation, so as to automatically respond to emergencies such as temporary speed limits, dynamically adjust the formation and operation plan, thereby improving the system's responsiveness and transportation efficiency.
[0005] To achieve the above objectives, this application provides the following solution.
[0006] This application provides a method for adjusting the timetable of heavy-haul trains with virtual formations under autonomous operation, including: Construct a virtual train formation model for adjusting the timetable of heavy-haul trains; The heavy-haul train timetable adjustment model is transformed into a reinforcement learning framework, and a knowledge-guided reinforcement learning algorithm is used to solve the heavy-haul train timetable adjustment model to obtain a heavy-haul train timetable with virtual formations under autonomous operation.
[0007] According to the specific embodiments provided in this application, the following technical effects are disclosed: This application discloses a method for adjusting the timetable of virtual-formation heavy-haul trains under autonomous operation. First, a timetable adjustment model for virtual-formation heavy-haul trains is constructed. Then, this model is transformed into a reinforcement learning framework, and a knowledge-guided reinforcement learning algorithm is used to solve the model, resulting in a timetable for virtual-formation heavy-haul trains under autonomous operation. This application improves the complexity of adjusting the timetable of virtual-formation heavy-haul trains under autonomous operation by transforming the model into a reinforcement learning algorithm framework and using a knowledge-guided reinforcement learning algorithm to explore and optimize, ultimately obtaining a conflict-free timetable for heavy-haul trains. By flexibly adjusting the virtual-formation plan and freight demand, it significantly improves the railway transportation system's responsiveness to emergencies, alleviates the direct pressure from temporary speed restrictions, maximizes the satisfaction of transportation needs, and improves the reliability and flexibility of the supply chain. The use of a knowledge-guided reinforcement learning algorithm to handle complex train timetable optimization problems effectively addresses the issues of high problem dimensionality and large real-time computational load. By combining knowledge rules to guide reinforcement learning methods, train schedules that meet requirements can be calculated relatively quickly while ensuring solution quality. Attached Figure Description
[0008] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0009] Figure 1 This is a schematic flowchart of a method for adjusting the timetable of a heavy-haul train under autonomous operation with virtual formation, provided in one embodiment of this application. Figure 2 A schematic diagram of a deep strategic gradient algorithm framework with dual neural networks; Figure 3 This is a schematic diagram of the training framework for a multi-agent algorithm. Figure 4 A schematic diagram illustrating the selection of unloading station unloading combinations for dynamic programming; Figure 5 A schematic diagram of a knowledge-guided reinforcement learning algorithm framework; Figure 6 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0010] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0011] The purpose of this application is to provide a method for adjusting the timetable of heavy-haul trains with virtual formations under autonomous operation, which aims to automatically respond to emergencies such as temporary speed limits, dynamically adjust the formations and operation plans, thereby improving the system's responsiveness and transportation efficiency.
[0012] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0013] In one exemplary embodiment, such as Figure 1 As shown, a method for adjusting the timetable of heavy-haul trains with virtual formations under autonomous operation is provided, including the following steps.
[0014] Step 1: Construct a virtual train formation timetable adjustment model for heavy-haul trains.
[0015] As an optional implementation method, the heavy-haul train timetable adjustment model includes: objective function and constraints; The objective function aims to maximize the freight schedule fulfillment deviation considering the importance of goods and minimize the total train travel time. The constraints include: loading station departure time constraints, temporary speed-limited section track capacity constraints, dynamic unloading site allocation constraints, cargo supply and demand matching constraints, train operation constraints, parking and waiting operation time constraints, safety tracking interval constraints, arrival and departure sequence constraints, effective virtual marshalling interval constraints, marshalling sequence constraints, effective demarcation constraints for different unloading sites, track occupancy constraints, and arrival demarcation constraints for different station operation modes.
[0016] As an optional implementation, the objective function includes: (1) (2) (3) in, The objective function value; Weighting of freight schedule fulfillment deviations to account for the importance of goods; To account for deviations in freight schedule fulfillment based on the importance of the goods; Weighted according to the total train travel time; Total train travel time; A collection of trains in a virtual formation; A collection of unloading stations; For train At the unloading station Unloading sign; For unloading station The delivery time of the planned goods; For train At the station The actual arrival time; For unloading station Importance of goods demand; For train At the loading station The actual departure time; A collection of all stations; For train At the station The actual departure time; For train At the station The actual departure time; For train At the station The actual arrival time; For train At the station The actual arrival time; For train With the train At the station The actual departure time interval For train With the train At the station The actual arrival time interval of the train With the train These are two adjacent trains.
[0017] As an optional implementation, the loading station departure time constraint means that the train must ensure the completion of loading operations, and the actual departure time of the train at the loading station must be greater than the planned departure time of the train at the loading station. The loading station departure time constraint includes: (4) in, For train At the loading station The scheduled departure time.
[0018] As an optional implementation method, the track capacity constraints for temporary speed-limited sections include: (5) (6) (7) (8) in, For train At the station The actual arrival time; For train At the station The actual departure time; It is a preset positive integer; For train At the station The value of is 1 when a line is started and 0 otherwise. For train At the station The speed limit sign indicates that the train must enter the area. Entering the station during temporary speed limit periods ,but ,on the contrary, ; Stations during temporary speed limit periods The minimum transit time; Stations during normal operating hours Minimum running time; This refers to the collection of stations within a temporary speed-limited section. For train At the station The actual departure time; This is the start time of the temporary speed limit period; For train At the station Departures after the start and departures before the temporary speed limit ends; For train At the station after the temporary speed limit period begins The departure sign, if the train At the station after the temporary speed limit period begins Departure, ,on the contrary, ; For train At the station before the temporary speed limit period ends The departure sign, if the train At the station before the temporary speed limit period ends Departure, ,on the contrary, ; This is the end time of the temporary speed limit period; For train At the station The speed limit sign indicates that the train must enter the area. Entering the station during temporary speed limit periods ,but ,on the contrary, .
[0019] Specifically, formula (5) indicates that under a sudden temporary speed limit, the temporary speed limit period is extended, and the train operation in the non-speed-limited section meets the minimum interval running time. This constraint is a conditional constraint and an AND constraint, and is linearized using the Big M method. Formula (6) guarantees that if the train... On the train At the station after the temporary speed limit period begins Departure, If the value is 0, then the value is 0. Formula (7) guarantees that if the train... At the station before the temporary speed limit period ends Departure, Conversely, it is 0. Formula (8) means that when hour, Trains will not be affected by the speed limit before or after the temporary speed limit period begins.
[0020] As an optional implementation, dynamic unloading site allocation constraints include: (9) (10) in, For train At the planned unloading station The unloading sign, if the train At the planned unloading station Unloading, ,on the contrary, ; For train At the station The scheduled departure time.
[0021] Specifically, since heavy-haul trains have the characteristic of 'weak planning', after a sudden interference event, the evaluation index is not based on the minimum delay time, but rather on the timely delivery of goods. Constraint (9) indicates that trains that have passed the temporary speed limit period arrive at the unloading point according to the plan, and constraint (10) indicates that the affected trains are dynamically allocated to unloading points according to transportation needs, and each train only arrives at one station to unload.
[0022] As an optional implementation method, the goods supply and demand matching constraint includes: (11) in, For train traction mass; For unloading station The freight demand.
[0023] Specifically, formula (11) indicates that the supply of goods at the unloading site must exceed the demand for goods at that site.
[0024] As an optional implementation method, train operation constraints include: (12) (13) (14) Specifically, formula (12) indicates the departure of the train at the station before loading. Formulas (13) and (14) indicate the dynamic allocation of unloading points for the train, and whether the train departs at the station depends on the unloading point.
[0025] As an optional implementation, the parking waiting time constraint includes: (15) (16) (17) (18) in, For train At the station The planned parking operation time; For train At the station Adjusted parking waiting time; This refers to the minimum stopping time for a train at a station.
[0026] Specifically, formula (15) indicates that if the train passes through the planned operation station, the planned stopping time is met. Formulas (16) and (17) indicate the adjusted train... At the station Stopping and waiting. Formula (18) indicates that the train must meet the minimum stopping operation time constraint while stopping and waiting.
[0027] As an optional implementation, the security tracking interval constraint includes: (19) (20) in, For train At the station The value of is 1 when a line is started and 0 otherwise. For train At the station The value of is 1 when a line is started and 0 otherwise. For train At the station The departure order is earlier than the train The sign, if the train At the station The departure order is earlier than the train ,but ,on the contrary, ; For the following trains in moving block The tracking interval of the operation; For train At the station and train The composition of the train formation, if the train At the station and train If a train formation is formed, then ,on the contrary, ; To trains in virtual formation The tracking interval of the operation; For train At the station The departure order is earlier than the train The sign, if the train At the station The departure order is earlier than the train ,but ,on the contrary, ; For the following trains in moving block The tracking interval of the operation; For train At the station and train The composition of the train formation, if the train At the station and train If a train formation is formed, then ,on the contrary, ; To trains in virtual formation The tracking interval of the operation; For train At the station The arrival order was earlier than the train The sign, if the train At the station The arrival order was earlier than the train ,but ,on the contrary, ; For train At the station and train Arrival formation markers, if the train At the station and train If the arrival group is formed, then ,on the contrary, ; For train At the station The arrival order was earlier than the train The sign, if the train At the station The arrival order was earlier than the train ,but ,on the contrary, ; For train At the station and train Arrival formation markers, if the train At the station and train If the arrival group is formed, then ,on the contrary, .
[0028] Specifically, whether a train is virtually formed needs to be dynamically determined during the adjustment process. Formula (19) indicates that if the train is running in virtual formation at departure, it will track within the formation using the virtual formation tracking interval; if the train is running in moving block formation, it will track between formations using the moving block tracking interval. Formula (20) indicates that if the train is running in virtual formation at arrival, it will track within the formation using the virtual formation tracking interval; if the train is running in moving block formation, it will track between formations using the moving block tracking interval.
[0029] Specifically, arrival / departure order constraints include: (twenty one) (twenty two) (twenty three) in, For train At the station The departure order is earlier than the train The sign, if the train At the station The departure order is earlier than the train ,but ,on the contrary, .
[0030] Formula (21) represents the train and train All at the station Only after the trains have started operating can there be a departure sequence between them. Formula (22) represents the train... and All at the station Only when the trains arrive at their destinations can there be an arrival order between them. Formula (23) means that trains cannot overtake each other within a section.
[0031] Effective virtual grouping interval constraints include: (twenty four) (25) in, For train Effective grouping interval; For train Effective grouping interval.
[0032] The effective virtual formation between trains is guaranteed by the time interval. If the time interval between trains is greater than the effective formation interval, the train is disassembled. Formula (24) represents the effective virtual formation interval constraint for train departure. Formula (25) represents the effective virtual formation interval constraint for train arrival.
[0033] Grouping order constraints include: (26) Formula (26) indicates that at the same station, only the following train and the preceding train can be grouped together.
[0034] Different effective uncoding constraints at the unloading location include: (27) in, For train At the station The unloading sign, if the train At the station Unloading, ,on the contrary, ; For train At the station The unloading sign, if the train At the station Unloading, ,on the contrary, .
[0035] Formula (27) indicates that if a train is unloaded at different unloading stations, it must be effectively decoupled at that station.
[0036] Lane occupancy constraints include: (28) (29) (30) in, For train At the station Regarding the plan to occupy the track The occupancy sign, if the train At the station Occupy plan occupies track ,but Conversely, it is 0; For stock market trends; For train At the station For the actual occupation of the lane The occupancy sign, if the train At the station Occupy actual occupied lane ,but Conversely, it is 0.
[0037] Formula (28) indicates that if the adjusted train passes through the planned operation station, it will occupy the planned allocated track. Formula (29) indicates that if the train is adjusted to stop and wait, a track will be allocated to the train and a train will only occupy one track at the station. Formula (30) indicates that when two trains are adjusted to occupy the same track, the departure and arrival interval must be met.
[0038] Different station operation methods result in different arrival decoupling constraints, including: (31) (32) (33) (34) in, For train At the station The parking waiting variable; For train At the station The parking waiting variable; For train and train At the station There is only one stop sign. If the train and train At the station If only one train stops, then Conversely, it is 0; For train and train At the station Only one train is occupying the track. The sign, if the train and train At the station Only one train is occupying the track. ,but Conversely, it is 0; For train At the station For the actual occupation of the lane The occupancy sign, if the train At the station Occupy actual occupied lane ,but Conversely, it is 0; For train At the station For the actual occupation of the lane The occupancy sign, if the train At the station Occupy actual occupied lane ,but Conversely, it is 0.
[0039] Step 2: Transform the heavy-haul train timetable adjustment model into a reinforcement learning framework, and use a knowledge-guided reinforcement learning algorithm to solve the heavy-haul train timetable adjustment model to obtain a virtual train timetable under autonomous operation.
[0040] Specifically, step 2 includes the following steps.
[0041] Step 21: Transform the heavy-haul train timetable adjustment model into a reinforcement learning framework. The detailed algorithm framework is as follows... Figure 5 As shown. The left side is the knowledge guidance module, and the right side is the deep strategic gradient reinforcement learning algorithm module.
[0042] The operation of a heavy-haul train according to a timetable, departing from the loading station, passing through intermediate stations, and finally stopping at the unloading station, can be viewed as a discrete process divided according to the stations. At each station, the train's state is determined by selecting an appropriate strategy until the train arrives at the unloading station. This process can be described as a Markov decision process. The heavy-haul train timetable adjustment model is then transformed into a reinforcement learning framework using a Markov decision process, as follows: (35) in, For the stage The train's status at the station, , To enhance the number of steps in learning execution; For the stage The train's chosen action, .
[0043] Step 22: Initialize the core parameter set. This includes the following steps.
[0044] Step 221: Input two types of preset parameters: ① MADDPG (Multi-Agent Deep Deterministic Strategy) algorithm parameters: number of training iterations Maximum capacity of the experience pool Number of experience replay samples Strategy Network Learning Rate Value network learning rate Soft update parameters Discount Factor ② Train timetable optimization parameters: information on temporary speed limit scenarios Schedule Line parameters (Number of stations, number of tracks, minimum station operating time), safe tracking interval between trains, supply at loading stations, and demand at unloading stations.
[0045] Step 222: Initialize the two types of preset parameters to ensure that the parameter format matches the requirements of the algorithm and simulation environment, and obtain the initialized algorithm parameter set and train timetable optimization parameter set.
[0046] Step 23: As Figure 3 Based on a multi-agent algorithm framework, environmental interaction and experience data storage are performed. Specifically, the steps include the following.
[0047] Step 231: Based on the line parameters and station information input in Step 2, determine the initial time phase. The status of all trains in the virtual train formation .
[0048] (36) (37) in, For all trains in the virtual formation at the stage The set of states, For train In the stage The local state; Stages No. The train's location at the station, arrival time, departure time, stop signs, track occupancy signs, operation signs, unloading signs, arrival formation signs, and departure formation signs.
[0049] Step 232: Invoke the initial policy network and approximate the obtained target value through the Actor network. , Let be the policy function, representing the strategy for choosing an action given a state. For the parameter set of the Actor target network, To stabilize the target network during training, "Parameters" is a collective term for all parameters. In order to be in The state corresponding to the stage. Generate initial actions with exploration noise. .
[0050] (38) (39) in, For all trains in the virtual formation at the stage A set of actions, For train In the stage Local actions.
[0051] Step 233: Apply the action space mask to block illegal actions and obtain legal actions.
[0052] The action probabilities output by the Actor neural network are the probabilities of all actions. However, when solving the problem of adjusting the train timetable for sudden temporary speed limit virtual formations under autonomous operation, the action is restricted. Therefore, an action mask is added to the action probabilities output by the Actor neural network to block illegal actions. This ensures the selection of reasonable actions while limiting the size of the action space and accelerating the training process. The actions that need to be blocked are as follows: 1) When the loading station is not the starting station of a train, the train will not depart from that station; 2) When the train has already unloaded at the unloading station, the train will not depart from subsequent stations; 3) When the station is not an unloading station, the unloading action cannot be selected.
[0053] Step 234: Apply adjusted knowledge rules.
[0054] 1) Selection of unloading actions: During the agent's exploration process, because the agent independently selects actions, it is difficult to satisfy the supply-demand matching constraints. If a negative reward is given and the agent returns to the starting point, the agent will still explore again, violating the constraints. Another approach is to choose an action that does not satisfy the constraints and then reselect the action at this step. However, the constraints for satisfying supply-demand matching increase exponentially, making it difficult to find matching actions and significantly reducing the algorithm's efficiency. Therefore, based on the goal of local optima, a dynamic programming method is designed to select unloading station unloading combinations, thereby guiding the algorithm to explore the state space more effectively. The principle is as follows: Figure 4 As shown. The current unloading station actions are guided by demand matching rules. The selection of the method avoids the agent's blind exploration and large negative rewards that cause the loss function to fail to converge, and ensures that the constraints are met until the train reaches the destination and does not return to the starting point midway.
[0055] 2) Generation of optimal state transitions between trains: In train schedule optimization, the interactions between trains have a significant impact on overall performance; therefore, it's insufficient to focus solely on the arrival and departure times of individual trains. Instead, the states and actions of other trains must be comprehensively considered. During state transitions, the states and actions of all agents can be obtained. This information is used to detect and resolve potential conflicts between trains and to determine whether trains can be grouped together. By optimizing the relationships between trains, conflicts can be avoided, and train grouping actions can be further optimized, thereby improving the efficiency and safety of the overall adjustment plan. In this process, the arrival order of trains at the next station is determined based on the current state, and the stopping, departure, and unloading actions of trains are determined based on the selected actions. The classic optimization problem-solving methods of branch and bound and cutting planes are used to find the optimal solution in one step, accelerating the training process of reinforcement learning.
[0056] 3) Calculate the reward and obtain the status of the next stage. : The reward function determines whether the loss function converges. The reward function can be designed as a discretized objective function, distributed across each action choice. When the train hasn't reached the unloading station, the system receives a small instantaneous reward for each step forward, depending on whether it stops, thus increasing the probability of the train reaching its destination. At each step, the system calculates the departure and arrival intervals between trains. Furthermore, once the train arrives at the unloading station, the system provides a positive reward to the agent and calculates the freight plan fulfillment deviation and single-train travel time. When the currently selected action violates the constraints, the system provides a large negative reward to the agent, reducing the probability of the next selection. The specific function is as follows: (40) in, To execute each stage The obtained state The instantaneous reward obtained below; The reward is selected based on the parking action. This refers to the train departure and arrival intervals during this phase; For the train in phase The reward for reaching the unloading station is in the stage. If the constraints are violated, a negative reward is given; otherwise, the freight plan deviation and the travel time of a single train at the station are calculated. The total reward obtained; In the stage Total rewards received.
[0057] Step 235: Transfer empirical data Stored in the experience replay buffer middle. This is the initial reward.
[0058] Step 236: Let Repeat steps 231-235 until the maximum capacity of the experience replay buffer is reached, thus obtaining an experience replay buffer filled with experience data. .
[0059] Step 24: Network training and parameter update. This includes the following steps.
[0060] Step 241: From the experience replay buffer Extraction One empirical sample of data.
[0061] Step 242: As Figure 2 The diagram shown is a framework diagram of the deep deterministic policy gradient reinforcement learning algorithm. According to the algorithm flow, the Critic network is updated: (41) (42) in, For the stage The expected cumulative reward for a state-action pair; For the stage The instantaneous reward obtained after performing an action; For the parameter set The Q-value function of the Critic network; This is the parameter set for the Critic network; For state Execute action The value of the next action. The value of the next action is approximately estimated using the Actor target network. .
[0062] Step 243: Calculate the loss according to equation (43), where, This represents the loss function used when updating the Critic network, which updates the network parameters using the mean squared error loss. The number of samples; For the stage Next state Execute action The value of the action.
[0063] Step 244: Update the Actor policy network: (43) The Actor target network is used to provide the policy for the next state, while the Actor training network provides the policy for the current state. Equation (43) combines the Q-value function of the Critic training network to obtain the policy gradient of the Actor network when updating parameters. For Actor policy networks in policy The policy gradient is used to guide the Actor network parameter updates to maximize the cumulative reward. . The Q-value function of the Critic network for state-action pairs gradient, For the Actor policy function, its parameters In state The gradient represents the effect of parameter changes on the policy output.
[0064] Step 245: Soft update the target network: (44) Every set number of steps, the parameters are softly updated according to equation (44). For the parameter set of the Critic target network; This is a soft update coefficient that controls the magnitude of the target network parameter update. This is the set of parameters for the Actor online network.
[0065] Step 246: Repeat steps 241-245 until the set number of training iterations is reached. Stop training and save the final Actor policy network. .
[0066] Step 25: Generate the adjusted train timetable. This includes the following steps.
[0067] Step 251: Set the initial state As input, enter Generate legal actions for the train at the first station. .
[0068] Step 252: Obtain the state based on the train state transition calculation. .
[0069] Step 253: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require enter Generate the action for the next station. Repeat step 252 to iteratively generate arrival and departure times and train formation data for each train at different stations.
[0070] Step 254: Integrate the arrival and departure times and train formation data of all trains at various stations to form a virtual train formation operation diagram for sudden temporary speed limit scenarios.
[0071] In one exemplary embodiment, a computer device is provided, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement a method for adjusting the timetable of heavy-haul trains with virtual formations under autonomous operation.
[0072] In one exemplary embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements a method for adjusting the timetable of heavy-haul trains with virtual formations under autonomous operation.
[0073] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements a method for adjusting the timetable of heavy-load trains with virtual formations under autonomous operation.
[0074] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 6 As shown, the computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a method for adjusting the timetable of a heavy-haul train under autonomous virtual formation operation.
[0075] Those skilled in the art will understand that Figure 6 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0076] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).
[0077] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0078] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0079] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0080] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A method for adjusting the timetable of heavy-haul trains with virtual formations under autonomous operation, characterized in that, The method for adjusting the timetable of heavy-haul trains under autonomous operation and virtual formation includes: Construct a virtual train formation model for adjusting the timetable of heavy-haul trains; The heavy-haul train timetable adjustment model is transformed into a reinforcement learning framework, and a knowledge-guided reinforcement learning algorithm is used to solve the heavy-haul train timetable adjustment model to obtain a heavy-haul train timetable with virtual formations under autonomous operation.
2. The method for adjusting the timetable of heavy-haul trains under autonomous operation and virtual formation as described in claim 1, characterized in that, The heavy-haul train timetable adjustment model includes: objective function and constraints; The objective function aims to maximize the freight plan fulfillment deviation considering the importance of goods and minimize the total train travel time. The constraints include: loading station departure time constraints, temporary speed-limited section track capacity constraints, dynamic unloading site allocation constraints, cargo supply and demand matching constraints, train operation constraints, parking and waiting operation time constraints, safety tracking interval constraints, arrival and departure sequence constraints, effective virtual formation interval constraints, formation sequence constraints, effective unloading constraints for different unloading sites, track occupancy constraints, and arrival unloading constraints for different station operation modes.
3. The method for adjusting the timetable of heavy-haul trains under autonomous operation with virtual formation as described in claim 2, characterized in that, The objective function includes: ; ; ; in, The objective function value; Weighting of freight schedule fulfillment deviations to account for the importance of goods; To account for deviations in freight schedule fulfillment based on the importance of the goods; Weighted according to the total train travel time; Total train travel time; A collection of trains in a virtual formation; A collection of unloading stations; For train At the unloading station Unloading sign; For unloading station The delivery time of the planned goods; For train At the station The actual arrival time; For unloading station Importance of goods demand; For train At the loading station The actual departure time; A collection of all stations; For train At the station The actual departure time; For train At the station The actual departure time; For train At the station The actual arrival time; For train At the station The actual arrival time; For train With the train At the station The actual departure time interval For train With the train At the station The actual arrival time interval of the train With the train These are two adjacent trains.
4. The method for adjusting the timetable of heavy-haul trains under autonomous operation with virtual formation as described in claim 3, characterized in that, The loading station departure time constraints include: ; in, For train At the loading station The scheduled departure time.
5. The method for adjusting the timetable of heavy-haul trains under autonomous operation and virtual formation as described in claim 4, characterized in that, The temporary speed-limited section line capacity constraints include: ; ; ; ; in, For train At the station The actual arrival time; For train At the station The actual departure time; It is a preset positive integer; For train At the station The value of is 1 when a line is started and 0 otherwise. For train At the station The speed limit sign indicates that the train must enter the area. Entering the station during temporary speed limit periods ,but ,on the contrary, ; Stations during temporary speed limit periods The minimum transit time; Stations during normal operating hours Minimum running time; This refers to the collection of stations within a temporary speed-limited section. For train At the station The actual departure time; This is the start time of the temporary speed limit period; For train At the station Departures after the start and departures before the temporary speed limit ends; For train At the station after the temporary speed limit period begins The departure sign, if the train At the station after the temporary speed limit period begins Departure, ,on the contrary, ; For train At the station before the temporary speed limit period ends The departure sign, if the train At the station before the temporary speed limit period ends Departure, ,on the contrary, ; This is the end time of the temporary speed limit period; For train At the station The speed limit sign indicates that the train must enter the area. Entering the station during temporary speed limit periods ,but ,on the contrary, 。 6. The method for adjusting the timetable of heavy-haul trains under autonomous operation and virtual formation as described in claim 5, characterized in that, The dynamic unloading site allocation constraints include: ; ; in, For train At the planned unloading station The unloading sign, if the train At the planned unloading station Unloading, ,on the contrary, ; For train At the station The scheduled departure time.
7. The method for adjusting the timetable of heavy-haul trains under autonomous operation with virtual formation as described in claim 6, characterized in that, The supply and demand matching constraints for goods include: ; in, For train traction mass; For unloading station The freight demand.
8. The method for adjusting the timetable of heavy-haul trains under autonomous operation with virtual formation as described in claim 7, characterized in that, The train operation constraints include: ; ; 。 9. The method for adjusting the timetable of heavy-haul trains under autonomous operation with virtual formation as described in claim 8, characterized in that, The parking waiting time constraints include: ; ; ; ; in, For train At the station The planned parking operation time; For train At the station Adjusted parking waiting time; This refers to the minimum stopping time for a train at a station.
10. The method for adjusting the timetable of heavy-haul trains with virtual formation under autonomous operation as described in claim 9, characterized in that, The secure tracking interval constraint includes: ; ; in, For train At the station The value of is 1 when a line is started and 0 otherwise. For train At the station The value of is 1 when a line is started and 0 otherwise. For train At the station The departure order is earlier than the train The sign, if the train At the station The departure order is earlier than the train ,but ,on the contrary, ; For the following trains in moving block The tracking interval of the operation; For train At the station and train The composition of the train formation, if the train At the station and train If a train formation is formed, then ,on the contrary, ; To trains in virtual formation The tracking interval of the operation; For train At the station The departure order is earlier than the train The sign, if the train At the station The departure order is earlier than the train ,but ,on the contrary, ; For the following trains in moving block The tracking interval of the operation; For train At the station and train The composition of the train formation, if the train At the station and train If a train formation is formed, then ,on the contrary, ; To trains in virtual formation The tracking interval of the operation; For train At the station The arrival order was earlier than the train The sign, if the train At the station The arrival order was earlier than the train ,but ,on the contrary, ; For train At the station and train Arrival formation markers, if the train At the station and train If the arrival group is formed, then ,on the contrary, ; For train At the station The arrival order was earlier than the train The sign, if the train At the station The arrival order was earlier than the train ,but ,on the contrary, ; For train At the station and train Arrival formation markers, if the train At the station and train If the arrival group is formed, then ,on the contrary, .