Heavy haul train virtual coupling rolling scheduling optimization method

CN122607394APending Publication Date: 2026-08-21SHUOHUANG RAILWAY DEV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610684061.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-18
Publication Date
2026-08-21

AI Technical Summary

Technical Problem

然而,在虚拟联挂条件下,重载列车间距离显著缩小,实时调度需确保安全间隔不被破坏、同时尽可能提高发车频率以缓解货物拥堵,此外还需兼顾晚点风险和能耗

Benefits of technology

[0018]上述重载列车虚拟联挂滚动调度优化方法、装置、计算机设备、计算机可读存储介质和计算机程序产品,通过获取当前时刻的调度环境数据构建状态图,该状态图包含列车节点和资源节点以及表征节点间交互关系的边,能够精确描述虚拟联挂场景下多列重载列车与轨道资源的占用关系以及列车之间的追踪或编组耦合关系,利用预训练的图神经网络模型对状态图中各节点的节点特征进行至少一层的传播,从而提取得到可以全面表征虚拟联挂场景的全局状态特征,基于预训练的策略网络根据全局状态特征生成当前时刻的调度动作决策,按照生成的调度动作决策对重载列车进行当前时刻的虚拟联挂调度控制,并滚动到下一时刻重复进行虚拟联挂调度控制,可以确保虚拟联挂调度控制能够精准匹配实时的调度环境,从而提升了重载列车虚拟联挂运行场景下的运输效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122607394A_ABST
    Figure CN122607394A_ABST
Patent Text Reader

Abstract

The application relates to a heavy-haul train virtual coupling rolling scheduling optimization method and device, computer equipment, a storage medium and a program product. The method comprises the following steps: acquiring scheduling environment data of a heavy-haul train at a current time, and constructing a state graph at the current time based on the scheduling environment data; nodes in the state graph comprise train nodes and resource nodes, and edges in the state graph are used for representing the interaction relationship between the nodes; based on a pre-trained graph neural network model, at least one layer of propagation is performed on the node features of the nodes in the state graph, and global state features are obtained; based on a pre-trained policy network, a scheduling action decision is generated at the current time according to the global state features; according to the scheduling action decision, virtual coupling scheduling control is performed on the heavy-haul train at the current time, and the virtual coupling scheduling control is rolled to the next time until the scheduling task for the heavy-haul train is completed. The method can improve the transportation efficiency of the heavy-haul train virtual coupling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of train control technology, and in particular to a method, apparatus, computer equipment, computer-readable storage medium, and computer program product for optimizing virtual coupling and rolling scheduling of heavy-haul trains. Background Technology

[0002] With the increase in freight volume on heavy-haul trains and the development of intelligent train control technology, virtual coupling has become a new concept for improving transport capacity and flexible scheduling. Virtual coupling refers to the use of wireless communication and collaborative control to enable multiple heavy-haul trains to form a virtual trainset and run closely together without physical connection. This technology can shorten the tracking interval of heavy-haul trains and improve the transport capacity of the line without significant infrastructure modifications. However, under virtual coupling conditions, the distance between heavy-haul trains is significantly reduced. Real-time scheduling needs to ensure that the safe interval is not violated, while maximizing the departure frequency to alleviate freight congestion. In addition, the risk of delays and energy consumption must also be considered.

[0003] Currently, traditional scheduling methods are unable to respond promptly to dynamic changes in freight volume and emergencies, resulting in low transportation efficiency for virtual coupling of heavy-haul trains. Summary of the Invention

[0004] Therefore, it is necessary to provide a method, apparatus, computer equipment, computer-readable storage medium, and computer program product for optimizing the virtual coupling and rolling scheduling of heavy-haul trains, which can improve transportation efficiency, in response to the above-mentioned technical problems.

[0005] Firstly, this application provides a method for optimizing the virtual coupling and rolling scheduling of heavy-haul trains, including:

[0006] Obtain the scheduling environment data of heavy-load trains at the current moment, and construct a state diagram at the current moment based on the scheduling environment data; the nodes in the state diagram include train nodes and resource nodes, and the edges in the state diagram are used to represent the interaction relationships between nodes;

[0007] Based on a pre-trained graph neural network model, the node features of each node in the state graph are propagated at least one layer to obtain global state features.

[0008] Based on a pre-trained policy network, scheduling action decisions for the current moment are generated according to global state features;

[0009] Based on the dispatching action decision, virtual coupling dispatching control is carried out for heavy-haul trains at the current moment, and then rolled to the next moment for virtual coupling dispatching control, until the dispatching task for heavy-haul trains is completed.

[0010] Secondly, this application also provides a virtual coupling and rolling scheduling optimization device for heavy-haul trains, comprising:

[0011] The state graph construction module is used to obtain the scheduling environment data of heavy-load trains at the current moment and construct the state graph at the current moment based on the scheduling environment data. The nodes in the state graph include train nodes and resource nodes, and the edges in the state graph are used to represent the interaction relationship between nodes.

[0012] The feature propagation module is used to propagate the node features of each node in the state graph at least one layer based on the pre-trained graph neural network model to obtain the global state features.

[0013] The action decision module is used to generate scheduling action decisions for the current moment based on the global state features and a pre-trained policy network.

[0014] The dispatch control module is used to perform virtual coupling dispatch control for heavy-haul trains at the current moment according to the dispatch action decision, and roll over to the next moment for virtual coupling dispatch control until the dispatch task for heavy-haul trains is completed.

[0015] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the method provided in the first aspect above.

[0016] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method provided in the first aspect above.

[0017] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps in the method provided in the first aspect above.

[0018] The aforementioned method, apparatus, computer equipment, computer-readable storage medium, and computer program product for optimizing virtual coupling and rolling scheduling of heavy-haul trains construct a state graph by acquiring the current scheduling environment data. This state graph includes train nodes, resource nodes, and edges representing the interaction relationships between nodes, accurately describing the occupancy relationship between multiple heavy-haul trains and track resources, as well as the tracking or formation coupling relationship between trains in the virtual coupling scenario. A pre-trained graph neural network model is used to propagate the node features of each node in the state graph through at least one layer, thereby extracting global state features that comprehensively represent the virtual coupling scenario. Based on the pre-trained policy network, a scheduling action decision for the current moment is generated according to the global state features. The generated scheduling action decision is used to perform virtual coupling scheduling control on the heavy-haul trains at the current moment, and this process is repeated until the next moment. This ensures that the virtual coupling scheduling control can accurately match the real-time scheduling environment, thereby improving the transportation efficiency in the virtual coupling operation scenario of heavy-haul trains. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a flowchart illustrating a method for optimizing the virtual coupling and rolling scheduling of heavy-haul trains in one embodiment.

[0021] Figure 2 This is a schematic diagram of the policy network update process in one embodiment;

[0022] Figure 3 This is a structural block diagram of a virtual coupling and rolling scheduling optimization device for heavy-load trains in one embodiment;

[0023] Figure 4 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0024] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0025] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.

[0026] The virtual coupling and rolling scheduling optimization method for heavy-haul trains provided in this application can be applied to computer equipment. Specifically, it can be executed independently by a terminal or server, or jointly by both. The computer equipment can be a server deployed in a dispatch center, an edge computing node, or an onboard controller on the train. The terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, drones, low-altitude aircraft, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, and projection devices. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted displays. Head-mounted displays can be virtual reality (VR) devices, augmented reality (AR) devices, and smart glasses. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0027] In one exemplary embodiment, such as Figure 1 As shown, a virtual coupling and rolling scheduling optimization method for heavy-haul trains is provided. In this embodiment, the method is applied to computer equipment as an example for illustration, including the following steps 102 to 108. Wherein:

[0028] Step 102: Obtain the scheduling environment data of the heavy-haul train at the current moment, and construct a state diagram at the current moment based on the scheduling environment data; the nodes in the state diagram include train nodes and resource nodes, and the edges in the state diagram are used to represent the interaction relationship between nodes.

[0029] Heavy-haul trains are used in railway transportation to transport bulk goods (such as coal, ore, and steel), and are typically long train formations consisting of multiple freight cars pulled by high-powered locomotives. Virtual coupling refers to the use of wireless communication and cooperative control technologies to allow multiple heavy-haul trains traveling in succession to form a logically closely following virtual formation without physical couplers. Within this virtual formation, the following train can receive real-time operating status information (such as speed, acceleration, and position) from the preceding train and automatically adjust its own status to maintain a convoy distance smaller than the traditional safety interval.

[0030] The current moment is the starting point for triggering a new scheduling decision calculation and execution. It is a discrete decision node within a rolling optimization scheduling cycle. The determination of the current moment can be based on various methods, such as a fixed time interval or triggered by a specific event (e.g., a train entering a specific area or receiving a new train request). The computer equipment can obtain the latest environmental state at the current moment and calculate and generate new scheduling actions to guide train operation over the next period. For example, if a decision is made every 5 seconds, at the current moment of 09:00:00, the computer equipment can calculate a scheduling instruction based on the data at that time; after advancing 5 seconds, 09:00:05 becomes the new current moment, and the computer equipment can perform calculations and decisions again.

[0031] Dispatch environment data is a collection of raw data reflecting the objective physical environment (railway lines, signaling equipment, train status, etc.) and subjective dispatching rules (such as timetables and priorities) of heavy-haul trains. Dispatch environment data can include real-time data streams or database snapshots from various sensors (such as track circuits and transponders), control systems (such as interlocking systems and train dispatching systems), and management information systems. In some embodiments, dispatch environment data may specifically include at least one of the following: track topology (length of each section, connection relationships), precise location, speed, acceleration, direction of travel, train length, load, train number, scheduled time, signal status, turnout position, and occupancy / idle status of each section.

[0032] A state diagram is a structured mathematical model based on graph theory that can be processed and learned by computers. It is the result of abstracting, reducing dimensionality, and modeling relationships within scheduling environment data. State diagrams can be used to accurately and concisely describe the key elements and their interdependencies in the scheduling environment at the current moment. A state diagram can be a binary tuple G(V, E) containing a set of nodes V and a set of edges E. For example, in a station throat area during peak hours, the state diagram contains 3 train nodes (A, B, C) and 5 track segment nodes (T1 to T5). When train A occupies segment T2, an "occupancy" edge is established between node A and node T2; when train B follows train A, an edge representing a tracking constraint is established between node A and node B. This state diagram describes the static distribution and dynamic relationships of the three trains and track resources at the current moment.

[0033] In a state diagram, a node is an abstract representation of a specific object with independent attributes and states in the scheduling environment. Each node can be attached with one or more feature vectors or labels to describe the object's various attributes (such as position, speed, priority, etc.) at the current moment. Nodes include train nodes and resource nodes. A train node represents a physically loaded train that is running, waiting to depart, or has just completed a task in the scheduling scenario. Resource nodes represent fixed infrastructure or immovable objects in the scheduling scenario that are requested, occupied, or released by trains, such as track section nodes (a section of track, a track within a station), platform nodes, turnout nodes, signal nodes, etc.

[0034] In a state diagram, an edge connects two nodes to represent a specific, real-time interaction between them. This interaction describes the dynamic, mutually influential, or strictly rule-bound relationship between the two nodes. Interactions can include, but are not limited to, occupancy, occupancy requests, train following, virtual coupling, uncoupling, overtaking, and control relationships such as locking and releasing between trains and resources. The meaning of an edge depends on the types of the two nodes it connects, encompassing operational rules and physical constraints. For example, an edge connecting "Train Node A" and "Section Node T1" with the attribute "occupancy" means "Train A is currently occupying section T1"; another edge connecting "Train Node A" and "Train Node B" with the attribute "virtual coupling following" means "Train B is running as a virtual train following Train A."

[0035] Optionally, the computer equipment can acquire the current dispatching environment data of heavy-haul trains. For example, the computer equipment can read the dispatching environment data from the real-time database or memory cache of the railway transportation management system; the computer equipment can also extract the dispatching environment data from real-time data streams subscribed to or requested by external systems such as trackside sensors (e.g., track circuits, transponders), signal interlocking systems, and automatic train protection systems via dedicated networks such as industrial Ethernet and 5G (5th Generation Mobile Communication Technology). The dispatching environment data can be a structured list of tuples containing each train's ID, precise location, speed, direction of travel, line section, and planned departure time. Based on the dispatching environment data, the computer equipment can construct a state graph at the current moment to transform the raw data into a graph structure. In some embodiments, the computer device can parse the scheduling environment data and identify all "participants" that need to be scheduled. For each heavy-load train that is running or waiting to be scheduled, the computer device can construct a corresponding train node. For each key resource such as track section, platform, and signal, a corresponding resource node can be constructed. The computer device can dynamically create edges between nodes based on the state and relationship between all objects at the current moment as described in the scheduling environment data, thereby obtaining the state graph at the current moment.

[0036] For example, if the data shows that train A is located within interval T1, the computer equipment can draw a directed or undirected edge between the node of train A and the resource node of interval T1, and mark the relationship type as "occupancy". If the data shows that train B is following train A in virtual coupling mode, the computer equipment can draw an edge between node B and node A, and mark the relationship as "virtual coupling". The computer equipment can combine all the extracted nodes and edges into a graph data structure, which fully describes the dynamic relationships and constraints between each train, each resource, and them at the current moment, thus obtaining the state graph at the current moment.

[0037] Step 104: Based on the pre-trained graph neural network model, propagate the node features of each node in the state graph at least one layer to obtain the global state features.

[0038] Graph neural network (GNN) models are deep learning models used to process graph-structured data. These models are pre-trained to learn how to extract high-level features relevant to scheduling tasks from the state graph. Structurally, a GNN model can consist of at least one propagation layer, where each layer exchanges information between nodes through a message-passing mechanism. For example, each node can receive features from its neighbors and update its own feature representation using aggregation functions (such as weighted summation) and nonlinear transformations. In some embodiments, the GNN model can employ a Graph Convolutional Network (GCN) or a Graph Attention Network (GAT), where the weights of neighboring nodes can be dynamically adjusted based on edge features or the attention mechanism.

[0039] Propagation is the process of information transfer and aggregation between adjacent nodes in a graph neural network model. Each propagation layer completes the operation of a node receiving information from its neighbors and updating its own representation. When the number of propagation layers is L, the final feature of each node can theoretically contain information from its neighbors within a range of L hops. In some embodiments, the computer device can, based on preset adjacency relationships and edge features, perform weighted aggregation of the features of its neighboring nodes in the current layer for each node (the weight normalization coefficients are determined by the edge features, such as using degree normalization or attention values), and then perform linear transformation and nonlinear activation with the features of the current layer itself to obtain the new features of the node in the next layer. At least one layer of propagation means that the model performs such message passing at least once to ensure that the node features can obtain information from its direct neighbors; more layers can capture topological relationships over greater distances.

[0040] Node features are numerical vectors associated with each node in the state graph, representing the node's state attributes at the current moment. Initial node features can be extracted from scheduling environment data and may include at least one of the following: current train speed, location coordinates (or odometer markers), remaining running time, whether it is in a waiting state, whether it is performing a coupling operation, planned departure time, and accumulated delay time. Initial resource node features may include at least one of the following: whether the section is occupied, remaining available length, current signal display status, and whether departure is permitted. After propagation through each layer of the graph neural network model, node features gradually incorporate information from neighboring nodes, thus updating into an embedding vector containing contextual semantics.

[0041] Global state features can be obtained based on the final features of all nodes in the state graph after multi-level propagation. For example, they can be obtained by reading out the final features of all nodes in the state graph after multi-level propagation. Reading out can involve averaging, summing, or taking the maximum value of all node features, or by weighted summation using a global attention mechanism. Global state features not only retain the state information of all nodes but also incorporate spatial topology and behavioral semantics through graph neural network propagation, thus comprehensively describing the overall situation of the current scheduling scenario, including train distribution, resource occupation conflict risk, and line congestion level.

[0042] For example, a computer device can extract initial features of each node and initial features of the edges from the state graph at the current moment. The data source for these initial features is real-time measurements or planned values ​​in the scheduling environment data, such as the position and speed reported by the train's onboard system, and the line status provided by the signaling system. The computer device can organize these initial features into a node feature matrix (the number of rows represents the number of nodes, and the number of columns represents the feature dimension), an edge feature matrix (stores the features of each edge, optional), and an adjacency matrix (indicating the connection relationships between nodes), and input them into a pre-trained graph neural network model.

[0043] The pre-trained graph neural network model has been trained on a large amount of historical scheduling data, and its model parameters have been fixed or optimized to be stable. The computer device invokes this graph neural network model and performs at least one layer of neural propagation on the input state graph. In each layer of propagation, for each node i in the state graph, the computer device finds all its neighboring nodes j according to the adjacency matrix and calculates aggregation weights based on the edge features. The computer device calculates the aggregation value of the neighbor messages for node i in that layer, which can be the sum of the neighbor node features after weight scaling. The computer device can combine the features (or initial features) of node i in this layer with the aggregation value to obtain new features for node i. For the first layer of propagation, the initial features of this layer are the initial features of that node in the state graph; for subsequent layers, the initial features of this layer are the output of the previous layer. After completing propagation for all preset layers, the computer device can obtain the target features of each node. The computer device can read out the target features of each node to obtain the global state features.

[0044] Step 106: Based on the pre-trained policy network, generate the scheduling action decision for the current moment according to the global state features.

[0045] The policy network can include a pre-trained machine learning model. It maps the current global state features to a set of scheduling actions or action probability distributions. The input to the policy network is the global state features, and its output can be a probability vector of discrete actions (e.g., through a Softmax layer) or the value of continuous actions (e.g., through a linear output layer), depending on how the scheduling actions are represented. For example, the policy network can consist of three fully connected layers. Its input layer is a 128-dimensional global state feature, the hidden layers are activated using ReLU (Rectified Linear Unit), and the output layer has 10 neurons corresponding to 10 discrete scheduling actions (e.g., departure, 5-second delay, 10% acceleration, 10% deceleration, establishing a virtual coupling, disengaging a virtual coupling, etc.). The output is then converted into the selection probability of each action using Softmax, thus obtaining the scheduling action decision.

[0046] Scheduling action decisions are a set of instructions generated by the policy network under given global state characteristics, used to guide specific scheduling operations at the current moment. Scheduling action decisions can include various control elements involving multiple heavy-haul trains and track resources, such as: determining whether to immediately depart or delay the departure of trains that have not yet departed, and the delay duration; issuing speed adjustment instructions (acceleration, deceleration, or maintaining speed) to trains currently in operation; and setting operations to establish or disengage virtual coupling relationships for trains in virtual coupling formations. Scheduling action decisions can be encapsulated as structured data packets, such as a multidimensional numerical vector (discrete or continuous), or a set of ordered instruction codes, each corresponding to a control action.

[0047] Optionally, the computer device can feed global state features into a pre-trained policy network, which has been trained on a large amount of scheduling experience data and whose weight parameters have been optimized. The policy network performs forward computation and outputs one or more numerical results representing scheduling actions. For a discrete action space, the output of the policy network can be a selection probability vector of each candidate action, and the computer device can then sample the probabilities or directly select the action with the highest probability as the scheduling action decision. For a continuous action space, the output of the policy network can be directly an action vector, which can be used as the scheduling action decision.

[0048] For example, the global state feature can be a 128-dimensional floating-point vector. The computer device inputs this vector into a policy network trained using the Proximal Policy Optimization (PPO) algorithm. This policy network is pre-trained, and its output layer is a 10-dimensional discrete action probability distribution. After obtaining the probabilities through Softmax, the computer device uses a greedy strategy to select action number 7, which has the highest probability. This number corresponds to the combined action of "establishing a virtual coupling between train B and train A and simultaneously delaying the departure of train C by 5 seconds." The computer device can then use action number 7 as the scheduling action decision for the current moment.

[0049] Step 108: According to the dispatching action decision, perform virtual coupling dispatching control for the heavy-haul train at the current moment, and roll over to the next moment to perform virtual coupling dispatching control until the dispatching task for the heavy-haul train is completed.

[0050] Virtual coupling and dispatching control refers to the actual control actions implemented by computer equipment on heavy-haul trains based on dispatching action decisions. For example, computer equipment can translate dispatching action decisions into standardized train control commands (e.g., via radio block centers or vehicle-to-ground communication protocols) and send them to the onboard system of the corresponding train. The onboard system then adjusts the train's traction, braking, cruising, or coupling mode accordingly. The controlled objects of virtual coupling and dispatching control can include, but are not limited to, departure times, operating speeds, start-up and stopping times, and coupling status switching. Dispatch tasks can cover the dispatching needs of all heavy-haul trains within a specific time period, or the complete transport of a specific group of trains from origin to destination. Dispatch tasks can include pre-set plans, such as "safely dispatching all 20 waiting heavy-haul trains and arriving at the designated section within two hours of the morning peak," or "completing the tracking operation of a group of virtually coupled trains on a certain line." The specific content of the dispatching task can be defined by the operator and input into the system. During the dispatching process, the computer equipment continuously checks whether the termination conditions are met (e.g., all trains have been dispatched, the time window has ended, or all trains have arrived at their destination). Once the conditions are met, the rolling process terminates, thus completing the execution of the dispatching task.

[0051] For example, computer equipment can translate scheduling action decisions into executable scheduling control instructions to perform virtual coupling scheduling control at the current moment according to the scheduling action decisions. In some embodiments, computer equipment can invoke a communication interface to send scheduling control instructions to the onboard equipment of each heavy-haul train according to standard protocols of the railway signaling system (e.g., through a radio block center or train control management system). The scheduling control instructions can be generated based on the scheduling action decisions, the content of which may include, but is not limited to, departure instructions, target speeds, acceleration limits, activation or deactivation of virtual coupling modes, etc. The execution of each scheduling control instruction will change the operating state of the train, generating new position, speed, and resource occupancy information.

[0052] After completing the virtual train coupling and scheduling control for the current moment, the computer equipment can move to the next decision moment (e.g., a fixed time step, such as 1 second or 5 seconds). At the new moment, the computer equipment again acquires the scheduling environment data (including updated states after instruction execution), reconstructs the state graph, executes the graph neural network model propagation, obtains new global state features, and generates new scheduling action decisions again through the policy network, then executes a new round of virtual train coupling and scheduling control. This virtual train coupling and scheduling control process continuously loops, forming a rolling optimization scheduling loop. The computer equipment can determine whether the scheduling task has been completed at each step. For example, if the scheduling task is "to safely transport a specified train formation from station A to station B," the computer equipment will check in each decision cycle whether all trains have arrived at station B and come to a complete stop; if the task includes a time window (e.g., "to guarantee the maximum number of departures within 30 minutes"), the computer equipment will stop the loop when the time limit is reached. Once the task completion conditions are met, the computer equipment terminates the rolling scheduling process; otherwise, the loop of virtual train coupling and scheduling control continues.

[0053] For example, consider the scheduling task "Three trains (A, B, and C) will virtually couple and depart from the station, then track each other to the next station 5 kilometers ahead." At the first decision time t0, the computer generates a scheduling action decision: "A departs immediately, B delays for 30 seconds, and C delays for 60 seconds." The computer immediately sends a departure instruction to A, and A begins its journey; it sends waiting instructions to B and C. At time t1 (10 seconds later), the computer collects new state data and finds that A is 500 meters from the starting point, while B is still in the station. The policy network generates a new decision: "B departs in 5 seconds, and a target speed is set to keep the interval between B and A within 5 seconds." The computer sends a pre-departure setting to B. Then, at time t2, the computer may perform similar control on C. This process repeats until all three trains reach the next station, at which point the computer detects task completion and stops rolling.

[0054] In the aforementioned method for optimizing the virtual coupling and rolling scheduling of heavy-haul trains, a state graph is constructed by acquiring the scheduling environment data at the current moment. This state graph includes train nodes, resource nodes, and edges representing the interaction relationships between nodes. It can accurately describe the occupancy relationship between multiple heavy-haul trains and track resources, as well as the tracking or formation coupling relationship between trains in the virtual coupling scenario. A pre-trained graph neural network model is used to propagate the node features of each node in the state graph at least one layer, thereby extracting global state features that can comprehensively represent the virtual coupling scenario. Based on the pre-trained policy network, a scheduling action decision for the current moment is generated according to the global state features. The virtual coupling scheduling control of heavy-haul trains is performed according to the generated scheduling action decision, and the virtual coupling scheduling control is repeated at the next moment. This ensures that the virtual coupling scheduling control can accurately match the real-time scheduling environment, thereby improving the transportation efficiency in the virtual coupling operation scenario of heavy-haul trains.

[0055] In an exemplary embodiment, based on a pre-trained graph neural network model, at least one layer of propagation is performed on the node features of each node in the state graph to obtain global state features, including: extracting initial features of each node and each edge in the state graph based on the pre-trained graph neural network model; performing at least one layer of propagation on the node features of each node based on the initial features of each node and each edge to obtain target features of each node; wherein, in the propagation at each layer, for each node, the initial features of the target node in the next layer are determined based on the initial features of the target node in this layer, the initial features of the neighboring nodes of the target node in this layer, and the edge features between the target node and the neighboring nodes; and the global state features are aggregated based on the target features of each node.

[0056] In a graph neural network, initial features are data representing the original or intermediate states of each node or edge in the graph at the input stage of a specific processing layer. For example, before inputting the state graph into the first layer of the graph neural network, the computer device can first assign an initial embedding vector to each node and each edge. This embedding vector is the result of numerically encoding the original attributes of the node or edge (such as the position coordinates, speed, and scheduled departure time of a train; the length and occupancy status of a track section; and the type of edge, such as "occupancy relationship" or "tracking relationship"). This initial vector is the starting point for the propagation of the first layer, i.e., the initial features of the first layer. In a multi-layer graph neural network, for any Lth layer other than the first layer, its initial features are the target features output by the previous layer (L-1th layer) after propagation and transformation; that is, the input of the Lth layer is the output of the L-1th layer.

[0057] The target feature is the node feature output by the graph neural network after it has completed all preset layers of propagation. The target feature is the result of the initial features after multiple layers of message passing, aggregation, and nonlinear transformation. The target feature integrates the structural and state information of the node itself and its multi-hop neighbor nodes, and is a high-level, global representation of the node's state and role in the entire scheduling scenario.

[0058] Neighbor nodes are other nodes directly connected to the currently targeted node via an edge. Neighbor nodes are the direct source of message passing, and the current node enriches its own representation by aggregating the features of its neighbors. For example, in a state graph, "train node A" is connected to "track section node R1" via an "occupancy edge" and simultaneously connected to the preceding "train node B" via a "tracking edge." For node A, its neighbor set includes R1 and B. When updating the features of A, the computer device aggregates the feature information of R1 and B. Edge features describe the attribute data of the edges connecting two nodes in the state graph, characterizing the type, strength, or state of the relationship between the nodes.

[0059] For example, a computer device can invoke a pre-trained graph neural network model. The first layer of the graph neural network model may include an input embedding layer. The computer device can use this input embedding layer to map the raw attribute data of each node and each edge in the state graph to a low-dimensional dense vector space, thereby generating initial features for each node and each edge. After obtaining the initial features of all nodes and edges, the computer device can initiate the propagation process of the graph neural network, which will be repeated at least one layer. After propagation through a preset L layers, the computer device uses the features output by the last propagation layer as the target features of each node. These target features contain the structural information and neighbor information of the node within the L-hop range.

[0060] In each layer of propagation, the computer device traverses every node in the state graph, performing aggregation and update operations on the currently targeted node. Specifically, the computer device finds all neighboring nodes of the current node, obtains their initial features for this layer, and also obtains the edge features connecting the current node to each neighboring node. The computer device can aggregate the features and edge features of the neighboring nodes according to a predefined aggregation function, then combine the aggregated information with the current node's own features, and through a nonlinear transformation, obtain the initial features of the node in the next layer. Once all nodes have completed this operation, the propagation for that layer ends, and the propagation for the next layer begins.

[0061] In some embodiments, during propagation at each layer, for any node i, the computer device may perform the following operations: determine the feature vector of node i at the start of propagation at this layer, i.e., the initial feature of the target node at this layer; determine the set of all neighboring nodes j of node i, and obtain the feature vector of each neighboring node j at this layer, i.e., the initial feature of the neighboring node at this layer; obtain the feature vector of the edge connecting node i and each neighboring node j, i.e., the edge feature; based on this information, the computer device calculates the feature vector of node i at the next layer using a predefined function F. The function F may contain two parts: one is to aggregate the information from the neighbors, and the other is to fuse the aggregated information with the node's own features. The calculated result is the initial feature of node i at the next layer.

[0062] After completing propagation across all layers and obtaining the target features of all nodes, the computer device can integrate this scattered node-level information into a unified vector that describes the entire scheduling scenario, i.e., the global state feature. In some embodiments, the computer device can apply a global aggregation function to pool the target features of all nodes. For example, it can take the average of all node features to obtain the global average feature, take the maximum value of all node features in each dimension to obtain the global maximum feature, or use an attention-based aggregation method to assign different importance weights to different nodes and then perform a weighted summation. The resulting fixed-dimensional vector is the global state feature. The global state feature condenses the structural information and node state information of the entire state graph and can be used as input to the subsequent policy network to generate scheduling decisions.

[0063] In this embodiment, through the multi-layer propagation mechanism of the graph neural network model, the features of each node can spread and interact with its neighboring nodes along the edges, so that the final aggregated global state features can comprehensively and dynamically reflect the situation of the entire scheduling scenario, thereby improving the transportation efficiency of virtual coupling scheduling control based on global state features.

[0064] In an exemplary embodiment, determining the initial features of the target node in the next layer based on the initial features of the target node in this layer, the initial features of the target node's neighboring nodes in this layer, and the edge features between the target node and its neighboring nodes includes: determining each weight adjustment factor based on the edge features between the target node and each neighboring node; and determining the initial features of the target node in the next layer based on the initial features of the target node in this layer, the initial features of the target node's neighboring nodes in this layer, and each weight adjustment factor.

[0065] The weight adjustment factor is a parameter calculated from the edge features between the current node and a neighboring node during message propagation in the graph neural network model. It quantifies the influence of the neighboring node's information on the current node's feature update. The weight adjustment factor can be a scalar (representing a single weight coefficient) or a vector (representing different weights based on feature dimensions). It adjusts the contribution ratio of information from different neighbors, enabling the graph neural network model to dynamically select neighbors based on the interaction relationships reflected by the edges when aggregating neighbor features. In some embodiments, the original features corresponding to each edge (e.g., edge type encoding, continuous attribute vectors, etc.) can be extracted from the current state graph. This is then processed through a weight calculation subnetwork or a predefined mapping function, such as linear transformation with normalization, cosine similarity, or ReLU activation followed by softmax normalization, to obtain the weight adjustment factor. The weight adjustment factor can be used to weight and sum neighbor features, combined with the current node's own features, and then processed by a nonlinear activation function to generate the node's initial features in the next layer.

[0066] Optionally, the computer device can acquire the state graph of the current layer (e.g., layer l), where each node i has a corresponding initial feature. Each edge connecting node i and its neighbor node j has a corresponding edge characteristic. Edge features can include distance, resource occupancy status, virtual attachment flags, track type, etc., which are generated based on scheduling environment data during state graph construction. For each node i, the computer device traverses all its neighboring nodes j∈N(i) and extracts the corresponding edge features. The computer device uses a pre-defined weight calculation unit (e.g., a trainable single-layer neural network with softmax normalization) to calculate the weights. By performing mapping, the weight adjustment factor is obtained. Edge features It is a direct decision Factors. After processing the neighbor edge features of all node i, the computer equipment obtains the weight adjustment factor corresponding to each pair of node-neighbors.

[0067] For node i, the computer device obtains the initial features of its current layer l. and the initial features of each neighbor j in the current layer And determine the node i related to (i is fixed, j traverses neighboring nodes), the computer device can synthesize the initial features. Initial features and weighting adjustment factor We obtain the initial features of node i in the next layer l+1. In some embodiments, the computer device can assign each neighbor's characteristics Multiply by the corresponding weight adjustment factor (If α is a scalar, multiply directly; if it is a vector, multiply and add the corresponding dimensions.) Then sum (or average) all the weighted neighbor features to obtain an aggregate vector, which is then combined with the initial features of node i itself. By combining these features, such as through addition followed by linear transformation and nonlinear activation functions, the initial features of node i in the next layer (layer l+1) can be obtained. After traversing each node in layer l and obtaining the initial features of each node in the next layer (layer l+1), propagation processing can be performed in the next layer until the propagation ends.

[0068] In this embodiment, the weight adjustment factor is determined based on the edge characteristics between the target node and its neighboring nodes, which can accurately reflect the strength of various relationships in train scheduling. Furthermore, the weight adjustment factor can be dynamically changed according to the edge characteristics, thereby improving the feature expression capability and ensuring the accuracy of the global state features.

[0069] In one exemplary embodiment, such as Figure 2 As shown, the virtual coupling and rolling scheduling optimization method for heavy-haul trains also includes strategy network update processing, specifically steps 202 to 206. Wherein:

[0070] Step 202: Obtain each current trajectory obtained from online sampling, and determine the first loss based on each current trajectory; the current trajectory is used to record the complete scheduling process of virtual linkage scheduling control for the corresponding scheduling task through the policy network.

[0071] Online sampling is the process of collecting real-time data on the execution of scheduling tasks generated by the current policy network during its interaction with the environment. Online sampling provides the policy network with the most recent experience reflecting its current level, enabling it to promptly perceive the latest changes in the environment and explore better actions. For example, in a virtual coupling scenario of heavy-haul trains, a computer can command the current policy network to handle a peak-hour scheduling task involving three trains, starting from time zero until all trains reach their destination or the time period ends. Throughout this process, the agent outputs an action based on the graph state, the environment returns a reward, and the process moves to the next state. This series of interactive data constitutes a trajectory obtained through online sampling. The current trajectory refers to the sequence of states, actions, rewards, and next states obtained through online sampling, recording the entire process of a complete scheduling task. Each current trajectory corresponds to a specific scheduling task, such as the entire process of all heavy-haul trains waiting to depart within a peak hour, from the start of departure to the completion of all departures and operations.

[0072] The trajectory can be represented as a time-step sequence { },in It is the state diagram feature at step t. It is a policy network based on Output scheduling action decisions (such as deciding whether a train should depart, speed adjustment value, coupling instructions, etc.). The immediate reward provided by the environment (e.g., based on conflict, energy saving, departure rhythm, etc.) is used to accumulate the total reward of the trajectory to evaluate overall performance. The completeness of the current trajectory is reflected in its coverage of all decision points of a scheduling task from its initial state to its final state, rather than fragmented segments. In heavy-haul train scheduling, a current trajectory can record the actual departure time, speed changes, conflict occurrences, and final transport volume and energy consumption results of three virtual coupled trains (A, B, and C). The first loss is calculated using a preset loss function based on each current trajectory obtained through online sampling. It measures the difference between the current policy's performance on this newly acquired data and the optimization direction. In some embodiments, the first loss can be calculated using the policy gradient loss in reinforcement learning, such as the shearing objective function of the proximal policy optimization (PPO).

[0073] Optionally, the computer device can run multiple complete scheduling tasks in a virtual coupled scheduling environment based on the current policy network. For each scheduling task, the computer device can construct a state graph, extract global state features, output actions from the current policy network and execute them, and the environment provides feedback on the next state and immediate reward. This process is repeated until the scheduling task ends. When the scheduling task ends, the computer device can obtain a current trajectory. Multiple runs can obtain multiple current trajectories; for example, running 20 times (each using a different random seed) yields 20 current trajectories. The computer device can calculate the cumulative reward for each current trajectory and calculate the relative advantage of each trajectory based on the rewards of similar tasks (e.g., all 20 peak scheduling tasks). The computer device can apply shearing constraints to the probability ratios at each time step using a near-end policy optimization algorithm to aggregate and obtain the first loss.

[0074] Step 204: Obtain each historical trajectory from the experience base and determine the second loss based on each historical trajectory; the historical trajectory is used to record the complete scheduling process of virtual linkage scheduling control for historical scheduling tasks.

[0075] The experience base stores historical scheduling trajectory data. It can contain complete trajectories generated by different versions of the policy network (including earlier versions of the current policy network, and even completely different algorithms) during previous training or actual operation. The experience base not only records the trajectory's state, actions, and reward sequences, but can also attach metadata such as the cumulative reward, task type, success rate metric, and behavioral entropy for each trajectory to facilitate effective filtering and sampling. In some embodiments, the experience base can contain tens of thousands of trajectories accumulated over the past few days or weeks, covering successful and failed cases under different cargo volumes, different route occupancy conditions, and different conflict modes. For example, the experience base can specifically categorize tasks as "peak-hour 6-vehicle scheduling" and store 20 historical trajectories under this category, with 3 selected trajectories marked for focused learning.

[0076] Historical trajectories are complete task decision sequences generated and recorded by a historical policy network and retrieved from an experience base. The data format of historical trajectories can be the same as current trajectories, containing a series of state-action-reward tuples, but the policy network that generated these trajectories is a past policy network and may differ significantly from the current policy network. Historical trajectories can include both high-performing preferred trajectories and average-performing ordinary trajectories. When retrieving historical trajectories from the experience base, the computer device can select samples with higher learning value based on certain sampling strategies (such as based on task difficulty and trajectory behavior entropy). The complete scheduling process of historical trajectories allows policy updates to reuse successful experiences and avoid catastrophic forgetting.

[0077] The second loss is calculated using a loss function based on historical trajectories obtained from the experience base. It measures the degree to which the current policy network can reproduce the behavior of these historical data or its optimization potential. The calculation of the second loss can be similar in form to the first loss, but different processing strategies can be adopted depending on the quality of the historical trajectories. For example, for high-value historical trajectories marked as preferred, the computer device may not prune the probability ratios, allowing the policy network to make large-scale adjustments to mimic the patterns of these excellent behaviors; while for other suboptimal historical trajectories, the same pruning loss as the first loss can be used to prevent the policy from deviating too much from the old policy. The introduction of the second loss allows policy updates to not only rely on the latest online experience but also extract effective knowledge from rich historical experience.

[0078] For example, a computer device can access a pre-built and continuously maintained experience base, which stores complete trajectories from past training or actual runs. Each trajectory includes metadata such as task type identifier, cumulative reward, success rate, and behavioral entropy. The computer device can execute sampling strategies based on this metadata, such as prioritizing sampling low-entropy preferred trajectories in medium-difficulty tasks, while also including a certain number of random trajectories to maintain diversity. For instance, the computer device selects a group from the experience base with a task type of "6-column virtual coupled standard scheduling," which contains 10 historical trajectories. From this group, the computer device selects the trajectory with the lowest behavioral entropy as the preferred trajectory, and then randomly selects 4 more as suboptimal trajectories, for a total of 5 historical trajectories. The computer device can then determine a second loss based on these historical trajectories.

[0079] Step 206: Obtain the target loss based on the first loss and the second loss, and update the parameters of the policy network based on the target loss to obtain the updated policy network; the updated policy network is used for virtual coupling scheduling control for future scheduling tasks.

[0080] The target loss is the final optimization objective obtained by combining the first and second losses. It can include a weighted sum of the first and second losses. The target loss represents the overall objective that the policy network needs to minimize (or maximize, depending on the notation convention) in the current update round. By weighting the first and second losses, the target loss emphasizes learning new policy networks from the latest online exploration while preserving the stability of inheriting useful patterns from historical experience. During training, the parameters of the policy network can be updated by minimizing the target loss (or optimizing its inverse) using gradient descent. The updated policy network is the policy network with new parameter values ​​obtained after optimizing the target loss once. The structure of the updated policy network can remain unchanged, but the internal learnable parameters such as the weight matrix and biases are updated through stochastic gradient ascent or the optimizer, so that under the new parameters, for the input global state features, the action distribution output by the updated policy network tends to obtain a higher expected reward. The updated policy network can be immediately used for scheduling decisions in the next time step, i.e., generating actions based on the new state graph, then continuing to interact with the environment and accumulate new trajectories, and then entering the next training cycle.

[0081] Optionally, the computer device can combine the first loss and the second loss to obtain the target loss, such as by adjusting the weights according to a set value. The target loss is obtained by combining the first loss and the second loss, with weights... The parameters can be dynamically adjusted based on the training phase. After obtaining the target loss, the computer can use a gradient-based optimizer (such as the Adam optimizer) to update the parameters of the policy network. The optimization direction is to reduce the target loss (if it is a loss function), thus obtaining an updated policy network. In subsequent scheduling decisions, when a new real-time state graph is received, the updated policy network will output scheduling action decisions based on the updated parameters, thereby improving scheduling performance.

[0082] In this embodiment, the current trajectory obtained from online sampling is acquired and a first loss is determined. Historical trajectories are obtained from the experience base and a second loss is determined. The first loss and the second loss are combined into a target loss to update the policy network. This allows the update process to achieve a balance between exploring new strategies and retaining old experiences. It can learn new changes from online data while maintaining the excellent scheduling capabilities it has mastered. This ensures that the policy network can adapt to complex scheduling scenarios with dynamic environmental changes, which is beneficial to improving the transportation efficiency of virtual coupling of heavy-haul trains.

[0083] In an exemplary embodiment, the virtual coupling and rolling scheduling optimization method for heavy-haul trains further includes: for each candidate trajectory corresponding to various types of scheduling tasks in the experience base, determining the behavioral entropy of each candidate trajectory based on the probability of selecting an actual scheduling action at each step in the candidate trajectory and the number of trajectory steps of the candidate trajectory; determining the target trajectory for various types of scheduling tasks based on the behavioral entropy of each candidate trajectory; obtaining each historical trajectory from the experience base, including: determining the target type from various types of scheduling tasks based on the task difficulty of various types of scheduling tasks; the task difficulty is determined based on the success rate of each candidate trajectory for various types of scheduling tasks; and obtaining at least the target trajectory corresponding to the scheduling task of the target type from the experience base to obtain each historical trajectory.

[0084] The candidate trajectory is a historical scheduling record stored in the experience base. Each candidate trajectory can contain a complete state-action-reward sequence from the start to the end of the scheduling task, representing the data trace left after a complete virtual coupling and scheduling process of heavy-load trains. At each time step of the candidate trajectory, the policy network outputs an action probability distribution based on the current state, representing the likelihood of each candidate action being adopted. The probability value corresponding to the action that is actually adopted is the probability of selecting the actual scheduling action at that step. This probability reflects the degree of certainty of the policy network regarding that action in that state; the higher the probability value, the higher the confidence of the policy network in the output action. The trajectory step count refers to the total number of decision steps contained in the candidate trajectory from the initial decision time to the completion of the task. A longer step count means that the scheduling task includes more scheduling points and a longer or more complex control process.

[0085] Behavioral entropy is an indicator that quantifies the overall decision-making certainty of candidate trajectories. It can be calculated by taking the negative logarithm of the probability of each step in the candidate trajectory choosing the actual scheduling action, and then averaging the results over all steps. A lower entropy value indicates that the action selection at each step in the candidate trajectory is concentrated within a high-probability range, resulting in strong overall decision-making logic and low randomness—high-quality and worthwhile experience to learn from. A higher entropy value indicates more random decision-making and relatively lower reference value. A target trajectory, for a specific scheduling task type, is one or more trajectories selected from all candidate trajectories of that type, representing the lowest behavioral entropy. Target trajectories are considered the preferred experience for that task type because they embody the most certain and reasonable decision path for the current policy network (or the policy network at storage) when solving such scheduling tasks.

[0086] Task difficulty is a quantitative evaluation of the ease or difficulty of completing a certain type of scheduling task. It can be determined by the proportion of successful trajectories among all candidate trajectories for that task type: a higher success rate indicates an easier task, while a lower success rate indicates a more difficult task. The success trajectory ratio is the proportion of candidate trajectories in the experience base that are judged as "successful" for a certain type of scheduling task. "Success" can be determined based on whether the cumulative reward exceeds a preset threshold, or whether specific criteria such as safety without conflict and controllable delays are met. The success trajectory ratio can be used to calculate task difficulty; for example, if 40 out of 50 candidate trajectories for a task type are successful, the success trajectory ratio is 0.8. The target type is the task type selected from multiple scheduling task types after comprehensively considering task difficulty, serving as the focus of current training sampling. In some embodiments, a medium-difficulty type with a success trajectory ratio close to 0.5 can be selected, as this type of task provides the most information for policy improvement.

[0087] For example, the computer device can retrieve all existing candidate trajectories from an experience base and group these candidate trajectories according to their respective scheduling task types. Each group corresponds to a scheduling task type, such as "six-vehicle formation during peak hours," "four-vehicle mixed formation during off-peak hours," and "special scheduling in rainy weather." For each candidate trajectory in each task type, the computer device can obtain the probability of selecting the actual scheduling action at each step of the candidate trajectory. These probabilities can be historical inference values ​​directly read from the candidate trajectory, or values ​​obtained by re-evaluating the state sequence of the candidate trajectory using the latest policy network. The computer device combines the number of trajectory steps of the candidate trajectory and calculates the behavioral entropy using the average negative logarithmic probability. For example, if a candidate trajectory has three steps, with action probabilities of 0.9, 0.8, and 0.95 for each step, then the behavioral entropy is -1 / 3(log0.9 + log0.8 + log0.95) ≈ 0.12. The lower the behavioral entropy, the more certain the decision of the candidate trajectory. Within each task type, the computer device can sort candidate trajectories based on their behavioral entropy and select one or more trajectories with the lowest behavioral entropy, marking them as the target trajectories for that task type. The number of target trajectories selected can be fixed (e.g., one) or dynamically determined according to quality quantiles (e.g., the top 5%). Each task type has one or more target trajectories representing its optimal decision-making mode.

[0088] After marking the target trajectory, the computer device can calculate the success rate of all candidate trajectories for each task type. Assuming a task type has 10 candidate trajectories, and 6 are deemed successful, the success rate is 0.6. The computer device can determine the task difficulty for that task type based on the success rate. For example, a rate close to 0.5 is considered medium difficulty, a rate significantly higher than 0.5 is considered easy, and a rate significantly lower than 0.5 is considered difficult. The computer device can determine one or more target types from various task types according to a preset strategy (e.g., always selecting medium difficulty task types). The computer device can obtain at least the target trajectory for the determined target type from the experience base as the final historical trajectory used for strategy updates. If more samples are needed, other candidate trajectories (non-target trajectories) within the same target type can also be obtained simultaneously, but at least the target trajectory must be included.

[0089] For example, suppose there are two task types in the experience base: Type P (peak-hour six-vehicle virtual coupling scheduling) with 40 candidate trajectories, and Type Q (off-peak-hour four-vehicle mixed formation scheduling) with 30 candidate trajectories. The computer first calculates the behavioral entropy of each candidate trajectory. For Type P, trajectory PT4 has a behavioral entropy of only 0.11, the lowest in its class; for Type Q, trajectory QT7 has a behavioral entropy of 0.08, also the lowest in its class. The computer marks PT4 and QT7 as the target trajectories for their respective types. Then, the computer calculates the success rate of trajectories: 28 out of the 40 trajectories for Type P are successful, a rate of 0.7; 15 out of the 30 trajectories for Type Q are successful, a rate of 0.5. Since the success rate for Type Q is closer to 0.5, the computer determines Type Q as the target type. Therefore, it retrieves the target trajectory QT7 (at least this one) for Type Q from the experience base as a historical trajectory. In this way, the historical samples used for updating come from the task type most suitable for increasing difficulty (medium difficulty), and are the most definitive demonstration trajectories within that type.

[0090] In this embodiment, target trajectories in each type of task are filtered out by behavioral entropy, and then task types with learning value are located by task difficulty. Target trajectories in this task type are extracted as historical trajectories for training, which improves the information density of training samples and balances the training opportunities for tasks of different difficulties, thereby improving the update training effect of the policy network.

[0091] In an exemplary embodiment, determining a first loss based on each current trajectory includes: for each current trajectory belonging to the same type of scheduling task, determining the relative advantage of each current trajectory based on its respective trajectory reward; the trajectory reward is determined based on the departure rhythm, path conflict, delay risk and energy consumption weighted average of the current trajectory; and determining the first loss based on the relative advantage of each current trajectory through a near-end policy optimization algorithm.

[0092] Trajectory rewards are used to evaluate the overall performance of a complete trajectory (i.e., the entire scheduling process from start to finish). Based on trajectory rewards, the policy network can be guided to update in a more optimal direction. Trajectory rewards can be determined based on a weighted average of departure rhythm, path conflict, delay risk, and energy consumption. For example, computer equipment can calculate the trajectory reward based on various data recorded during the scheduling process, according to preset weight coefficients. Departure rhythm is an indicator of scheduling efficiency, reflecting the number of trains successfully dispatched per unit time or the compactness of the departure sequence. For example, departure rhythm can include the total number of trains successfully dispatched within a scheduling cycle. Path conflict is an indicator of scheduling safety, reflecting the number of times train running paths overlap or interfere in time and space during the scheduling process. By monitoring the position and occupied section information of each train during the scheduling process, it is possible to detect whether two or more trains simultaneously request to occupy the same track resource; each occurrence is counted. Delay risk is an indicator of scheduling punctuality, reflecting the degree of deviation between the actual train running time and the planned time. For example, delay risk can include the sum of all train delay times. Energy consumption is an indicator of scheduling economy and environmental friendliness, reflecting the total energy consumed by trains during operation.

[0093] For example, in a scheduling task, if 10 trains are successfully dispatched, 0 conflicts occur, the total delay time is 30 seconds, and the total energy consumption is 500 units, with weighting coefficients α=1, β=100, γ=0.5, and δ=0.1, then the trajectory reward can be calculated as 1×10 - 100×0 - 0.5×30 - 0.1×500, resulting in -40. As another example, if another trajectory successfully dispatches 12 trains, experiences 1 conflict, has a total delay time of 60 seconds, and a total energy consumption of 600 units, then its trajectory reward is 1×12 minus 100×1 minus 0.5×60 minus 0.1×600, resulting in -118.

[0094] Relative advantage measures the performance of a trajectory relative to the average level of similar tasks. It distinguishes the value of different trajectories during policy network updates, providing positive incentives for trajectories performing better than the average and negatively suppressing those performing worse. For a set of current trajectories belonging to the same type of scheduling task, a computer device first calculates the trajectory reward for each trajectory, then calculates the average reward for the set, and finally subtracts this average from the reward for each trajectory to obtain its relative advantage. Proximal policy optimization (PFO) is a policy gradient algorithm in reinforcement learning used to control the update step size when updating the policy network, preventing policy performance collapse due to excessively large single updates. PFO can introduce a pruning mechanism to limit the ratio of action selection probabilities between the old and new policy networks within a reasonable range, thus ensuring the stability of policy updates. In some embodiments, the computer device can calculate a pruned loss value based on the relative advantage and probability ratio of the current trajectory. This probability ratio is the ratio of the probability of the current policy network selecting a certain action to the probability of the old policy network selecting the same action.

[0095] For example, the computer device can classify current trajectories, grouping trajectories belonging to the same type of scheduling task together. For instance, all trajectories for "peak-hour three-train formation" tasks can be grouped into one group, and all trajectories for "off-peak-hour single-train" tasks into another. For each current trajectory within a group, the computer device can calculate the trajectory reward for each trajectory. The trajectory reward can be obtained by weighting and summing the recorded departure rhythm, number of path conflicts, cumulative delay time, and total energy consumption within the current trajectory, combined with preset weighting coefficients. The computer device can calculate the average reward of all current trajectories within the group. For each current trajectory within the group, the computer device can subtract the group average from its own trajectory reward to obtain the relative advantage of that current trajectory. For example, assuming there are three trajectories in a group with rewards of 90, 70, and 50 respectively, and an average of 70, the relative advantages are 20, 0, and -20 respectively. This relative advantage value reflects the performance of the trajectory relative to the average level of the group; a positive value indicates better than the average, and a negative value indicates worse than the average.

[0096] After obtaining the relative advantage of each current trajectory, the computer device can calculate the probability ratio corresponding to each current trajectory. This probability ratio is the ratio of the probability that the current policy network chooses to actually execute an action at each decision point of the current trajectory to the probability that the old policy network before the update chooses the same action. For a current trajectory containing multiple decision steps, the computer device can calculate the probability ratio for each decision step and finally obtain a comprehensive trajectory-level probability ratio. The computer device can apply the pruning objective function of the proximal policy optimization algorithm. For each current trajectory, the computer device can calculate the unpruned loss, i.e., the probability ratio multiplied by the relative advantage; the computer device can also calculate the pruned loss, i.e., first limiting the probability ratio to the range of 1 minus ε to 1 plus ε, and then multiplying it by the relative advantage. The computer device can take the smaller of these two losses as the final loss for the current trajectory. The computer device can sum or average the final losses of all current trajectories in the group to obtain the first loss. The first loss is used to update the parameters of the policy network, so that when the policy network faces similar scheduling tasks in the future, it will be more inclined to produce high-reward decisions and avoid producing low-reward decisions.

[0097] In this embodiment, the first loss is determined based on the trajectory reward of each current trajectory belonging to the same type of scheduling task by using the near-end policy optimization algorithm. This can reduce the impact of the difference in task difficulty and prevent the policy performance from fluctuating drastically, thereby ensuring the stability and reliability of the policy network training.

[0098] In an exemplary embodiment, each historical trajectory includes a first historical trajectory belonging to the target trajectory and a second historical trajectory not belonging to the target trajectory. Determining a second loss based on each historical trajectory includes: for each historical trajectory belonging to the same type of scheduling task, determining the relative advantage of each historical trajectory based on its respective trajectory reward; the trajectory reward is determined based on the departure rhythm, path conflict, delay risk, and energy consumption weighting of the historical trajectory; obtaining the loss of each first historical trajectory based on its relative advantage and probability ratio; obtaining the loss of each second historical trajectory based on its relative advantage using a near-end policy optimization algorithm; and determining the second loss based on the losses of each first historical trajectory and the losses of each second historical trajectory.

[0099] The first historical trajectory refers to the target trajectory selected from the experience base that belongs to the target type of scheduling task. The first historical trajectory is an experience trajectory considered to have high learning value after hierarchical screening and priority sampling. In some embodiments, the computer device can, for each type of scheduling task, sort the candidate trajectories from low to high behavioral entropy and select one or more trajectories with the lowest entropy value as the target trajectory for that type. The second historical trajectory refers to other historical trajectories in the experience base that belong to the same type of scheduling task as the first historical trajectory but are not target trajectories. Although the second historical trajectory is not selected as a target trajectory, it still contains certain experience information and can assist in policy training. The probability ratio refers to the ratio of the action selection probability of the current policy network and the old policy network in the same state-action pair. For each decision step in a historical trajectory, the probability of actually executing the action in that step state can be calculated based on the current policy network and the old policy network respectively. The probability ratio for that step is obtained based on the ratio between the probability of the current policy network actually executing the action in that step state and the probability of the old policy network actually executing the action in that step state. Based on the probability ratios of each step in the trajectory, the overall probability ratio of the trajectory can be obtained.

[0100] Optionally, the computer device can retrieve historical trajectories from an experience database, which can be grouped according to the type of scheduling task they belong to. For each group of K historical trajectories belonging to the same type of scheduling task, the computer device can perform analysis on each historical trajectory. Calculate its trajectory reward The trajectory reward is the result of a comprehensive reward function, specifically comprising: the number of successfully dispatched trains multiplied by weight α, minus the number of path conflicts multiplied by maximum weight β, minus the cumulative delay time multiplied by weight γ, and minus energy consumption multiplied by weight δ. These components are standardized and then weighted and summed to obtain a scalar reward value. The computer equipment can calculate the relative advantage based on the rewards of all historical trajectories within the group. This relative advantage measures the performance of the historical trajectory relative to the group's average level; a positive value indicates performance better than average, and a negative value indicates performance worse than average.

[0101] For a first historical trajectory (i.e., the target trajectory) belonging to the same type of scheduling task, the computer device can extract the state-action pairs at each time step of the first historical trajectory and calculate the probability ratio of the current policy network relative to the old policy network. The computer device can then determine the relative advantage of each first historical trajectory. Multiplying by the probability ratio yields the uncut importance sampling loss term. If a first historical trajectory contains multiple time steps, the losses of each time step can be averaged or summed. The computer can process the first historical trajectory as a whole, for example, by using trajectory-level probability ratios (such as the product or average of the probability ratios of each step) or by calculating step by step and then summing to obtain the loss for each first historical trajectory. This loss directly reflects the performance of the current policy network on that first historical trajectory; the larger the value, the better the consistency between the policy and successful experiences.

[0102] For second historical trajectories (i.e., non-target trajectories) belonging to the same type of scheduling task, the computer device can calculate the probability ratio of each second historical trajectory at each time step and use the pruning mechanism of Proximal Policy Optimization (PPO) to obtain the loss of the second historical trajectory based on the relative advantage of each second historical trajectory. In some embodiments, for each time step, the computer device can calculate the pruned target item based on the Proximal Policy Optimization algorithm, and the computer device summarizes the losses of all time steps (e.g., by averaging) as the loss of that second historical trajectory. The computer device can combine the losses of all first historical trajectories in the same group (denoted as L1) with the losses of all second historical trajectories (denoted as L2) to obtain the second loss for this type of scheduling task. In some embodiments, the combination method can be a simple summation or a weighted summation (e.g., averaging by the number of trajectories).

[0103] In this embodiment, for the first historical trajectory, the loss is determined by directly using the relative advantage and probability ratio without shearing. For the second historical trajectory, the corresponding loss is determined by the near-end policy optimization algorithm. This allows the policy network to be updated using the experience base, so that the best experience can be fully utilized to quickly improve performance, while a large number of ordinary experiences can be handled reliably, thereby achieving efficient and stable policy network optimization.

[0104] In one exemplary embodiment, the virtual coupling scheduling control is rolled to the next moment until the scheduling task for the heavy-haul train is completed, including: determining the next moment as the current moment, and returning to the step of obtaining the scheduling environment data of the heavy-haul train at the current moment, until the scheduling task for the heavy-haul train is completed.

[0105] In this context, the "next moment" refers to the next decision-making point generated relative to the current moment during the rolling optimization process. The next moment represents the next state acquisition and decision execution point that the rolling loop will advance to. For example, after completing the scheduling action at the current moment and waiting for the next moment, the computer device can check whether the scheduling task has been completed. If the scheduling task is not completed, the computer device assigns the previously determined next moment to the current moment and jumps back to the initial step, that is, re-executes the step of acquiring the scheduling environment data of the heavy-load train at the current moment, and sequentially completes operations such as state diagram construction, graph neural network embedding, policy network generation, and execution of scheduling control. This loop repeats until the conditions for completing the scheduling task are met.

[0106] In some embodiments, train departure tasks can be completed cyclically with a fixed step size. Assuming the computer device is a dispatch center server, the dispatch task is to continuously dispatch thirty heavy-load trains within a station within 15 minutes, with each train using a virtual coupling method and a minimum interval of 5 seconds. The server sets the decision step size to 1 second. At the initial time t0, the server acquires environmental data, constructs a state graph containing all trains to be dispatched and track resources, and generates a decision through graph embedding and policy network: immediately dispatch the first and second trains, with a 5-second interval. The server sends the instruction to the trains. Subsequently, the server determines the next time as t0+1 seconds and checks whether the task is completed (not completed, there are still twenty-eight trains to be dispatched). The server then assigns the next time to the new current time and returns the environmental data acquired at time t0+1 seconds. At this point, the first train has traveled 300 meters, and the second train is preparing to depart. Based on the new state, the server decides not to dispatch the trains temporarily to maintain a safe interval. The server continues to scroll every second, updating the current time, continuously collecting data and making decisions. After the 30th train departs, the server checks that the queue of trains to be dispatched is empty and that all dispatched trains have entered automatic tracking mode. It then determines that the dispatching task is complete and the loop terminates.

[0107] In this embodiment, by determining the next moment as the current moment and returning to obtain the scheduling environment data of the heavy-haul train at the current moment, the virtual coupling rolling scheduling has the ability to respond to dynamic changes in real time. This ensures that the virtual coupling scheduling control can accurately match the real-time scheduling environment, thereby improving the transportation efficiency in the virtual coupling operation scenario of heavy-haul trains.

[0108] This application also provides an application scenario in which the above-mentioned heavy-haul train virtual coupling and rolling scheduling optimization method is applied. Specifically, the application of the heavy-haul train virtual coupling and rolling scheduling optimization method in this scenario is as follows:

[0109] Traditional timetable scheduling methods for fixed or moving block systems become inadequate in the face of such high-density, tightly coupled, and multi-objective scenarios. Conventional rule-based control or optimization algorithms often rely on pre-defined safety bays and fixed scheduling plans, making it difficult to respond promptly to dynamic changes in freight volume and unforeseen circumstances, and unable to learn and improve strategies from past scheduling processes. Therefore, it is necessary to research an efficient learning-based intelligent scheduling system that maximizes departure efficiency and reduces delays and energy consumption, while considering the virtual coupling of heavy-haul trains.

[0110] Based on this, the virtual coupling and rolling scheduling optimization method for heavy-haul trains provided in this application can solve the problems of difficulty in balancing safety and efficiency, difficulty in balancing multiple objectives, and lack of self-learning ability in existing scheduling schemes. By introducing a reinforcement learning framework and an experience-based optimization mechanism, the scheduling system can optimize the departure rhythm of heavy-haul trains while ensuring no line conflicts, dynamically adjust the operation plan to reduce the risk of delays for heavy-haul trains, and control the acceleration, deceleration, and virtual formation process of heavy-haul trains, thereby saving energy. This application achieves real-time scheduling decision updates through rolling optimization, enabling the system to adapt to changes in freight volume and emergencies, thereby improving transportation efficiency, reducing freight waiting time, and saving energy.

[0111] Specifically, the virtual coupling and rolling scheduling optimization method for heavy-haul trains provided in this application includes processes such as graph trajectory construction, structure embedding, task priority normalization, experience reorganization, and strategy generation, wherein:

[0112] 1. Graph trajectory construction

[0113] First, the current scheduling scenario is abstracted as graph-structured data. A state graph is constructed considering the line topology, train positions, and their interrelationships. For example, it can be assumed that... A diagram representing the scheduling environment: It is a set of nodes, including train nodes and resource nodes (track / section nodes), etc. The graph is a set of edges representing the relationship between heavy-haul trains and track resources, or the interaction between trains. Each train is a node in the graph, and each track section, signal block, or platform can also be considered a node. If a train occupies a section, an edge is established between the corresponding train node and the section node to represent the resource occupancy relationship. Different train nodes are connected by edges to represent potential tracking interval constraints or virtual coupling relationships. This graph structure representation unifies the modeling of train operating status, infrastructure occupancy, and coupling relationships between trains. During the graph trajectory construction process, it is also necessary to integrate time-dimensional information, discretizing time into rolling window frames at each decision moment. Construct the corresponding state diagram For each train currently in operation or awaiting departure, its historical trajectory can be recorded as an attribute of that train node; while trains that have not yet departed can be represented in the graph through special nodes or additional attributes. In virtual coupling scenarios, dynamic train formation and decoupling may occur, and changes in train formation relationships can be reflected by adding or removing edges on the graph.

[0114] In some embodiments, the graph trajectory construction process can be replaced with other forms of state representation as needed. For example, the position of a heavily loaded train can be represented by a time-space grid, or the states of multiple trains can be encoded simultaneously using sequence vectors.

[0115] 2. Structural embedding

[0116] The state graph is input into a pre-trained graph neural network model to extract high-level feature representations. Specifically, a graph convolutional network or graph attention network can be used to encode features of each node and edge in the state graph. Then, message passing or aggregation algorithms are used to obtain the embedding representation of the entire state graph, i.e., the global state features. The hidden vector representation of each node is obtained through iterative updates. A single propagation update can be formalized as:

[0117]

[0118] in, Indicates the first Layer Time Node The representation of (i.e., the first) Layer Time Node (initial characteristics) It is a node The set of neighboring nodes, It is a node The set of neighboring nodes The neighboring node is at the _th ... Initial characteristics of the layer; The adjacency normalization coefficient is... , , For trainable weight matrix, It is a non-linear activation function. For the first Layer Time Node The initial characteristics.

[0119] After several layers of propagation, the graph neural network model obtains node embeddings containing adjacency relationships and trajectory interaction information. Then, it performs a readout operation on the entire graph or a specific subgraph to generate global state features. , serving as the state representation of the policy network. Among them, For read operations, For nodes go through The target features obtained after layer propagation This is the set of nodes in the state graph at the current moment.

[0120] In some embodiments, emphasis can be placed on encoding path-level trajectories. For each train node, its complete running path can be regarded as a sequence. By designing path aggregation operators in the graph neural network model or extracting path features by combining sequence models, the embedded representation can cover the global travel information of the train. At the same time, through behavioral semantic modeling, some patterns in the state graph are specially processed. When two train nodes are detected to be connected by an edge and are very close, the feature is given extra weight to reflect the behavioral semantic of "virtual coupling and following". If a train node has a long-term stopped state, the feature label "waiting" semantic can be introduced. Through the computation of structural embedding processing, the final state feature vector contains rich spatial topology and behavioral semantic information.

[0121] In some embodiments, structural embedding processing can also employ Transformer networks to model trajectory sequence features, or use traditional machine learning feature extraction methods instead of graph neural networks. If obtaining the graph structure is inconvenient, a centralized state vector can be directly input into the neural network for decision-making, but this will result in the loss of some topological information.

[0122] 3. Task priority unification

[0123] Specifically, based on multi-objective scheduling requirements, a unified evaluation index and reward function can be designed to clarify the objectives of policy network optimization. For the scheduling task, a comprehensive reward can be defined. This approach unifies the evaluation of objectives such as departure schedule, route conflicts, delay risks, and energy consumption under a single benchmark. Specifically, a weighted sum approach can be used.

[0124]

[0125] in, For comprehensive rewards, This indicates the number of trains successfully dispatched within the scheduling cycle. Indicates the number of path conflicts that occurred. This indicates the total accumulated delay time of the train. This indicates the total amount of energy consumed. The weighting coefficients are used to reflect the priority of each objective. To ensure safety, a very high weight can be assigned to the conflict penalty term, ensuring that any scheduling strategy receives a very low reward if a conflict occurs, thus enabling the reinforcement learning AI to spontaneously avoid path conflicts. In some embodiments, other weights can be adjusted according to the operator's priorities; for example, they can be increased for routes with high punctuality requirements. It can improve efficiency in scenarios with high energy-saving requirements. In some embodiments, to achieve normalization of indices with different dimensions, the above items need to be appropriately standardized to maintain a comparable numerical scale. Through this comprehensive reward, the complex multi-objective optimization problem is transformed into a single scalar objective for evaluation by the reinforcement learning value function. Task priority normalization means that the weights of each objective can be dynamically adjusted according to actual needs, with priority given to increasing them during peak periods. The weighting system encourages shorter departure intervals, while safety-related weights can be temporarily increased in special circumstances. This flexible weighting strategy ensures that different objectives are measured within a unified framework and allows for highlighting key areas as needed, thus forming a clear direction for optimization.

[0126] In some embodiments, the reward design can be modified according to changes in operational objectives. For example, in certain application scenarios, energy consumption can be... Replace these with indicators such as wear and tear costs of heavy-haul trains or carbon emission indicators, and incorporate them into the comprehensive evaluation. Weighting The weights are not fixed and can be adjusted using adaptive weighting algorithms. In multi-objective optimization, the concept of Pareto optimality can also be introduced, allowing the policy network to learn to find equilibrium solutions among different objectives rather than simply weighted sums.

[0127] 4. Experience Reorganization

[0128] Specifically, the experience trajectories accumulated during reinforcement learning training can be classified, filtered, and recombined for use, which can improve learning efficiency and policy convergence speed in sparse reward scenarios. This can be achieved through three steps: experience storage, hierarchical filtering, and priority sampling.

[0129] (1) The policy network continuously generates policy trajectories during its interaction with the scheduling environment. Each trajectory contains a series of state-action-reward sequences, recording a scheduling decision process. One trajectory can correspond to several train departure scheduling schemes and their results for a complete peak period. For each trajectory, its cumulative reward is extracted. And the completion status of each objective. The trajectory and related indicators are stored in the experience base. In reserve. When rail transit scenarios have repetitive task characteristics, multiple trajectories of the same type of scheduling task are grouped together for subsequent unified evaluation.

[0130] (2) Regularly review the experience database The analysis was conducted, and the trajectories were stratified based on the difficulty and success rate of the tasks they corresponded to. A task success rate metric was defined to measure trajectory quality. This indicates that the experience base is for the first The success rate of similar scenarios. Tasks are categorized by... The difficulty level is categorized into ranges: scenarios above a certain threshold are considered easy, those between medium and medium difficulty are considered medium difficulty, and those close to zero are considered difficult. Furthermore, simple tasks that have been fully mastered are marked as retired and no longer require training. Within each scenario category, trajectories are further filtered based on other quality metrics, and the behavioral entropy of each trajectory is calculated to measure its determinism and reasonableness. The behavioral entropy of a trajectory is defined as:

[0131]

[0132] in, For trajectory, For trajectory behavioral entropy, For trajectory The number of steps in the trajectory, The time step index represents the trajectory. The Middle At the moment of decision-making, The policy network represents the trajectory The Middle Step to select actual action The probability, For the first The environmental state at the time of step (global state characteristics). The lower the entropy value of the behavior entropy, the more certain the behavior selection of the trajectory, and the more low the randomness and high the logic of the strategy. This application regards low-entropy trajectories as high-value experiences. In each task category of the experience base, one or more trajectories with the lowest entropy value are selected and marked as the preferred experiences of that category. (i.e., the target trajectory).

[0133] (3) During each iteration of reinforcement learning training, prioritize sampling tasks of moderate difficulty that contain high-value trajectories from the experience base for learning. When sampling a mini-batch of tasks, the task... Assigning the probability of being selected:

[0134]

[0135] in, For the task The probability of being sampled. For the task Success rate in the experience base To achieve the target success rate, To control the distribution width, It is an exponential function. This distribution makes... The task that is too easy or too difficult receives the highest selection probability, while the task that is too easy or too difficult receives a lower sampling probability.

[0136] After selecting a task, it is derived from its corresponding set of experience trajectories. (No. The first in the class of scheduled tasks Samples are drawn from a set of empirical trajectories, with priority given to selecting the preferred trajectories from that set. (No. The target trajectory of the scheduling task is the primary learning object. If multiple trajectory samples are needed, the remaining samples can also be selected by sorting them by entropy value.

[0137] In some embodiments, the experience-first sampling mechanism can be replaced by the classic experience-first replay algorithm. The experience-first replay algorithm determines the sampling weight based on the error of each experience, thus prioritizing the learning of experiences where policy estimation is inaccurate, unlike selection based on success rate and entropy. A forgetting mechanism can also be used to manage the experience base, periodically discarding outdated data to prevent interference from old experiences as the policy evolves over time.

[0138] 5. Strategy Generation

[0139] Specifically, based on the current policy network and the selected experience batches, reinforcement learning can be used to optimize and update the policy, generating new scheduling decision policies. For example, a hybrid policy optimization objective function can be employed, combining the latest trajectories obtained from real-time exploration in the environment with the optimal trajectories replayed from the experience base to update the policy network parameters. .definition Let the empirical samples represent the proportion of the entire training batch, then the target loss can be expressed as:

[0140]

[0141] in, It is the target loss. The loss (first loss) is calculated using environmental data sampled online in real time. The loss (second loss) is calculated using replay data from the experience base, and the two are weighted and summed to achieve a balance between exploration and utilization.

[0142] (1) For the on-orbit strategy optimization objective This paper employs the idea of ​​proximal policy optimization (PPO) to update policies for online samples, specifically combining it with a group relative policy optimization framework to improve the stability of advantage estimation in multi-task scenarios. For the current policy network... A batch of scheduling sequences is sampled from the environment, and each scheduling task in the sequence is treated as a group. For the first task within a group... Trajectory Calculate its advantage estimate It can be obtained by subtracting from the baseline:

[0143]

[0144] in, It is the first Trajectory, It is a trajectory The advantage estimate, It is a trajectory Cumulative rewards; It is the number of trajectories contained in the task group. Each task group corresponds to a type of scheduling scenario (such as a set of departure tasks on a specific route or during peak hours). Each trajectory in the group is generated by different strategies (or exploring noise) under the same or similar conditions. For all within the task group The cumulative reward summation for each trajectory. Advantage function. The trajectory was measured Compared to the average performance of the task, it can effectively reduce the difference in benefits caused by the difference in difficulty between different tasks.

[0145] Define the probability ratio of the policy This indicates the current strategy relative to the old strategy in terms of trajectory. The ratio of the probabilities of choosing the above action. Wherein, The probability ratio of the strategies. The current policy network has the following parameters: , For the old policy network, the parameters are: , In the trajectory The actual actions taken by the policy network at a certain time step. In the trajectory Central strategy network takes action The corresponding global state features; For the current policy network in state Select action The probability, For the old policy network in state Select action The probability of.

[0146] The shear target of PPO can then be written as:

[0147]

[0148] in, The first loss, This represents the number of trajectories contained in a task group. For trajectory The estimated advantage (relative advantage). It is the shear threshold constant. The operation limits the probability ratio to Within the range.

[0149] (2) For experience-based optimization objectives These optimized trajectories sampled from the experience base are also used to update the strategy. Since these trajectories may be generated by old strategies or even different strategies, distribution differences need to be corrected during the update. Here, a hybrid strategy objective is adopted. For batches of experience samples, not only is a PPO shearing loss similar to that described above used, but special weights are also assigned to the best trajectory in each task group. Let the optimized trajectory sampled from each group of experience tasks be... ,for Corresponding probability ratio Instead of applying shearing to fully utilize its high-dominance information, the loss of other suboptimal trajectories is still calculated using the shearing method. This can be represented as the sum of the uncut update items and the cut update items, i.e.

[0150]

[0151] in, This is the second loss. It is the ratio of the probability of the preferred trajectory. The function, It is the preferred trajectory The importance sampling probability ratio of (first historical trajectory) It is the preferred trajectory The estimate of the dominance function, It is the set of all trajectories in the experience batch, extracted from the experience database by a priority sampling mechanism. It is the alternative to the preferred trajectory In addition, experience batches Other alternative trajectories (secondary historical trajectories) in the data. It is the second historical trajectory The probability ratio, It is the second historical trajectory The estimate of the dominance function, Probability Ratio Perform a clipping operation. We can use the identity function or a function that is appropriately amplified but avoids infinity. By maximizing... The strategy will draw inspiration from experience, especially to reproduce decision-making patterns from successful experiences as much as possible, while using a cut-off mechanism to avoid significant differences from the old strategy.

[0152] Ultimately, the target loss The gradient is as described above. and The policy parameters are updated using a weighted sum of partial gradients and stochastic gradient ascent. After one update, the improved policy network is obtained. .

[0153] Through iterative training of policy generation processing, the agent's scheduling strategy will gradually become optimized. Guided by a large number of simulations and experience reorganizations, it learns more compact control of heavy-haul train departure rhythm, more effective conflict prevention schemes, more reasonable delay trade-offs, and more energy-efficient driving strategies.

[0154] In actual operation, the trained policy network can receive the current state graph in real time. Global state features are extracted through the aforementioned structural embedding. The policy network then outputs the scheduling action decisions for each train at the current moment. Typical actions include deciding whether to depart immediately for each train to be dispatched, issuing speed adjustment commands to trains in operation, and whether to perform virtual coupling operations, etc. Then, the system enters the interactive environment to execute and observe the next state. And possible reward feedback, thus forming a rolling scheduling closed loop.

[0155] In some embodiments, in addition to the PPO / GRPO (Group Relative Policy Optimization) paradigm, other reinforcement learning algorithms can be used to achieve similar functionality. For example, deep Q-networks can be used to learn the Q-values ​​of scheduling actions and leverage priority replay to accelerate convergence.

[0156] In some embodiments, assuming a heavy-haul train line, a virtual coupling train control system is configured, allowing two heavy-haul trains to depart consecutively with an ultra-small interval of 5 seconds. Now, during peak hours, three trains, A, B, and C, need to be deployed in close coordination, with A as the lead train and B and C as follow trains. Initial state: Heavy-haul train A departs according to schedule. Bus B departs, with B following closely behind. Sending in seconds, C is The system constructs a graph based on train location information and the planned status, including 3 train nodes and multiple section nodes, with each train node connected to its corresponding section node. After train A departs, train B prepares to depart after 30 seconds. At this point, train A is approximately 20 seconds away from the exit of the next section. The agent, through graph embedding, recognizes that if train B departs at 60-second intervals, it will be too close to train A in section 2, and train A may be slightly delayed due to increased passenger volume. Considering these factors, the strategy decides to delay train B's departure by 10 seconds. This ensures that train B... When the train departed, car A had just made enough time to avoid being too close during the coupling and following process. The departure interval of car C could then be appropriately shortened to 50 seconds. In subsequent operations, cars A and B maintained a virtual coupling, with car B automatically adjusting its speed to maintain an interval of approximately 5 seconds with car A, and arriving at stations simultaneously. Car C, arriving slightly later, was not directly coupled. As a result, the three cars successfully and safely completed departures at an average interval of approximately 55 seconds, 15 seconds less than the originally planned total interval, without any delays such as waiting at stations or stopping at signals.

[0157] The virtual coupling and rolling scheduling optimization method for heavy-haul trains provided in this application reduces the minimum departure interval under virtual coupling conditions through reinforcement learning intelligent optimization strategies. While ensuring safety, the system automatically finds the optimal combination of different train departure sequences, significantly increasing the number of heavy-haul trains passing through per unit time and improving capacity during peak periods. By severely penalizing conflict events and using graph networks to finely characterize the distance relationships between heavy-haul trains, the strategy ensures that collision avoidance is always the priority. Even in emergency situations, the strategy can quickly adjust subsequent trains to maintain a stable safe interval. Energy-saving optimization is achieved through intelligent control; the strategy learns to utilize synchronization and coordination between heavy-haul trains to reduce unnecessary braking and acceleration.

[0158] Compared to pre-designed rules, the virtual coupling and rolling scheduling optimization method for heavy-haul trains provided in this application leverages reinforcement learning to continuously learn and improve from historical data. Experience-based restructuring ensures that the strategy draws lessons from past successes and failures, such as identifying which scheduling modes are prone to bottlenecks and which train sequence combinations are more efficient. Therefore, the system is adaptive to different freight modes and line parameters, eliminating the need for frequent manual parameter adjustments. When the environment changes, rolling optimization can quickly replan and restore stable operation; this intelligent adaptability greatly enhances the robustness and responsiveness of the scheduling system.

[0159] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.

[0160] Based on the same inventive concept, this application also provides a heavy-haul train virtual coupling and rolling scheduling optimization device for implementing the above-mentioned heavy-haul train virtual coupling and rolling scheduling optimization method. The solution provided by this device is similar to the solution described in the above method. Therefore, the specific limitations of one or more embodiments of the heavy-haul train virtual coupling and rolling scheduling optimization device provided below can be found in the limitations of the heavy-haul train virtual coupling and rolling scheduling optimization method above, and will not be repeated here.

[0161] In one exemplary embodiment, such as Figure 3 As shown, a virtual coupling and rolling scheduling optimization device 300 for heavy-haul trains is provided, comprising: a state diagram construction module 302, a feature propagation module 304, an action decision module 306, and a scheduling control module 308, wherein:

[0162] The state diagram construction module 302 is used to obtain the scheduling environment data of the heavy-load train at the current moment and construct the state diagram at the current moment based on the scheduling environment data; the nodes in the state diagram include train nodes and resource nodes, and the edges in the state diagram are used to represent the interaction relationship between nodes.

[0163] The feature propagation module 304 is used to propagate the node features of each node in the state graph at least one layer based on the pre-trained graph neural network model to obtain global state features.

[0164] Action decision module 306 is used to generate scheduling action decisions for the current time based on global state features using a pre-trained policy network.

[0165] The dispatch control module 308 is used to perform virtual coupling dispatch control for heavy-haul trains at the current moment according to the dispatch action decision, and roll to the next moment for virtual coupling dispatch control until the dispatch task for heavy-haul trains is completed.

[0166] In some embodiments, the feature propagation module 304 is further configured to extract initial features of each node and each edge in the state graph based on a pre-trained graph neural network model; based on the initial features of each node and each edge, perform at least one layer of propagation on the node features of each node to obtain the target features of each node; wherein, in the propagation of each layer, for each node, the initial features of the target node in the next layer are determined based on the initial features of the target node in this layer, the initial features of the neighboring nodes of the target node in this layer, and the edge features between the target node and the neighboring nodes; and the global state features are aggregated based on the target features of each node.

[0167] In some embodiments, the feature propagation module 304 is further configured to determine each weight adjustment factor based on the edge features between the target node and each neighbor node; and to determine the initial features of the target node in the next layer based on the initial features of the target node in this layer, the initial features of the neighbor nodes of the target node in this layer, and each weight adjustment factor.

[0168] In some embodiments, a policy update module is further included, configured to acquire each current trajectory obtained from online sampling, and determine a first loss based on each current trajectory; the current trajectory is used to record the complete scheduling process of virtual coupled scheduling control for the corresponding scheduling task through the policy network; acquire each historical trajectory from the experience base, and determine a second loss based on each historical trajectory; the historical trajectory is used to record the complete scheduling process of virtual coupled scheduling control for historical scheduling tasks; obtain a target loss based on the first loss and the second loss, and update the parameters of the policy network based on the target loss to obtain an updated policy network; the updated policy network is used for virtual coupled scheduling control for future scheduling tasks.

[0169] In some embodiments, the policy update module is further configured to: determine the behavioral entropy of each candidate trajectory for each candidate trajectory corresponding to various types of scheduling tasks in the experience base, based on the probability of selecting an actual scheduling action at each step in the candidate trajectory and the number of trajectory steps of the candidate trajectory; determine the target trajectory for each type of scheduling task based on the behavioral entropy of each candidate trajectory; determine the target type from various types of scheduling tasks based on the task difficulty of various types of scheduling tasks; determine the task difficulty based on the success rate of each candidate trajectory for each type of scheduling task; and obtain the target trajectory corresponding to at least the target type of scheduling task from the experience base to obtain each historical trajectory.

[0170] In some embodiments, the policy update module is further configured to determine the relative advantage of each current trajectory based on its respective trajectory reward for each current trajectory belonging to the same type of scheduling task; the trajectory reward is determined based on the departure rhythm, path conflict, delay risk and energy consumption weighted average of the current trajectory; and the first loss is determined based on the relative advantage of each current trajectory through the near-end policy optimization algorithm.

[0171] In some embodiments, each historical trajectory includes a first historical trajectory belonging to the target trajectory and a second historical trajectory not belonging to the target trajectory; the policy update module is further configured to determine the relative advantage of each historical trajectory based on its respective trajectory reward for each historical trajectory belonging to the same type of scheduling task; the trajectory reward is determined based on the departure rhythm, path conflict, delay risk and energy consumption weighting of the historical trajectory; the loss of each first historical trajectory is obtained based on the relative advantage and probability ratio of each first historical trajectory; the loss of each second historical trajectory is obtained based on the relative advantage of each second historical trajectory through a near-end policy optimization algorithm; and a second loss is determined based on the loss of each first historical trajectory and the loss of each second historical trajectory.

[0172] In some embodiments, the scheduling control module 308 is further configured to determine the next moment as the current moment and return to the step of obtaining the scheduling environment data of the heavy-load train at the current moment, until the scheduling task for the heavy-load train is completed.

[0173] Each module in the aforementioned virtual coupling and rolling scheduling optimization device for heavy-haul trains can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.

[0174] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 4 As shown, the computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores various data involved in the virtual coupling and rolling scheduling optimization method for heavy-haul trains. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a virtual coupling and rolling scheduling optimization method for heavy-haul trains.

[0175] Those skilled in the art will understand that Figure 4The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0176] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0177] In one embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0178] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0179] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0180] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0181] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0182] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for optimizing the virtual coupling and rolling scheduling of heavy-haul trains, characterized in that, The method includes: Obtain the scheduling environment data of heavy-load trains at the current moment, and construct a state diagram at the current moment based on the scheduling environment data; the nodes in the state diagram include train nodes and resource nodes, and the edges in the state diagram are used to represent the interaction relationships between the nodes; Based on a pre-trained graph neural network model, the node features of each node in the state graph are propagated at least one layer to obtain global state features. Based on a pre-trained policy network, a scheduling action decision for the current moment is generated according to the global state features; According to the dispatching action decision, virtual coupling dispatching control is performed on the heavy-haul train at the current moment, and then rolled to the next moment for virtual coupling dispatching control, until the dispatching task for the heavy-haul train is completed.

2. The method according to claim 1, characterized in that, The pre-trained graph neural network model performs at least one layer of propagation on the node features of each node in the state graph to obtain global state features, including: The initial features of each node and edge in the state graph are extracted based on a pre-trained graph neural network model. Based on the initial features of each node and each edge, the node features of each node are propagated at least one layer to obtain the target features of each node; wherein, in the propagation at each layer, for each node, the initial features of the target node in the next layer are determined based on the initial features of the target node in this layer, the initial features of the neighboring nodes of the target node in this layer, and the edge features between the target node and the neighboring nodes. Based on the target features of each node, the global state features are aggregated.

3. The method according to claim 2, characterized in that, The step of determining the initial features of the target node in the next layer based on the initial features of the target node in this layer, the initial features of the target node's neighboring nodes in this layer, and the edge features between the target node and its neighboring nodes includes: Based on the edge characteristics between the target node and its neighboring nodes, determine the weight adjustment factors; Based on the initial characteristics of the target node in this layer, the initial characteristics of the target node's neighboring nodes in this layer, and each of the weight adjustment factors, the initial characteristics of the target node in the next layer are determined.

4. The method according to claim 1, characterized in that, The method further includes: Each current trajectory obtained from online sampling is acquired, and a first loss is determined based on each current trajectory; the current trajectory is used to record the complete scheduling process of virtual coupled scheduling control for the corresponding scheduling task through the policy network; Each historical trajectory is obtained from the experience base, and a second loss is determined based on each historical trajectory; the historical trajectory is used to record the complete scheduling process of virtual coupled scheduling control for historical scheduling tasks; The target loss is obtained based on the first loss and the second loss, and the parameters of the policy network are updated based on the target loss to obtain the updated policy network; the updated policy network is used for virtual coupling scheduling control for future scheduling tasks.

5. The method according to claim 4, characterized in that, The method further includes: For each candidate trajectory corresponding to various types of scheduling tasks in the experience base, the behavioral entropy of each candidate trajectory is determined based on the probability of selecting an actual scheduling action at each step in the candidate trajectory and the number of trajectory steps of the candidate trajectory. Based on the behavioral entropy of each candidate trajectory, the target trajectory of various types of scheduling tasks is determined. The process of obtaining historical trajectories from the experience base includes: Based on the task difficulty of various types of scheduling tasks, a target type is determined from the various types of scheduling tasks; the task difficulty is determined based on the proportion of successful trajectories of each candidate trajectory of each type of scheduling task. From the experience base, at least the target trajectory corresponding to the scheduling task of the target type is obtained to obtain each historical trajectory.

6. The method according to claim 4, characterized in that, The determination of the first loss based on each of the current trajectories includes: For each current trajectory belonging to the same type of scheduling task, the relative advantage of each current trajectory is determined based on its respective trajectory reward; the trajectory reward is determined based on the departure rhythm, path conflict, delay risk and energy consumption weighted average of the current trajectory. The first loss is determined by using a near-end strategy optimization algorithm based on the relative advantages of each current trajectory.

7. The method according to claim 4, characterized in that, Each of the historical trajectories includes a first historical trajectory belonging to the target trajectory and a second historical trajectory not belonging to the target trajectory; the determination of the second loss based on each of the historical trajectories includes: For each historical trajectory belonging to the same type of scheduling task, the relative advantage of each historical trajectory is determined based on its respective trajectory reward; the trajectory reward is determined based on the departure rhythm, path conflict, delay risk and energy consumption weighted average of the historical trajectory. Based on the relative advantage and probability ratio of each first historical trajectory, the loss of each first historical trajectory is obtained; By using a proximal strategy optimization algorithm, the loss of each second historical trajectory is obtained based on the relative advantages of each second historical trajectory. The second loss is determined based on the loss of each of the first historical trajectories and the loss of each of the second historical trajectories.

8. The method according to any one of claims 1 to 7, characterized in that, The process of rolling to the next moment for virtual coupling and scheduling control until the scheduling task for the heavy-haul train is completed includes: The next moment is determined as the current moment, and the process returns to the step of obtaining the scheduling environment data of the heavy-haul train at the current moment, until the scheduling task for the heavy-haul train is completed.

9. A virtual coupling and rolling scheduling optimization device for heavy-haul trains, characterized in that, The device includes: A state graph construction module is used to acquire the scheduling environment data of heavy-load trains at the current moment, and construct a state graph at the current moment based on the scheduling environment data; the nodes in the state graph include train nodes and resource nodes, and the edges in the state graph are used to represent the interaction relationships between the nodes; The feature propagation module is used to propagate the node features of each node in the state graph at least one layer based on the pre-trained graph neural network model to obtain global state features. An action decision module is used to generate a scheduling action decision for the current moment based on the global state features, using a pre-trained policy network. The scheduling control module is used to perform virtual coupling scheduling control for the heavy-haul train at the current time according to the scheduling action decision, and roll to the next time to perform virtual coupling scheduling control until the scheduling task for the heavy-haul train is completed.

10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 8.