Train operation distribution method and device, computer equipment, medium and product
By constructing joint tensors and node agents, and combining local action inference and extreme scenario analysis, the consistency and robustness issues of cross-regional scheduling in railway train operation scheduling are solved. This achieves multi-round federated collaboration with consistent information and controlled latency, thereby improving the real-time performance and cross-regional resource optimization capabilities of railway train scheduling.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-03
- Publication Date
- 2026-04-07
AI Technical Summary
Existing railway train operation scheduling schemes suffer from poor real-time consistency, insufficient robustness in cross-regional coupling and extreme scenarios in large-scale, cross-regional scenarios. They are unable to provide a unified and real-time yielding strategy when multiple nodes experience train track occupancy conflicts simultaneously, and the lack of information consistency leads to frequent backtracking at boundary stations.
By acquiring hourly streaming data to construct a joint tensor, node agents are obtained. Combining local action inference and extreme scenario perturbation analysis, candidate action sets are determined, and differential incentive broadcasting and priority iterative sorting are performed until the overall pressure is less than a preset threshold. Combining policy gradient global weighting and emergency action sets, an executable plan and capacity margin parameters are generated to achieve information consistency and latency-controlled multi-round federated collaboration.
It significantly improves the robustness of cross-regional coupling and extreme scenarios, realizes multi-round federated collaboration with consistent information and controlled latency, and enhances the real-time consistency of railway train scheduling and the optimization capability of cross-regional resources.
Smart Images

Figure CN121799476A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of train dispatching technology, and in particular to a train operation allocation method, apparatus, computer equipment, medium and product. Background Technology
[0002] With the rapid development of science and technology, railway transportation has become an indispensable mode of transport. The development of algorithms for railway transportation resource allocation has roughly gone through a gradual process from centralized rule tables and mathematical programming to distributed intelligent agents. Although railway train operation and track allocation have progressed from centralized rule tables to mixed integer programming and then to distributed multi-agent reinforcement learning, existing algorithm systems still show multiple limitations in large-scale, cross-regional scenarios.
[0003] Centralized rule-based scheduling relies on fixed priorities and manual intervention, making it impossible to provide a unified and real-time yielding strategy when multiple nodes experience simultaneous train track occupancy conflicts. Insufficient information consistency leads to frequent backtracking at boundary stations. While integer programming models on spatiotemporal networks can approximate the global optimum under static data, the dimensionality of variables expands exponentially with the line and time period. Engineering implementation necessitates regional partitioning or rolling solutions, resulting in weakened cross-regional resource coupling and delayed re-optimization during extreme blockages or demand surges, making it difficult to meet minute-level scheduling windows.
[0004] Therefore, railway train operation scheduling schemes still suffer from defects such as poor real-time consistency, cross-regional coupling, and varying robustness in extreme scenarios. Summary of the Invention
[0005] Therefore, it is necessary to provide a train operation allocation method, device, computer equipment, medium, and product to address the aforementioned technical problems.
[0006] This application provides a train operation allocation method, the method comprising: acquiring hourly stream data, constructing a joint tensor and writing it into a scheduling shared memory unit; wherein, the hourly stream data includes at least one of train position, track occupancy, port unloading quota, mine loading cycle time, locomotive cycle time and temporary blockade; segmenting node agents according to the joint tensor, and performing local action deduction based on the node agents to determine a candidate action set; synchronously receiving the candidate actions and the joint tensor, calculating differential excitation and comprehensive pressure, broadcasting the differential excitation into the candidate action set, performing priority iterative sorting on the candidate action set until the comprehensive pressure is less than a preset pressure threshold, and determining the corrected action set; and based on the target historical event model... Using a template and random parameters, extreme scenario disturbance analysis and emergency game analysis are performed to determine the extreme case tensor and emergency action set. Global weighting of the strategy gradient is performed based on the extreme case tensor, and cyclic convergence analysis is conducted using the corrected action set and the emergency action set to determine a preliminary train allocation strategy. This preliminary train allocation strategy is then fed into an elastic train formation compressor to generate an executable plan and capacity margin parameters, which are written into the scheduling release channel and pushed to the field execution terminal. The executable plan and capacity margin parameters are monitored based on real-time tensors, and if preset trigger conditions are met, the process returns to the previous steps of performing extreme scenario disturbance analysis and emergency game analysis based on the target historical event template and random parameters to determine the extreme case tensor and emergency action set.
[0007] In one embodiment, constructing the joint tensor includes: retaining the original timestamps of the hourly stream data and writing them into a cache unit in the order of receipt; performing coordinate mapping on the data written into the cache unit to determine the coordinate mapping relationship; and filling tensor units sequentially according to the coordinate mapping relationship in a monotonic order of the time axis, spatial node axis, and resource axis to obtain a three-dimensional joint tensor.
[0008] In one embodiment, the step of obtaining node agents based on the joint tensor segmentation and determining candidate action sets by combining the node agents with local action deduction includes: segmenting the three-dimensional joint tensor along the spatial node axis to obtain different node slices; generating node agents by combining the mapping function and the node slices; wherein, the node slice represents all resource parameters of the same spatial node in the same time window; concatenating the local node and two-hop neighborhood slices of the node agent, and then performing graph attention and gated loop processing to obtain a local temporal vector; performing priority deduction calculation based on the local temporal vector and preset constraints to obtain an action set; verifying the feasibility of the action set to determine a feasible action set; performing cross-node impact assessment analysis based on the feasible action set to determine a cross-node impact weight matrix; and determining a candidate action set by combining the cross-node impact weight matrix and the feasible action set.
[0009] In one embodiment, the process of synchronously receiving the candidate actions and the joint tensor, calculating differential incentives and overall pressure, broadcasting the differential incentives into the candidate action set, iteratively prioritizing the candidate action set until the overall pressure is less than a preset pressure threshold, and determining the corrected action set includes: writing the candidate action set and the joint tensor uploaded by each node agent into a federated critic; dividing the joint tensor into regional blocks according to preset rules, and calculating the overall pressure by combining the skylight pressure, track congestion, and empty vehicle shortage; calculating the differential incentives of relevant nodes in the pressure-overhead section, and injecting the differential incentives into the node agents to determine the corrected priority score of each candidate action set; re-sorting the candidate action set according to the corrected priority score to obtain a global action set; writing the global action set into the federated critic for iterative calculation until the overall pressure is less than a preset pressure threshold, then using the global action set as the corrected action set and writing it into the scheduling and publishing channel.
[0010] In one embodiment, the step of performing extreme scenario perturbation analysis and emergency game analysis based on the target historical event template and random parameters to determine the extreme case tensor and emergency action set includes: performing Hadamard perturbation analysis based on the target historical event template and random parameters to determine the extreme case tensor; writing the extreme case tensor into the federated critic, and combining the node agent and risk penalty vector to perform risk calculation for each region block to determine a comprehensive risk score; combining the comprehensive risk score to perform game adversarial activities to generate emergency actions; and performing audit analysis on the emergency actions, and when the error compression rate meets a preset threshold compression rate, combining the trigger condition vector and the current emergency action to determine the emergency action set and write it into the scheduling and release channel.
[0011] In one embodiment, the step of performing global policy gradient adjustment based on the extreme case tensor and performing cyclic convergence analysis based on the corrected action set and the emergency action set to determine the preliminary train allocation strategy includes: determining local action triples based on the extreme case tensor and the preset incentive weights of each node agent; writing the local action triples and the extreme case tensor into the federated critic, updating the node incentive weights using policy gradients, and determining the delay increment; if the delay increment is greater than or equal to a preset increment threshold, calling the perturbation generator to apply pressure and returning to execute the step of performing extreme scenario perturbation analysis and emergency game analysis based on the target historical event template and random parameters to determine the extreme case tensor and the emergency action set, until the delay increment is less than the preset increment threshold, and then determining the preliminary train allocation strategy based on the current emergency action set and the corrected action set.
[0012] In one embodiment, the step of sending the preliminary train allocation strategy into the flexible formation compressor to generate an executable plan and capacity margin parameters, writing it into the scheduling release channel and pushing it to the field execution terminal includes: sending the preliminary train allocation strategy into the flexible formation compressor and locating the cross-regional meeting section; establishing a time window sliding index for the meeting section, sequentially performing sliding, single-time shrinkage insertion and rule verification, generating an executable plan and capacity margin parameters, writing it into the scheduling release channel and pushing it to the field execution terminal.
[0013] In one embodiment, the method further includes: comparing the real-time tensor periodicity of train operation with the executable plan to determine the regional empty train rate, the key port congestion index, and the train serial waiting time; locking the tensor snapshot when a preset trigger condition is met based on the regional empty train rate, the key port congestion index, and the train serial waiting time; and, based on the tensor snapshot, returning to the step of performing extreme scenario disturbance analysis and emergency game analysis based on the target historical event template and random parameters to determine the extreme case tensor and emergency action set, so as to generate a corrected executable plan and capacity margin parameters within a preset time.
[0014] A train operation allocation device, comprising: a tensor construction module for acquiring hourly stream data, constructing a joint tensor, and writing it into a scheduling shared memory unit; wherein the hourly stream data includes at least one of train position, track occupancy, port unloading quota, mine loading cycle time, locomotive cycle time, and temporary blockade; a local decision module for segmenting node agents based on the joint tensor, and performing local action deduction based on the node agents to determine a candidate action set; a pressure assessment module for synchronously receiving the candidate actions and the joint tensor, calculating differential excitation and comprehensive pressure, broadcasting the differential excitation into the candidate action set, performing priority iterative sorting on the candidate action set until the comprehensive pressure is less than a preset pressure threshold, and determining a corrected action set; and an extreme scenario deduction module for performing deduction based on a target historical event template and random events. The system comprises several modules: a train allocation module and a train decomposition module. The train allocation module performs extreme scenario disturbance analysis and emergency game analysis based on machine parameters to determine the extreme case tensor and emergency action set. A convergence optimization module performs global weighting of the strategy gradient based on the extreme case tensor and performs cyclic convergence analysis based on the corrected action set and the emergency action set to determine a preliminary train allocation strategy. A plan generation module sends the preliminary train allocation strategy to a flexible train formation compressor to generate an executable plan and capacity margin parameters, writes them to the scheduling release channel, and pushes them to the field execution terminal. A real-time monitoring module monitors the executable plan and capacity margin parameters based on real-time tensors and, when preset trigger conditions are met, returns to control the extreme scenario deduction module to execute the operation of performing extreme scenario disturbance analysis and emergency game analysis based on the target historical event template and random parameters to determine the extreme case tensor and emergency action set.
[0015] A computer device includes a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps of the method described above.
[0016] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described above.
[0017] A computer program product includes a computer program that, when executed by a processor, implements the steps of the method described above.
[0018] The aforementioned train operation allocation method, device, computer equipment, storage medium, and computer program product first establish a joint tensor by combining at least one of the following: train position, track occupancy, port unloading quota, mine loading cycle time, locomotive cycle time, and temporary blockade. This tensor is then written into the scheduling shared memory unit, enabling one-time alignment of the entire network state within milliseconds. This provides a synchronous, frame-free decision-making foundation for parallel node agents and lays the groundwork for a low-latency shared data environment for subsequent global-local coupling. Next, node agents are obtained by segmenting the joint tensor, and local action deduction is performed using these agents to determine candidate action sets. Candidate actions and the joint tensor are received synchronously, and differential excitations and comprehensive pressure are calculated. The differential excitations are broadcast into the candidate action sets, and the candidate action sets are iteratively prioritized until the comprehensive pressure is less than a preset pressure threshold. A corrected action set is then determined, achieving information consistency and latency-controlled multi-round federated collaboration.
[0019] Furthermore, target historical event templates and random parameters are introduced to conduct extreme scenario perturbation analysis and emergency game analysis, determining the extreme case tensor and emergency action set. Finally, the extreme case tensor is combined with global policy gradient adjustment, and the modified action set and emergency action set are combined with cyclic convergence analysis to determine the preliminary train allocation strategy. The preliminary train allocation strategy is then sent to the elastic train formation compressor to generate an executable plan and capacity margin parameters, which are written into the scheduling release channel and pushed to the field execution end. This significantly improves the robustness of cross-regional coupling and extreme scenarios.
[0020] Furthermore, real-time tensors can be used to monitor the executable plan and capacity margin parameters, and when preset trigger conditions are met, the reused link can be triggered to update the executable plan and capacity margin parameters, thereby achieving continuous operation of the algorithm-scheduling integration. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a schematic diagram of a train operation allocation method in one embodiment of this application;
[0023] Figure 2 This is a schematic diagram of the tensor construction process in one embodiment of this application;
[0024] Figure 3 This is a schematic diagram of a partial decision-making process in one embodiment of this application;
[0025] Figure 4 This is a schematic diagram of the stress assessment process in one embodiment of this application;
[0026] Figure 5 This is a schematic diagram of the network architecture in one embodiment of this application;
[0027] Figure 6 This is a schematic diagram of an extreme simulation process in one embodiment of this application;
[0028] Figure 7 This is a schematic diagram of the planned generation process in one embodiment of this application;
[0029] Figure 8 This is a schematic diagram of the train operation allocation device structure in one embodiment of this application;
[0030] Figure 9 This is an internal structural diagram of a computer device in one embodiment of this application. Detailed Implementation
[0031] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0032] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.
[0033] The train operation allocation method provided in this application embodiment can be applied to the railway transportation system to schedule and regulate train operation. Specifically, it can be applied in the train dispatching center.
[0034] Please see Figure 1 This application provides a train operation allocation method, which includes steps 101, 102, 103, 104, 105, 106 and 107.
[0035] Step 101: Obtain hourly streaming data, construct a joint tensor, and write it into the scheduling shared memory unit.
[0036] Step 102: Obtain node agents based on joint tensor segmentation, and perform local action inference based on node agents to determine candidate action sets.
[0037] This week, 103, candidate actions and joint tensors are received simultaneously. Differential excitations and comprehensive pressure are calculated. The differential excitations are broadcast into the candidate action set. The candidate action set is iteratively sorted by priority until the comprehensive pressure is less than the preset pressure threshold. The corrected action set is then determined.
[0038] Step 104: Based on the target historical event template and random parameters, perform extreme scenario perturbation analysis and emergency game analysis to determine the extreme case tensor and emergency action set.
[0039] Step 105: Perform global weighting of the policy gradient by combining the extreme case tensor, and perform cyclic convergence analysis by combining the corrected action set and the emergency action set to determine the preliminary train allocation strategy.
[0040] Step 106: The preliminary train allocation strategy is sent to the flexible formation compressor to generate an executable plan and capacity margin parameters, which are then written into the scheduling release channel and pushed to the field execution terminal.
[0041] Step 107: Monitor the executable plan and capacity margin parameters based on real-time tensors.
[0042] If the preset triggering conditions are met, return to step 104.
[0043] Specifically, hourly streaming data includes at least one of the following: train location, track occupancy, port unloading quota, mine loading cycle time, locomotive turnaround time, and temporary closure. Train location refers to the real-time geographical location of a train within the railway network, typically tracked precisely by the railway dispatching system using technologies such as track circuits, GPS positioning, or transponders. Track occupancy refers to the status of a specific numbered track (railway) being occupied by a train or locomotive within the railway transportation system. Port unloading quota is a restrictive indicator set by port management departments to balance cargo throughput and optimize resource allocation. It specifies the maximum quantity of a certain type of cargo (such as coal or containers) that a port can unload within a specific period (such as a month or quarter). Mine loading cycle time is a key efficiency indicator in the mining transportation system, referring to the time interval required to complete one mine car loading operation per unit of time. It reflects the collaborative efficiency of loading equipment (such as loaders and conveyor belts) and transportation processes (such as mine car dispatching and track operation). Locomotive turnaround time (average total locomotive turnaround time) is a core indicator for measuring the efficiency of railway locomotive operation, referring to the average time required for a locomotive to complete one round trip operation within its traction section. Temporary closures are emergency safety measures taken in railway operations to respond to unforeseen events (such as equipment failures, natural disasters, and safety accidents). It refers to closing a section of track or station for a specific period, prohibiting train traffic to ensure personnel safety and facilitate repair work. In a more detailed embodiment, hourly stream data simultaneously includes train location, track occupancy, port unloading quotas, mine loading cycle time, locomotive cycle time, and temporary closures.
[0044] Scheduling shared memory units is an efficient inter-process communication (IPC) mechanism that allows multiple processes to directly access the same physical memory region, enabling high-speed data sharing and transfer. A joint tensor, also known as a tensor obtained by combining different dimensions, is a multilinear mapping on a vector space and its dual space. It can be represented as a multidimensional array, and its core characteristics include order (rank), data type, and covariance under coordinate transformations. Node agents are the core components of multi-agent systems. They are distributed nodes with autonomy, interactivity, and task processing capabilities, capable of perceiving the environment, executing decisions, and collaborating with other nodes or the central system to complete complex tasks. In practical scenarios, for the same node agent, candidate action sets can be obtained by splicing together each node and its two-hop neighborhood slices, and then analyzing and calculating them separately.
[0045] In this embodiment, the federated critic can simultaneously receive candidate actions and joint tensors, and then divide the region according to certain rules (such as according to administrative boundaries or the direction of railway trunk lines), solve the comprehensive pressure for each region, broadcast differential excitation, and iteratively sort the candidate actions until the comprehensive pressure is less than a preset pressure threshold. The iteratively sorted set of candidate actions is then used as the corrected set of actions.
[0046] The target historical event template refers to a template library built based on target historical events. These templates will vary depending on the specific target historical event, as long as they can represent the railway operating status under extreme scenarios. For example, in one embodiment, the target historical event is the worst 5% of historical events, that is, the events representing the worst 5% of railway operating events. Random parameters are the pre-set random parameters in the disturbance generator. Their type is not unique; in one embodiment, they include parameters with three dimensions: random address, duration, and resource reduction degree.
[0047] Therefore, the perturbation generator By combining the above three-dimensional parameters and the target historical event template, segment mapping can be performed to obtain the preservation mask. This generates an extreme case tensor. Subsequently, the extreme case tensor and the set of emergency actions are simultaneously fed into the federated critic, creating a real-time adversarial landscape of punishment-driven, emergency-response hedging, ultimately resulting in an adversarial game that produces the emergency action set.
[0048] After that, the federated critic updates the incentive weights of the nodes using policy gradients, and repeatedly performs the above operation by increasing the resource reduction parameter when the delay is still high, until the delay of the entire network can no longer be significantly increased, and obtains a converged global action set (including the corrected action set obtained in normal scenarios and the emergency action set obtained in extreme cases). Finally, a preliminary train allocation strategy is established with the global action set to achieve the dual goals of controlled delay and balanced resources in extreme scenarios.
[0049] Then, the initial train allocation strategy is fed into the flexible formation compressor. Each train section is indexed by its operating area code and track affiliation number. Cross-regional intersection sections are located one by one, and a time-slip index is established. After performing slip, single-time shortening and insertion, and rule verification processes in sequence, the final executable plan and capacity margin parameters can be obtained. These are then written into the dispatching release channel and pushed to the field execution terminal, thereby realizing the dispatching of train operations.
[0050] During scheduling, the executable plan and capacity margin parameters can be monitored using real-time tensors, and preset trigger conditions can be established based on preset trigger thresholds. The system continuously compares the real-time tensors and the executable plan to determine if the trigger conditions are met. Upon triggering, the system returns to step 104, reusing steps 104, 105, and 106 to generate an updated plan (i.e., the updated executable plan and capacity margin parameters) within a preset time (e.g., 10 minutes) and then executes it.
[0051] The aforementioned train operation allocation method first establishes a joint tensor by combining at least one of the following: train position, track occupancy, port unloading quota, mine loading cycle time, locomotive cycle time, and temporary blockade. This tensor is then written into the scheduling shared memory unit, enabling one-time alignment of the entire network state within milliseconds. This provides a synchronous, frame-free decision-making foundation for parallel node agents and lays the groundwork for a low-latency shared data environment for subsequent global-local coupling. Next, node agents are segmented using the joint tensor, and local action deduction is performed to determine a candidate action set. Candidate actions and the joint tensor are received synchronously, and differential excitations and overall pressure are calculated. The differential excitations are broadcast into the candidate action set, and the candidate action set is iteratively prioritized until the overall pressure is less than a preset pressure threshold. A corrected action set is then determined, achieving information consistency and latency-controlled multi-round federated collaboration.
[0052] Furthermore, target historical event templates and random parameters are introduced to conduct extreme scenario perturbation analysis and emergency game analysis, determining the extreme case tensor and emergency action set. Finally, the extreme case tensor is combined with global policy gradient adjustment, and the modified action set and emergency action set are combined with cyclic convergence analysis to determine the preliminary train allocation strategy. The preliminary train allocation strategy is then sent to the elastic train formation compressor to generate an executable plan and capacity margin parameters, which are written into the scheduling release channel and pushed to the field execution end. This significantly improves the robustness of cross-regional coupling and extreme scenarios.
[0053] Furthermore, real-time tensors can be used to monitor the executable plan and capacity margin parameters, and when preset trigger conditions are met, the reused link can be triggered to update the executable plan and capacity margin parameters, thereby achieving continuous operation of the algorithm-scheduling integration.
[0054] Please see Figure 2 In one embodiment, the construction of the joint tensor includes steps 201, 202 and 203.
[0055] Step 201: Retain the original timestamps of the hourly stream data and write them into the buffer unit in the order of receipt.
[0056] Step 202: Perform coordinate mapping on the data written to the cache unit to determine the coordinate mapping relationship.
[0057] Step 203: Based on the coordinate mapping relationship, fill the tensor units sequentially according to the monotonic order of the time axis, spatial node axis, and resource axis to obtain the three-dimensional joint tensor.
[0058] Specifically, this embodiment uses the time axis, spatial node axis, and resource axis as three axes to establish a three-dimensional tensor, that is, a three-dimensional joint tensor.
[0059] First, the six types of hourly streaming data scattered across various regions on the scheduling data bus—namely, train position, track occupancy, port unloading quota, mine loading cycle time, locomotive cycle time, and temporary blockade—are written into a unified buffer according to the order of receipt. Assign a unique sequence number to each written record, preserving the original timestamp. .
[0060] Then the buffer Each record Coordinate mapping is performed using a triplet of time period index, node index, and resource index. Used as the unique reference coordinate; where, This is a continuous time-period index with hourly granularity. Existing station and industrial / mining / port codes, This corresponds to the entity categories of vehicles, tracks, and loading / unloading points. During the mapping process, all heterogeneous names such as train positions and track occupancy are aligned to [the appropriate categories]. A coordinate system enables cross-business system mutual recognition.
[0061] After obtaining the complete coordinate mapping, according to Axis (time), Axis (spatial node) The monotonic sequence of axes (resource types) is used to fill tensor cells sequentially, generating a three-dimensional joint tensor. ;in, Indicates the time period ,node ,resource The only possible state value.
[0062] Will Write to the scheduled shared memory ,in, (This is a shared memory area for all intelligent agents). Based on this, the intelligent agents at each node in the mining area, trunk line, and port directly reference this when formulating vehicle yielding, track allocation, and train priority adjustment. A unified platform. This eliminates conflicts arising from information silos, such as competition for tracks and duplicate shunting orders, at the decision-making level.
[0063] Please see Figure 3 In one embodiment, step 102 includes steps 301, 302, 303, 304 and 305.
[0064] Step 301: Divide the three-dimensional joint tensor along the spatial node axis to obtain different node slices, and combine the mapping function and node slices to generate a node agent.
[0065] Among them, a node slice represents all resource parameters of the same spatial node in the same time window.
[0066] Step 302: After splicing the local node and two-hop neighborhood slices of the node agent, the local temporal vector is obtained through graph attention and gated loop processing.
[0067] Step 303: Priority calculation is performed based on the local time-series vector and preset constraints to obtain the action set. The feasibility of the action set is verified to determine the feasible action set.
[0068] Step 304: Perform cross-node impact assessment analysis based on the action set to determine the cross-node impact weight matrix.
[0069] Step 305: Combine the cross-node influence weight matrix and the action set to determine the candidate action set.
[0070] Specifically, the three-dimensional joint tensor Cut along the spatial axis to obtain node slices ;in, A unique index for the node. Represents a node The status of all resources in the current time window. Hand over to the mapping function Generate unique agent instances And write the category tag "mining area trunk port / electricity" according to the node's business attributes. Through the aforementioned one-to-one mapping binding, all loading docks in mining areas, trunk line hubs, and ports / power plants simultaneously obtain independent decision-making entities, replacing the traditional single-center queuing logic and reserving precise node identity anchor points for cross-regional vehicle borrowing.
[0071] For those in the running state The local graph attention and gating loop is as follows: with two-hop neighborhood slices The local state matrix is constructed by splicing together the data. Then, by weighting its execution graph, we obtain... ,in, For nodes The weighted feature matrix; It is a non-linear activation function; For nodes The set of two-hop neighborhoods; The attention coefficient is calculated by combining the train occupancy conflict level, track remaining capacity, and loading and unloading demand. It is a trainable linear mapping matrix; neighboring nodes The state matrix.
[0072] Will The timing vector is obtained by feeding it into the ingress control loop unit. ( For nodes (The aggregated description of adjustable resources and potential bottlenecks within the current window) enables the agent to focus only on the three key business characteristics required for decision-making at this node, avoiding interference from irrelevant cross-regional information on the decision-making rhythm.
[0073] Will The input to the decision-making level involves parallel simulations of three types of operations: train formation sequence, track occupancy time, and vehicle relocation direction; for each candidate action... Calculate the overall priority score:
[0074]
[0075] in, Candidate actions Priority; To execute The length of the subsequent release of the stock channel period; To execute The capacity can be increased by using empty vehicles in the subsequent cross-regional areas; To execute Potential local train waiting times; The weighting constants are used to address the three demands: freeing up lanes, sharing vehicles, and reducing queuing.
[0076] For all according to Arrange in descending order to form an action set After including the executable time window and a list of affected nodes, it is written into the incentive broadcast channel. This enables the federal critic to capture high-conflict actions in real time on a global scale, thereby accurately injecting differential incentives and proactively resolving the predicament of border crossing blockades and vehicle idleness.
[0077] The action set obtained in the previous sub-step Each physical and technological constraint at this node is checked and verified, and an upper bound vector of the constraints is constructed. ;in This is the upper limit of the train length that a track can accommodate. This refers to the upper limit of the operating pressure that loading and unloading machinery can withstand. The upper limit of the available locomotive weekly budget, This is the maximum allowed power generation waiting time limit. For each action... Calculate the resource consumption vector Then perform a feasibility assessment: ,in, For action Feasibility indicator; For action The usage of four resources: length, pressure, cycle time, and waiting time; For nodes The upper bound of the corresponding resources that can be allocated within the same window; This is a vector element-wise non-greater than relation.
[0078] Only when At that time, retain Proceed to subsequent collaboration; otherwise, be eliminated. Ultimately, actions that do not meet the track length or locomotive cycle time limits will be eliminated at the source to prevent unfeasible plans from being broadcast globally and to avoid queuing at hubs due to secondary backtracking.
[0079] All retained actions (That is, the action set after removal) calculates the ratio of its saving or occupying of window space for adjacent nodes, and defines the cross-node influence weight matrix. ,in , For action neighboring nodes Sunroof release - occupancy ratio; For action The length of the time window for the track released at this node; For the same action, they are adjacent nodes. The length of the freed-up track window (if occupied, take a negative value); To prevent extremely small positive numbers with a denominator of zero.
[0080] Will After being packaged into action description frames (i.e., candidate action sets), excitation is broadcast. Synchronously update node private memory This allows the negotiated skylight information to be inherited in the next scrolling window. The matrix-based impact description can be directly used by the federated critic to identify high-conflict sections, quickly inject targeted incentives for yielding and vehicle sharing, and resolve two persistent problems: intersection blockages and resource idleness.
[0081] Please see Figure 4 In one embodiment, step 103 includes steps 401, 402, 403, 404 and 405.
[0082] Step 401: Write the candidate action set and joint tensor uploaded by each node agent into the federated critic.
[0083] Step 402: Divide the joint tensor into regional blocks according to preset rules, and calculate the overall pressure by combining the sunroof compression degree, track congestion degree, and empty car shortage degree.
[0084] Step 403: Calculate the differential excitation of the relevant nodes in the overburdened section, inject the differential excitation into the node agent, and determine the correction priority score of each candidate action set.
[0085] Step 404: Reorder the candidate action set according to the corrected priority score to obtain the global action set.
[0086] Step 405: Write the global action set into the federated critic for iterative calculation until the overall pressure is less than the preset pressure threshold. Then, use the global action set as the corrected action set and write it into the scheduling and release channel.
[0087] Specifically, references can be drawn together. Figure 5 Each node agent can analyze and obtain a corresponding set of candidate actions, and the candidate action sets submitted by all node agents are then processed. With global state tensor (That is, the aforementioned three-dimensional joint tensor) synchronously input into the federated critic .
[0088] right The region blocks are obtained by dividing the area into two levels: administrative boundaries and trunk road directions (the default rule in this embodiment is administrative boundaries and trunk road directions). Then for each In continuous time period Calculate the three business pressure indicators as the sunroof squeeze degree Stockway congestion Insufficient empty vehicles .
[0089] Using a weighted coupling formula: ,in, For the region During the period The overall stress score; To address the management weight constants for the three pressures of sunroof, track, and empty train, and to meet... ; For the region During the period The overlap rate of sunroofs; For the region During the period The occupancy saturation ratio of the stock lanes; For the region During the period The availability of empty vehicles.
[0090] For all Extract the coordinates of the maximum value in descending order. Subsequently, differential excitation was calculated for the relevant nodes in the overburdened section. ,in , Complete the closed loop of conflict localization → incentive quantification, and lock in the key contact points of mutual blockage and idle vehicles.
[0091] Differential excitation via broadcast bus Inject the value assessment module into each node's intelligent agent; for each node Original action score in the current window Superimposed incentives form a corrected priority score: ,in, To incentivize the revised action priority score; For nodes During the period The original ranking score for local actions; To target indicators Assign nodes The differential excitation quantity.
[0092] Through the above linear superposition, each node incorporates the reduction of sunroof conflicts and the improvement of empty vehicle sharing into the decision-making basis during the sorting stage, driving adjacent areas to spontaneously give way and vehicles to complement each other, avoiding two-way lane occupation at the intersection and long-distance empty vehicle return trips.
[0093] Given a set of actions that have been injected with stimuli and rearranged, each node agent selects a new optimal solution. (That is, the global action set) and send it back. ;Criticism of the entire set Recalculate the remaining comprehensive pressure .
[0094] like Still above the target threshold ,but The incentive weights of the corresponding nodes—indicators—are adaptively adjusted, and the differential incentives are rebroadcast; this process is iterated until… Then the final motion set (That is, the revised action set) is written into the scheduling and publishing channel. This enables the coordinated placement of train and vehicle flows across multiple regions within the same window, completely eliminating barriers to cross-regional command exclusion.
[0095] Based on the above scheme, node agents are automatically instantiated along the spatial axis slices. Each agent concatenates its own node slice and two-hop neighborhood slices, then inputs them into a graph attention network to extract topologically relevant features. This is followed by gated loop modeling of temporal dependencies, resulting in a sparse vector containing only key components such as occupancy conflict degree, remaining capacity, and loading / unloading demand. Candidate actions are evaluated simultaneously using a linear weighting function, considering lane release volume, empty vehicle sharing increment, and local waiting time. Cross-node window impact coefficients are embedded within the action frame and sent after hard constraint filtering. The federated critic synchronously receives actions and a global tensor, calculates three indicators—window squeeze, lane congestion, and insufficient empty vehicles—and broadcasts differential incentives. This achieves consistent information and time-controlled multi-round federated collaboration, unlike traditional multi-agent frameworks that rely on a single global reward and struggle to quickly distribute credit.
[0096] Please see Figure 6 In one embodiment, step 104 includes steps 601, 602, 603 and 604.
[0097] Step 601: Perform Hadamard perturbation analysis based on the target historical event template and random parameters to determine the tensor for extreme cases.
[0098] Step 602: Write the extreme case tensor into the federated critic, combine the node agent and the risk penalty vector to calculate the risk of each region block and determine the comprehensive risk score.
[0099] Step 603: Combine the comprehensive risk score to conduct game-based confrontation and generate emergency actions.
[0100] Step 604: Audit and analyze the emergency actions. When the error compression rate meets the preset threshold compression rate, determine the emergency action set and write it into the scheduling and release channel by combining the trigger condition vector and the current emergency action.
[0101] Specifically, the target historical event template is the worst-case 5% event template. Historical scheduling logs are extracted by node and time period coordinates, identifying three high-risk segments: typhoon landfall, mainline outage, and sudden demand surge, and written into the worst-case 5% event template library. ;
[0102] For the disturbance generator Set random address Duration and resource reduction intensity The three-dimensional parameters (i.e., random parameters) are used to perform segment remapping on the selected template to obtain a capacity-preserving mask. And generate extreme tensors using the Hadamard product:
[0103]
[0104] in, Based on the three-dimensional joint tensor; It is the product of elements of the same dimension; For time period ,node ,resource The capacity retention factor; This represents the number of templates selected in this perturbation iteration; For the first The reduction factor for each template, with values ranging from... ; template The corresponding time period set; template The corresponding set of nodes; template The set of resource categories involved; This is an indicator function.
[0105] Through this template, parameters, and masking link, any extreme event scenario can be injected into the global state, ensuring that subsequent game calculations have a consistent and controllable high-pressure input.
[0106] Will Simultaneously sent to the Federal Critic With all node agents Enable risk penalty vector for critics For each region block Time period Calculate the risk score: ,in, For the region During the period The overall risk score; This refers to the additional delay time after introducing the disturbance; This is to reduce the available track time after the introduction of disturbances; To account for the increase in the shortage of empty vehicles after the disturbance; The penalty weights are assigned to the three risks of delays, track closures, and empty trains.
[0107] Criticism tool Replace routine indicators in action evaluation; at the same time, for each Limit search depth It requires that candidates be selected from only three emergency action pools: congestion relief, decentralized loading, and early return. The action descriptions must include positive and negative transmission quantities to neighboring nodes, thereby creating a real-time confrontational pattern of punishment-driven and emergency countermeasures.
[0108] Emergency actions generated by adversarial games The critic will audit each item individually; if the target node's delay compression rate is high... Below the threshold Callback immediately promote or Applying pressure to the same segment leads to the next round of the game; if all If so, the current action set is frozen, and the condition vector will be triggered. and (That is, the emergency action set) is written into the scenario memo. And send it to the scheduling and publishing channel.
[0109] The solution in this embodiment, through a progressive chain of risk re-pressure, action re-examination, and condition solidification, can obtain an emergency script that can be directly implemented during the calculation stage, providing minute-level elastic recovery capability for real network operation.
[0110] In one embodiment, step 105 includes: combining the extreme case tensor and the preset incentive weights of each node agent to determine local action triples; writing the local action triples and the extreme case tensor into the federated critic, updating the node incentive weights using policy gradients, and determining the delay increment.
[0111] If the delay increment is greater than or equal to the preset increment threshold, the disturbance generator is invoked to increase pressure in Jining and return to execute the steps of extreme scenario disturbance analysis and emergency game analysis based on the target historical event template and random parameters to determine the extreme case tensor and emergency action set until the delay increment is less than the preset increment threshold. Then, the preliminary train allocation strategy is determined by combining the current emergency action set and the corrected action set.
[0112] Specifically, the tensor of extreme cases Node mapping (Obtained during the above mapping process) and synchronously sent to all node agents. For each agent, a predetermined incentive weight vector is set. Let it be within a limited iteration step size Internal output local action triplet: ,in, For nodes This round of decision-making action vector; The result of rearranging the vehicle grouping order; To allow for the start and end times of the sunroof to be adjusted; This refers to the direction and number of vehicle groups to be borrowed across regions.
[0113] Node agents simultaneously label resource correction amounts ,in For track occupancy correction time, This is a correction amount for loading and unloading rhythm, used to directly offset sudden queues triggered by typhoon blockades or surges in demand.
[0114] All local actions and Submit to the Federal Critic For each region block Summary of Delay Increment And identify the node with the longest delay chain. The critic employs a policy gradient: ,in, For nodes New incentive weights; For nodes The old incentive weights; The global learning rate; For nodes The gradient operator for the activation vector; This represents the total delay across the entire network.
[0115] The critic will The signal is sent back to the intelligent agent via the broadcast bus to enhance the understanding of... Unleash the value of sunroofs and shared vehicles, and prioritize compressing global latency peaks in the next iteration.
[0116] If the critic feedback indicates that there is still Above the threshold That is, the period of continuous increase in delays Call the perturbation generator Increase pressure and increase its resource reduction parameters And refresh the tensor for extreme cases. Then, the above operation is retried. This process is repeated multiple times until no new disturbance can significantly increase the level. The critic freezes the current global action set, combines the current emergency action set with the revised action set, and outputs an action that satisfies... Conditional vehicle allocation strategy The dispatch and release channel is available for real-world network calls, achieving the dual goals of controlled delays and balanced resources in extreme scenarios.
[0117] Please see Figure 7 In one embodiment, step 106 includes steps 701 and 702.
[0118] Step 701: The preliminary train allocation strategy is sent to the flexible train formation compressor, and the cross-regional meeting section is located.
[0119] Step 702: Establish a time window sliding index for the intersection section, and sequentially perform sliding, single-time shrinking insertion and rule verification to generate an executable plan and capacity margin parameters, write them into the scheduling release channel and push them to the field execution terminal.
[0120] Specifically, Feed into flexible group compressor Each train section is indexed according to its operating area code and track affiliation number to locate the inter-regional meeting sections one by one; a time window sliding index array is established for each meeting section. ,in Record the smallest time granularity at which forward or reverse sliding is possible. This is achieved by writing within the strategy. This preserves a precise operational channel for rearrangement and peak shaving, and resolves the conflict of competing for time windows between regional dispatch centers from a mechanism perspective.
[0121] Time window index for intersection segment The step-by-step sliding process employs a sequential scanning method to adjust train pairs with potential track-blocking conflicts to open positions. Simultaneously, it locks the original loading / unloading cycle time field to prevent swaying during loading / unloading due to sliding. Then, the boundary of the yield window after sliding is compared with adjacent dispatch windows to ensure it has not been compressed below a threshold. This chain-like action of single-step sliding, cycle time locking, and boundary correction releases cross-regional track openings without affecting the overall train schedule structure, reducing the risk of mutual blockages along the line.
[0122] The shifted train set is matched with the over-limit window list. For trains that still exceed the capacity threshold, a combination of reduction and insertion is used for compression. Specifically, the train formation length field is reduced to meet the remaining capacity vector of the local track. The minimum value is then used to insert trains according to the index order of available tracks, and the trains in the same direction are constrained to only undergo one reduction in number within a day to avoid insufficient loading caused by continuous reductions. This achieves peak shaving and valley filling of train flow during peak periods without disrupting the balance of actual freight volume on the network.
[0123] Compare the slip and reduction results with the track capacity table. and port shift schedule The three-way verification process involves immediately removing any entry if it has a loading / unloading sequence conflict or if a track is over-occupied, and rolling it back to the previous valid state. Entries that pass all verifications are then written into the executable plan. Simultaneously, the corresponding time window and capacity surplus item are recorded in the elasticity margin list. .
[0124] The published plan With real-time train operation status tensor A monitoring system is established to compare three indicators at fixed intervals: regional empty train rate, key port congestion index, and train waiting time. When the empty train rate imbalance exceeds [a certain threshold], a trigger is triggered. The port congestion index is greater than Or serial waiting time exceeds Every minute, the corresponding node and time period are marked as a scenario requiring intervention, and scenario tags are generated simultaneously. .
[0125] Will The specified current tensor snapshot, along with the node identity mapping, is sent to the perturbation generator at once. Federal critic It establishes links with all node intelligent agents; it conducts three consecutive rounds of adversarial calculations against the three parties, strictly limiting the total time to a preset duration (e.g., 10 minutes). Upon completion of the calculations, it directly generates an updated plan that includes three types of measures: vehicle rearrangement, yield window adjustment, and cross-regional borrowing, and immediately issues it for execution through the scheduling release channel to ensure that traffic flow, vehicle flow, and track flow can quickly restore balance in extreme scenarios.
[0126] In one embodiment, the method further includes: comparing the real-time tensor periodicity of train operation with the executable plan to determine the regional empty train rate, the congestion index of key ports, and the train serial waiting time; locking the tensor snapshot when a preset trigger condition is met based on the regional empty train rate, the congestion index of key ports, and the train serial waiting time; and, based on the tensor snapshot, returning to the step of performing extreme scenario disturbance analysis and emergency game analysis based on the target historical event template and random parameters to determine the extreme case tensor and emergency action set, so as to generate a corrected executable plan and capacity margin parameters within a preset time.
[0127] Specifically, the real-time train operation global state tensor (i.e., real-time tensors) and published executable plans Perform periodic comparisons; analyze regional empty vehicle rates. Congestion index at key ports Train waiting time Establish threshold triggers and set the criteria respectively. , , (minutes). For After index mapping, if two consecutive frames satisfy any criterion, the corresponding node and time period will be marked as... And lock the tensor snapshot Metric extraction, threshold comparison, and snapshot locking enable the system to capture the precise temporal and spatial coordinates of operational imbalances within minutes.
[0128] right Specified The computational chain of steps 104-106, which employs a three-way multiplexing path of perturbation generator → critic → multi-agent, is as follows: Send in Generate a risk pressure tensor; synchronously input the risk tensor. and Three rounds of adversarial iterations were initiated, with a limited total duration. Minutes; convergence at the iteration endpoint yields the updated measure package. It includes entries for vehicle rearrangement, yield window fine-tuning, and cross-regional borrowing. The established process of snapshot injection, three-round confrontation, and measure aggregation ensures that the system automatically generates a new, implementable solution within ten minutes in extreme scenarios, resolving the problems of window competition and short-distance vehicle return caused by fragmented regional scheduling.
[0129] Ultimately, Feed into flexible group compressor The execution involves three levels of compression: time window sliding, shrinking and insertion, and rule verification, to generate a revised plan. With the balance list The data is then written into the scheduling and publishing channel and pushed to the on-site execution terminal.
[0130] To enhance robustness in extreme scenarios, the above scheme introduces a perturbation generator based on the worst-case percentile template and Hadamard mask on the critic side. This generator reduces the capacity of the same tensor to achieve consistent adversarial training of the model input. A risk penalty vector replaces the normal incentive, driving each node to search only in a limited-depth emergency action pool, thus reducing network latency peaks within several rounds. The converged action set enters an elastic grouping compressor, which releases the capacity of intersection segments through time window sliding and single-stage shrinking and insertion. After rule verification, it is directly converted into an executable plan. Real-time closed-loop monitoring reuses the link after triggering a threshold, enabling adaptive recalculation within ten minutes and achieving continuous operation of the algorithm-scheduling integration.
[0131] It is understood that in other embodiments, tensor sharing can be replaced by batch asynchronous event streams, policies can be trained separately at each node and global reward averaging can be used; or convolutional networks or transformers can be used to replace graph attention, or integer programming can be used to form a hybrid decision by superimposing heuristics. The solutions of the embodiments of this application can also be implemented, and no specific limitation is made.
[0132] In a more detailed embodiment, the above train operation allocation method is verified:
[0133] At the end of 2023, a railway logistics group launched a project to improve the efficiency of integrated port, mining, and power transportation. The group operates three trunk lines exceeding 1,000 kilometers in length, six coal mine loading yards, two large ports, and one coastal power plant, with an average annual shipment volume exceeding 140 million tons. For a long time, the loading pace at the mines and the unloading speed at the ports have been mismatched, exacerbated by frequent monsoon season closures, leading to significant train queues at inter-regional border crossings and empty car return trips often exceeding 36 hours. The dispatch center urgently needed a collaborative optimization method that could be implemented within minutes. The group's information center, in collaboration with a research institution, introduced the group's integrated railway vehicle allocation method proposed in this application into a live network test. The data window selected over 3,000 train operation logs from October 2023 to February 2024, along with six types of hourly streaming data collected simultaneously.
[0134] During the data aggregation phase, the Flink stream processing platform receives six streams, including train location and track occupancy, in batches with a 10-second granularity. All streams are mapped to a three-axis coordinate system of time period, node, and resource. A 128×94×6-dimensional joint tensor is generated every hour and stored in Redis-SharedMemory. The average time to complete the mapping is 43 milliseconds, which meets the requirements of the rolling window.
[0135] The agent training employed a Ray cluster deployment of 102 agent nodes. Each agent fed its own node and two-hop neighborhood slices into a PyTorch-Geometric version of the GAT-GRU network. The policy network was trained on 2×A100 GPUs with a learning rate of 5e-4 for 18 hours until convergence. An average of 14.7 candidate actions were generated, of which 9.3 were retained after constraint filtering and simultaneously written to the incentive broadcast channel. The federated critic divided the region into 12 blocks based on administrative regions and trunk line directions, with a comprehensive pressure threshold set at 0.25. After three rounds of differential incentive iterations, the regional pressure decreased to below 0.17, and the computation-communication round-trip latency was controlled within 2.8 seconds.
[0136] Extreme disturbance verification selected 67 typhoon-related track closures and mainline bridge speed limit events from 2019 to 2022 to construct a worst-case 5% template library; the disturbance generator injected 11 capacity masks at once, and the critic output an emergency action set in 6 steps under risk penalty-driven depth search; the flexible train formation compressor reduced the overlap rate of the junction skylight from 0.46 to 0.28 through time window sliding and single-time reduction and insertion, and shortened the peak train serial waiting time by 42%.
[0137] Three months after its official launch, the average empty train return time from the mining area to the port decreased from 35.8 hours to 24.1 hours, the average queue length at the border decreased by 31%, and the maximum waiting time for loading ships at the port decreased from 11.5 hours to 5.2 hours. During the double typhoon lockdown in January 2024, the system automatically recalculated and issued executable instructions within 7 minutes and 40 seconds, maintaining the continuity of the train schedule without manual intervention. Real-time KPI monitoring showed that the number of consecutive triggers of the three indicators—empty train imbalance, port congestion index, and serial waiting time—decreased by 68% compared to the same period last year, fully verifying the applicability and stability of the proposed method in large network, extreme disturbance, and cross-regional coupling scenarios.
[0138] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.
[0139] Based on the same inventive concept, this application also provides a train operation allocation device for implementing the train operation allocation method described above. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations in one or more train operation allocation device embodiments provided below can be found in the limitations of the train operation allocation method described above, and will not be repeated here.
[0140] Please see Figure 8 A train operation allocation device includes: a tensor construction module 801, a local decision-making module 802, a stress assessment module 803, an extreme simulation module 804, a convergence optimization module 805, a plan generation module 806, and a real-time monitoring module 807.
[0141] Tensor construction module 801 is used to acquire hourly streaming data, construct a joint tensor, and write it into the scheduling shared memory unit; local decision module 802 is used to divide the joint tensor into node agents, and combine the node agents to perform local action inference to determine the candidate action set; stress assessment module 803 is used to synchronously receive candidate actions and joint tensor, calculate differential excitation and comprehensive stress, broadcast the differential excitation into the candidate action set, perform priority iterative sorting of the candidate action set until the comprehensive stress is less than the preset stress threshold, and determine the corrected action set; extreme scenario inference module 804 is used to perform extreme scenario disturbance analysis and emergency game analysis based on the target historical event template and random parameters to determine the extreme case tensor and emergency action set; The convergence optimization module 805 is used to perform global weighting of the strategy gradient by combining the extreme case tensor, and to perform cyclic convergence analysis by combining the corrected action set and the emergency action set to determine the preliminary train allocation strategy; the plan generation module 806 is used to send the preliminary train allocation strategy into the flexible formation compressor to generate an executable plan and capacity margin parameters, write them into the scheduling release channel and push them to the field execution terminal; the real-time monitoring module 807 is used to monitor the executable plan and capacity margin parameters according to the real-time tensor, and when the preset trigger conditions are met, it returns to the control extreme simulation module 804 to perform extreme scenario disturbance analysis and emergency game analysis based on the target historical event template and random parameters to determine the extreme case tensor and emergency action set.
[0142] In one embodiment, the tensor construction module 801 is also used to retain the original timestamps of the hourly streaming data and write them into the cache unit in the order of receipt; to perform coordinate mapping on the data written into the cache unit to determine the coordinate mapping relationship; and to fill the tensor unit in the monotonic order sequence of the time axis, spatial node axis and resource axis according to the coordinate mapping relationship to obtain a three-dimensional joint tensor.
[0143] In one embodiment, the local decision module 802 is further configured to segment the three-dimensional joint tensor along the spatial node axis to obtain different node slices, and generate node agents by combining the mapping function and the node slices; after splicing the local node and two-hop neighborhood slices of the node agent, the local temporal vector is obtained through graph attention and gating loop processing; priority inference calculation is performed based on the local temporal vector and preset constraints to obtain an action set; the feasibility of the action set is verified to determine the feasible action set; cross-node impact assessment analysis is performed based on the feasible action set to determine the cross-node impact weight matrix; and candidate action sets are determined by combining the cross-node impact weight matrix and the feasible action set.
[0144] In one embodiment, the pressure assessment module 803 is further configured to write the candidate action sets and joint tensors uploaded by each node agent into the federated critic; divide the joint tensor into regional blocks according to preset rules, and calculate the comprehensive pressure by combining the skylight compression degree, track congestion degree, and empty vehicle shortage degree; calculate the differential excitation of the relevant nodes in the pressure-overhead section, and inject the differential excitation into the node agent to determine the correction priority score of each candidate action set; reorder the candidate action sets according to the correction priority score to obtain the global action set; write the global action set into the federated critic for iterative calculation until the comprehensive pressure is less than the preset pressure threshold, then use the global action set as the corrected action set and write it into the scheduling release channel.
[0145] In one embodiment, the extreme case simulation module 804 is further configured to perform Hadamard perturbation analysis based on the target historical event template and random parameters to determine the extreme case tensor; write the extreme case tensor into the federated critic, and combine it with the node agent and risk penalty vector to perform risk calculation for each region block to determine the comprehensive risk score; combine the comprehensive risk score to conduct game-theoretic adversarial analysis to generate emergency actions; perform audit analysis on the emergency actions, and when the error compression rate meets the preset threshold compression rate, combine the trigger condition vector and the current emergency action to determine the emergency action set and write it into the scheduling and release channel.
[0146] In one embodiment, the convergence optimization module 805 is further configured to combine the extreme case tensor and the preset incentive weights of each node agent to determine the local action triplet; write the local action triplet and the extreme case tensor into the federated critic; update the node incentive weights using policy gradients; and determine the delay increment.
[0147] If the delay increment is greater than or equal to the preset increment threshold, the disturbance generator Jining pressure is invoked and returned to execute. Based on the target historical event template and random parameters, extreme scenario disturbance analysis and emergency game analysis are performed to determine the operation of the extreme case tensor and emergency action set until the delay increment is less than the preset increment threshold. Then, combined with the current emergency action set and the corrected action set, the preliminary train allocation strategy is determined.
[0148] In one embodiment, the plan generation module 806 is also used to send the preliminary train allocation strategy into the flexible formation compressor and locate the cross-regional meeting section; establish a time window sliding index for the meeting section, and sequentially perform sliding, single-time shrinkage insertion and rule verification to generate an executable plan and capacity margin parameters, write them into the scheduling release channel and push them to the field execution terminal.
[0149] In one embodiment, the real-time monitoring module 807 is further configured to compare the real-time tensor periodicity of train operation with the executable plan to determine the regional empty train rate, key port congestion index, and train serial waiting time; and, based on the regional empty train rate, key port congestion index, and train serial waiting time, lock the tensor snapshot when a preset trigger condition is met; and, based on the tensor snapshot, return to execute the operation of performing extreme scenario disturbance analysis and emergency game analysis based on the target historical event template and random parameters to determine the extreme case tensor and emergency action set, so as to generate a corrected executable plan and capacity margin parameters within a preset time.
[0150] Each module in the aforementioned train operation allocation device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.
[0151] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 9 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores train operation-related data in the railway system. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a train operation allocation method.
[0152] Those skilled in the art will understand that Figure 9 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0153] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:
[0154] Hourly streaming data is acquired, a joint tensor is constructed, and written into the scheduling shared memory unit. Node agents are segmented based on the joint tensor, and local action deduction is performed using these agents to determine a candidate action set. Candidate actions and the joint tensor are received synchronously, and differential incentives and overall pressure are calculated. The differential incentives are broadcast into the candidate action set, and the candidate action set is iteratively prioritized until the overall pressure is less than a preset pressure threshold, thus determining the corrected action set. Extreme scenario disturbance analysis and emergency game analysis are performed based on the target historical event template and random parameters to determine the extreme case tensor and emergency action set. Global weighting of the policy gradient is performed using the extreme case tensor, and cyclic convergence analysis is conducted using the corrected action set and emergency action set to determine the preliminary train allocation strategy. The preliminary train allocation strategy is sent to the elastic train compression unit to generate an executable plan and capacity reserve parameters, which are written into the scheduling release channel and pushed to the field execution end. The executable plan and capacity reserve parameters are monitored based on the real-time tensor. If the preset triggering conditions are met, return to execute the operation based on the target historical event template and random parameters to perform extreme scenario disturbance analysis and emergency game analysis, and determine the extreme case tensor and emergency action set.
[0155] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:
[0156] Hourly streaming data is acquired, a joint tensor is constructed, and written into the scheduling shared memory unit. Node agents are segmented based on the joint tensor, and local action deduction is performed using these agents to determine a candidate action set. Candidate actions and the joint tensor are received synchronously, and differential incentives and overall pressure are calculated. The differential incentives are broadcast into the candidate action set, and the candidate action set is iteratively prioritized until the overall pressure is less than a preset pressure threshold, thus determining the corrected action set. Extreme scenario disturbance analysis and emergency game analysis are performed based on the target historical event template and random parameters to determine the extreme case tensor and emergency action set. Global weighting of the policy gradient is performed using the extreme case tensor, and cyclic convergence analysis is conducted using the corrected action set and emergency action set to determine the preliminary train allocation strategy. The preliminary train allocation strategy is sent to the elastic train compression unit to generate an executable plan and capacity reserve parameters, which are written into the scheduling release channel and pushed to the field execution end. The executable plan and capacity reserve parameters are monitored based on the real-time tensor. If the preset triggering conditions are met, return to execute the operation based on the target historical event template and random parameters to perform extreme scenario disturbance analysis and emergency game analysis, and determine the extreme case tensor and emergency action set.
[0157] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, performs the following steps:
[0158] Hourly streaming data is acquired, a joint tensor is constructed, and written into the scheduling shared memory unit. Node agents are segmented based on the joint tensor, and local action deduction is performed using these agents to determine a candidate action set. Candidate actions and the joint tensor are received synchronously, and differential incentives and overall pressure are calculated. The differential incentives are broadcast into the candidate action set, and the candidate action set is iteratively prioritized until the overall pressure is less than a preset pressure threshold, thus determining the corrected action set. Extreme scenario disturbance analysis and emergency game analysis are performed based on the target historical event template and random parameters to determine the extreme case tensor and emergency action set. Global weighting of the policy gradient is performed using the extreme case tensor, and cyclic convergence analysis is conducted using the corrected action set and emergency action set to determine the preliminary train allocation strategy. The preliminary train allocation strategy is sent to the elastic train compression unit to generate an executable plan and capacity reserve parameters, which are written into the scheduling release channel and pushed to the field execution end. The executable plan and capacity reserve parameters are monitored based on the real-time tensor. If the preset triggering conditions are met, return to execute the operation based on the target historical event template and random parameters to perform extreme scenario disturbance analysis and emergency game analysis, and determine the extreme case tensor and emergency action set.
[0159] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0160] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0161] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A train operation allocation method, characterized in that, The method includes: Hourly streaming data is acquired, a joint tensor is constructed, and written into the scheduling shared memory unit; wherein, the hourly streaming data includes at least one of train position, track occupancy, port unloading quota, mine loading cycle time, locomotive cycle time, and temporary blockade. The node agents are obtained by segmenting the joint tensor, and local action inference is performed based on the node agents to determine the candidate action set. The candidate actions and the joint tensor are received synchronously, the differential excitation and the comprehensive pressure are calculated, the differential excitation is broadcast and injected into the candidate action set, the candidate action set is sorted by priority iteratively until the comprehensive pressure is less than a preset pressure threshold, and the corrected action set is determined. Based on the target historical event template and random parameters, extreme scenario perturbation analysis and emergency game analysis are performed to determine the extreme case tensor and emergency action set; The policy gradient is globally adjusted by combining the extreme case tensor, and a cyclic convergence analysis is performed by combining the corrected action set and the emergency action set to determine the preliminary train allocation strategy. The preliminary train allocation strategy is sent to the flexible formation compressor to generate an executable plan and capacity margin parameters, which are then written into the scheduling release channel and pushed to the field execution terminal. The executable plan and capacity margin parameters are monitored based on real-time tensors. When preset trigger conditions are met, the process returns to the step of performing extreme scenario disturbance analysis and emergency game analysis based on the target historical event template and random parameters to determine the extreme case tensor and emergency action set.
2. The train operation allocation method according to claim 1, characterized in that, The construction of the joint tensor includes: The original timestamps of the hourly stream data are retained and written to the buffer unit in the order of receipt; The data written to the cache unit is mapped to coordinates to determine the coordinate mapping relationship; Based on the coordinate mapping relationship, tensor units are filled sequentially according to the monotonic order of the time axis, spatial node axis, and resource axis to obtain a three-dimensional joint tensor.
3. The train operation allocation method according to claim 2, characterized in that, The step of obtaining node agents based on the joint tensor segmentation, and performing local action inference using the node agents to determine a candidate action set includes: The three-dimensional joint tensor is segmented along the spatial node axis to obtain different node slices. The node agent is generated by combining the mapping function and the node slices. The node slice represents all resource parameters of the same spatial node in the same time window. After concatenating the local node and two-hop neighborhood slices of the node agent, a local temporal vector is obtained through graph attention and gated loop processing. Based on the local time-series vector and preset constraints, priority deduction calculations are performed to obtain the action set; The feasibility of the action set is verified to determine the feasible action set; Based on the set of possible actions, conduct a cross-node impact assessment and analysis to determine the cross-node impact weight matrix; The candidate action set is determined by combining the cross-node influence weight matrix and the action set.
4. The train operation allocation method according to claim 1, characterized in that, The process involves synchronously receiving the candidate actions and the joint tensor, calculating the differential excitation and the overall pressure, broadcasting the differential excitation into the candidate action set, iteratively prioritizing the candidate action set until the overall pressure is less than a preset pressure threshold, and determining the corrected action set, including: Write the candidate action set uploaded by each node agent and the joint tensor into the federated critic; The joint tensor is divided into regional blocks according to preset rules, and the overall pressure is calculated by combining the sunroof compression degree, the track congestion degree, and the empty car shortage degree. Calculate the differential excitation of the relevant nodes in the overpressed section, and inject the differential excitation into the node agent to determine the correction priority score of each candidate action set; The candidate action set is reordered based on the corrected priority score to obtain the global action set; The global action set is written into the federated critic for iterative calculation until the overall pressure is less than a preset pressure threshold. Then, the global action set is used as the corrected action set and written into the scheduling and release channel.
5. The train operation allocation method according to claim 1, characterized in that, The process involves performing extreme scenario perturbation analysis and emergency game analysis based on the target historical event template and random parameters to determine the extreme case tensor and emergency action set, including: Hadamard perturbation analysis is performed based on the target historical event template and random parameters to determine the tensor for extreme cases; The extreme case tensor is written into the federated critic, and combined with the node agent and risk penalty vector, risk calculation is performed on each region block to determine the comprehensive risk score; The comprehensive risk score is used to conduct game-based confrontation and generate emergency actions; The emergency actions are audited and analyzed. When the error compression rate meets the preset threshold compression rate, the emergency action set is determined and written into the scheduling and release channel by combining the trigger condition vector and the current emergency action.
6. The train operation allocation method according to claim 1, characterized in that, The process of globally adjusting the policy gradient by combining the extreme case tensor, and performing cyclic convergence analysis by combining the corrected action set and the emergency action set to determine the preliminary train allocation strategy includes: By combining the extreme case tensor and the preset incentive weights of each node agent, local action triples are determined; Write the local action triplet and the extreme case tensor into the federated critic, update the node incentive weights using policy gradient, and determine the delay increment. If the delay increment is greater than or equal to a preset increment threshold, the disturbance generator Jining is invoked and the process returns to execute the steps of performing extreme scenario disturbance analysis and emergency game analysis based on the target historical event template and random parameters to determine the extreme case tensor and emergency action set, until the delay increment is less than the preset increment threshold. Then, the preliminary train allocation strategy is determined by combining the current emergency action set and the corrected action set.
7. The train operation allocation method according to claim 1, characterized in that, The step of sending the preliminary train allocation strategy into the flexible formation compressor to generate an executable plan and capacity margin parameters, writing it into the scheduling release channel, and pushing it to the field execution terminal includes: The preliminary train allocation strategy is fed into the flexible train formation compressor, and the cross-regional meeting section is located. A time window sliding index is established for the intersection section. Sliding, single-time shrinking insertion and rule verification are performed in sequence to generate an executable plan and capacity margin parameters, which are written into the scheduling release channel and pushed to the field execution terminal.
8. The train operation allocation method according to claim 7, characterized in that, The method further includes: The real-time tensor periodicity during train operation is compared with the executable plan to determine the regional empty train rate, the congestion index of key ports, and the train serial waiting time. If the preset triggering conditions are met based on the regional empty car rate, the key port congestion index, and the train serial waiting time, lock the tensor snapshot; Based on the tensor snapshot, return to the step of performing extreme scenario perturbation analysis and emergency game analysis based on the target historical event template and random parameters to determine the extreme case tensor and emergency action set, so as to generate a corrected executable plan and capacity margin parameters within a preset time.
9. A train operation distribution device, characterized in that, The device includes: The tensor construction module is used to acquire hourly stream data, construct a joint tensor, and write it into the scheduling shared memory unit; wherein, the hourly stream data includes at least one of train position, track occupancy, port unloading quota, mine loading cycle time, locomotive cycle time, and temporary blockade. The local decision-making module is used to obtain node agents based on the joint tensor segmentation, and to perform local action inference based on the node agents to determine the candidate action set. The stress assessment module is used to synchronously receive the candidate actions and the joint tensor, calculate the differential excitation and the overall stress, broadcast the differential excitation into the candidate action set, perform priority iterative sorting on the candidate action set until the overall stress is less than a preset stress threshold, and determine the corrected action set. The extreme scenario simulation module is used to perform extreme scenario perturbation analysis and emergency game analysis based on the target historical event template and random parameters, and to determine the extreme case tensor and emergency action set. The convergence optimization module is used to perform global weighting of the policy gradient by combining the extreme case tensor, and to perform cyclic convergence analysis by combining the corrected action set and the emergency action set to determine the preliminary train allocation strategy. The plan generation module is used to send the preliminary train allocation strategy into the flexible formation compressor, generate an executable plan and capacity margin parameters, write them into the scheduling release channel and push them to the field execution terminal; The real-time monitoring module is used to monitor the executable plan and capacity margin parameters based on real-time tensors, and when preset trigger conditions are met, it returns to control the extreme simulation module to perform the operation of performing extreme scenario disturbance analysis and emergency game analysis based on the target historical event template and random parameters to determine the extreme case tensor and emergency action set.
10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 8.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.
12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.