Asynchronous control method and system for heavy haul train
Through layered multi-agent decision-making model and reinforced learning technology, the vertical impulse problem of heavy-loaded trains in complex environments is solved, achieving safe and stable operation and intelligent improvement.
Patent Information
- Application Number
- CN202510230907.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-20
AI Technical Summary
The existing asynchronous control technology of heavy-load trains cannot effectively solve the problem of longitudinal impulse caused by terrain differences in the growing group, and lacks fine-grained control algorithms and comprehensive experimental verification, which cannot meet the needs of complex operating environments.
The layered multi-agent decision-making model is adopted to obtain train operation information in real time, build a reinforcement learning model, and provide comprehensive decision-making support for trains from top to bottom, and dynamically adjust driving decisions based on the actual environment and status.
The safe and smooth operation of heavy-load trains has been achieved, the level of intelligence has been improved, the ability to operate independently has been improved, vertical impulse has been reduced, operating costs have been reduced, and the service life of key equipment has been extended.
Smart Images

Figure CN120178720A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of rail vehicles, and particularly relates to a heavy-haul train asynchronous control method and system. Background Art
[0002] Heavy-haul trains are widely used in freight transportation. They are characterized by long formation lengths and large load capacities, which can significantly improve transportation efficiency and reduce costs. Currently, heavy-haul trains mainly adopt two braking strategies: traditional air braking and the electronically controlled pneumatic braking system (ECP).
[0003] Among them, traditional air braking transmits braking commands by using compressed air, and each vehicle brakes in sequence. This braking method is simple, but has poor synchronization, and the braking time interval between the front and rear vehicles is long, resulting in large longitudinal impulses and low braking efficiency. ECP technology transmits braking commands through electrical signals, achieving synchronous braking and release of the entire train, shortening the coasting time and braking distance, and reducing longitudinal impulses. However, ECP technology still has limitations in long-formation heavy-haul trains, and cannot effectively solve the longitudinal impulse problem caused by terrain differences, and lacks asynchronous control algorithm support.
[0004] In recent years, some research-based driving strategies in the prior art have attempted to optimize train control through methods such as reinforcement learning. For example, the DQN method is used to train synchronous strategies or neural network predictive control curves. However, these methods are currently mostly limited to simple line environments, lack adaptability to the operating conditions of long trains, and do not optimize the coupler force. In addition, existing research mostly focuses on locomotive control, lacking refined control of each vehicle in the entire train, especially the realization of asynchronous braking and coupler force optimization. Therefore, although the electronically controlled pneumatic braking system (ECP) provides a technical basis for the asynchronous control of heavy-haul trains, due to the lack of fine-grained control algorithms and comprehensive experimental verification, the prior art still cannot meet the requirements of the complex operating environment of heavy-haul trains.
[0005] Therefore, there is an urgent need for a heavy-haul train asynchronous control method that can achieve train-level autonomous driving, refined control of each vehicle, and optimization of coupler forces. Summary of the Invention
[0006] The purpose of the present invention is to solve one of the above technical problems, and provide a heavy-haul train asynchronous control method and system. By constructing a hierarchical multi-agent decision-making model, it can provide comprehensive decision-making support for the train from top to bottom based on real-time operating environment information, and can dynamically adjust the driving decision according to the actual operating environment and state of the train, ensuring the safe and stable operation of the heavy-haul train.
[0007] To achieve the above purpose, the technical solution adopted by the present invention is:
[0008] An asynchronous control method for heavy-haul trains, comprising the following steps:
[0009] Obtain the real-time operation information of each vehicle in the heavy-haul train; the real-time operation information includes the speed information, position information, gradient information, and environmental information of each train;
[0010] Construct a hierarchical multi-agent decision-making model, and establish a reinforcement learning model for each agent in the decision-making model; the multi-agent decision-making model includes a decision-making layer and an execution layer;
[0011] The decision-making layer contains multiple levels; among them, the first level has one agent; except for the first level, the number of agents in each level is determined by the number of agents in the upper level and the decision-making number of each agent in the upper level; the agent in the first level receives the real-time operation information of the vehicle and outputs the decision for the agents in the second level; except for the agents in the first level, each agent receives the decision output by the agent in the upper level and refines and splits it into the decision input for the lower level;
[0012] The execution layer includes the on-vehicle controllers of each vehicle in the heavy-haul train. The execution layer receives the decision output by the agent in the last level of the decision-making layer and translates it into an actuator instruction to control the train to run;
[0013] Iteratively train the reinforcement learning models of the agents in the hierarchical multi-agent decision-making model based on the real-time operation information using the proximal policy optimization algorithm;
[0014] Deploy the trained hierarchical multi-agent decision-making model to obtain the control decisions of each vehicle in the heavy-haul train and act on the traction and braking systems of the heavy-haul train.
[0015] In some embodiments of the present invention, the following steps are further included:
[0016] When the number of decisions given by the agent in the upper level cannot fully match the number of agents in the lower level, resampling is performed using the interpolation method to match the number of agents in the lower level.
[0017] In some embodiments of the present invention, each agent is preset with a built-in basic decision. The basic decision takes the current speed, maximum speed limit, and gradient of the vehicle as inputs and outputs the decision with the required granularity;
[0018] For each agent, when the upper-level decision instruction exists, execute the upper-level decision instruction. When the upper-level decision instruction does not exist, use the built-in basic decision as the upper-level decision instruction.
[0019] In some embodiments of the present invention, the method for establishing a reinforcement learning model for each agent in the decision-making model includes the following steps:
[0020] Define the decision space of each agent wherein, is the decision of each agent, k1 is the predetermined maximum braking threshold, k2 is the predetermined maximum traction threshold, and N n is the number of decisions made by the current agent to control the agents at the next level;
[0021] Define the state space of each agent which includes the positions s of all vehicles s , speeds s v , gradients s trac , image preprocessing features s img , decisions at the previous time step and the vehicle range s controlled by the current agent mask ;
[0022] Design the reward for each agent where μ k is the weight of different rewards; includes longitudinal impact reward speed reward upper-level instruction reward and safety limit reward
[0023] The calculation formula for the longitudinal impact reward is:
[0024]
[0025] The calculation formula for the speed reward is:
[0026]
[0027] The calculation formula for the upper-level instruction reward is:
[0028]
[0029] The calculation formula for the safety limit reward is:
[0030]
[0031] where, is the maximum coupler force among the vehicles controlled by the agent , and and are the coupler forces before and after the vehicle group controlled by the agent respectively, is the maximum coupler force received by each vehicle, and v limit are the average speed and the maximum limit speed of the vehicle group controlled respectively, is the decision made by the superior agent ; For the intelligent agent The decision response made Is the maximum safety limit of the preset coupler force.
[0032] In some embodiments of the present invention, the reinforcement learning model of each intelligent agent adopts the Actor-Critic architecture, including the behavior network Act θ (s) and the critic network V θ (s). The method for establishing the reinforcement learning model for each intelligent agent in the decision-making model further includes the following steps:
[0033] Define the policy of each intelligent agent Wherein, Is the function that maps the environmental observation result To the execution decision Of the function, Is the neural network parameter of the superior intelligent agent Of.
[0034] In some embodiments of the present invention, it includes the following steps:
[0035] During the iterative training of the reinforcement learning model of each intelligent agent, a stable loss function based on the proximal policy optimization algorithm is used to alleviate the training oscillation of the policy of each intelligent agent;
[0036] The expression of the stable loss function based on the proximal policy optimization algorithm is:
[0037]
[0038] Wherein, E t Is the expectation of the calculation function for each time step t, π θ (a t |s t ) is the probability that the policy generates the action a t Under the parameters θ and the environment s t , given by the behavior network Act θ (s), Is the probability that the policy generates the action a old Under the parameters θ t And the environment s t , given by the behavior network Act θ (s), Is the advantage function value at time step t, Is the importance weight, representing the ratio of the new policy to the old policy, and ∈ is the clipping range, used to limit the amplitude of the policy update.
[0039] In some embodiments of the present invention, it includes the following steps:
[0040] Establish the judgment criteria for similar agents, determine the agents that meet the judgment criteria as similar agents, and share network parameters among similar agents during the iterative training of the reinforcement learning models of each agent.
[0041] In some embodiments of the present invention, the method for iteratively training the reinforcement learning models of each agent in the hierarchical multi-agent decision-making model using the proximal policy optimization algorithm includes the following steps:
[0042] Adopt a network structure composed of an encoder and a decoder. Use the encoder to map the values in the observation space to a hidden encoding, and use the decoder to map the hidden encoding to a decision to construct the action network Act θ (s);
[0043] Adopt a temporal feature extraction layer combined with a fully connected layer to construct the critic network V θ (s);
[0044] Construct the action network Act θ (s) and the critic network V θ (s);
[0045] Set the optimization objective of the action network Act θ (s):
[0046] L Act = L CLIP (θ) + KL(π θ );
[0047] Among them, L CLIP (θ) is a stable loss function based on the proximal policy optimization algorithm, and KL(π θ ) is the KL divergence of the current action distribution, which is used to measure the decision diversity of the decision-making;
[0048] Set the optimization objective of the critic network
[0049] Among them, is the advantage function value calculated using the generalized advantage estimation method at time step t, The calculation formula of
[0050] Among them
[0051] Among them, γ ∈ [0.9, 1] is the discount factor, which is used to consider the discount of future rewards when calculating the cumulative reward, and λ ∈ [0, 1] is the hyperparameter of the generalized advantage estimation method, which is used to control the trade-off between bias and variance, is the temporal difference error at time step t, r t is the reward obtained at time step t, V(s t ) and V(st+1 ) are the state value estimates for states s t and s t+1 respectively;
[0052] Set the total training objective to be minimized as L=-L Act +L cri , and iteratively train the model based on the minimized total training objective.
[0053] In some embodiments of the present invention, the method for iteratively training the model based on the minimized total training objective includes the following steps:
[0054] Each agent interacts with the interaction object, makes a decision based on the observation, and obtains a reward;
[0055] After each agent collects a predetermined amount of data, it packs the decision, observation, and reward and sends them to the proximal policy optimization algorithm, and trains the policy, Act θ (s) and V θ (s) based on the algorithm process until the proximal policy optimization algorithm converges.
[0056] Some embodiments of the present invention further provide a heavy-haul train asynchronous control system for implementing the above heavy-haul train asynchronous control method, including:
[0057] The locomotive end includes a decision-making module, and the decision-making module includes a hierarchical multi-agent decision-making model for giving real-time train operation decisions based on the real-time train operation information;
[0058] The vehicle end includes a measurement module, a communication module, and a power module; the measurement module is used to obtain the real-time train operation information; the communication module is used to transmit the real-time train operation information and decision-making information between the vehicles of the train; the power module is used to output corresponding traction force and / or braking force based on the decision of the decision-making module;
[0059] The cloud end includes a training module, and the training module is used to pre-train the deep reinforcement learning models of each agent based on the initial data, continuously update the hierarchical multi-agent decision-making model based on the supplementary data uploaded by the train, and perform self-training optimization on the hierarchical multi-agent decision-making model.
[0060] The beneficial effects of the present invention are as follows:
[0061] 1. The hierarchical multi-agent decision-making model constructed by the present invention can provide comprehensive decision support for the train from top to bottom based on the real-time operation environment information, realizes fine operation decisions at each vehicle level, micro-optimizes the forces on the couplers, ensures the safe and stable operation of the heavy-haul train, and significantly improves the intelligent level of train operation;
[0062] 2. The hierarchical multi-agent decision-making model constructed by the present invention enables the train to autonomously perceive and adapt to the operating environment, thereby making scientific and reasonable decisions without excessive manual intervention, greatly improving the train's autonomous operation ability. Moreover, since the driving decision can be dynamically adjusted according to the actual operating environment and state of the train, the train can operate stably in a complex and changeable environment.
[0063] 3. Compared with the traditional driver operation mode, the present invention ensures that the train operates at a better speed and strategy on the premise of ensuring safety through precise algorithm control, thereby effectively reducing the correction time caused by repeated manual operations and significantly improving the train's operation efficiency.
[0064] 4. The present invention allows the algorithm model to continuously learn and update, which not only ensures that the system can continuously adapt to new operating environments and requirements, but also realizes the maintenance and repair throughout the life cycle, greatly reducing the operating cost.
[0065] 5. The hierarchical multi-agent decision-making model constructed by the present invention effectively optimizes the driving strategy, which helps to extend the sustainable service life of key equipment such as couplers. Moreover, the design of a penalty term is introduced into the reward function, which can effectively prevent the train from deviating from the desired state or touching the safety boundary during operation, ensuring the safety of train operation while guaranteeing the operation efficiency.
[0066] 6. When facing emergencies, the hierarchical multi-agent decision-making model constructed by the present invention can quickly collect comprehensive data and make comprehensive judgments and responses, and timely adjust the operation strategy, further enhancing the stability and safety of the train in different operating environments.
[0067] Other features and advantages of the present invention will be described in the following specification, and some of them will become obvious from the specification or be understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures pointed out in the specification, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will describe the specific embodiments of the present invention in detail with reference to the drawings. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0069] Figure 1 It is a flowchart of a heavy-haul train asynchronous control method;
[0070] Figure 2 It is a schematic structural diagram of the hierarchical multi-agent decision-making model provided by the embodiment of the present invention;
[0071] Figure 3 Schematic diagram of the training process of the hierarchical multi-agent decision-making model provided by the embodiment of the present invention;
[0072] Figure 4 System architecture diagram of a heavy-haul train asynchronous control system. Detailed implementation manners
[0073] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be described and explained below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments provided in the present application without creative efforts shall fall within the scope of protection of the present application.
[0074] It should be noted that the terms used herein are only for describing the specific implementation manners and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless otherwise clearly specified in the context, the singular form is also intended to include the plural form. In addition, it should be understood that the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0075] Without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0076] The technical solutions of the present invention will be described in detail below with reference to specific embodiments and the accompanying drawings of the specification.
[0077] As shown in the attached Figure 1 - attached Figure 3 In a schematic embodiment of a heavy-haul train asynchronous control method of the present invention, the control method includes the following steps.
[0078] S1: Obtain the real-time operation information of each vehicle and the surrounding environment of the heavy-haul train at each sampling time. Among them, the real-time operation information includes the speed information, position information, slope information of each train, and the environmental information such as weather and obstacles transmitted through the external image recognition interface.
[0079] Preprocess the obtained real-time operation information and normalize it into a feature vector that follows a normal distribution.
[0080] It should be noted that in this embodiment, the environmental information input through the image recognition interface is used as the input of the decision-making model, which facilitates adjusting the decision-making estimation of the model according to the image during the training process, and further allows dynamic adjustment of the decision-making according to external factors. For example, when it is raining or snowing, the speed is reduced and braking is advanced in advance, and emergency braking is performed when an obstacle appears, etc.
[0081] S2: Construct a hierarchical multi-agent decision-making model, and establish a reinforcement learning model for each agent in the decision-making model to provide driving decisions for the heavy-haul train.
[0082] As shown in the appendix Figure 2 As shown, the multi-agent decision-making model includes a decision-making layer and an execution layer.
[0083] Among them, the decision-making layer I contains n levels, and its split structure is represented as I = {I1, I2, I3,... I n}.
[0084] Among them, the first level I1 has an agent A 1 . The input of the agent at the first level is the feature vector of the real-time running information of the vehicle after preprocessing, and the output is the decision for the agent at the second level.
[0085] Except for the first level, the number of agents in each level is determined by the number of agents in the previous level and the number of decisions of each agent in the previous level. That is, starting from the second level I2, each level has I n = I n-1 * N n-1 agents Among them, I n-1 is the number of agents in the previous level, and N n-1 represents the number of decisions of each agent in the previous level; i1, i2,... i n represents the subordinate relationship of the agent.
[0086] For the agent It is controlled by the agent at the first layer The agent at the second layer …, the agent at the n-1th layer Control, and it is the i n th agent under this control path, so it is named
[0087] Except for the agents at the first level, each agent receives the decisions output by the agents at the previous level, refines and splits them into the decision inputs of multiple agents at the next level. For the agents at the last level, they directly output the control decisions for each vehicle in the execution layer.
[0088] Specifically, for each non-top-level agent The input is the feature vector obtained by preprocessing the real-time operation information of the vehicle, which is the same as that of the top-level agent, and its corresponding superior agent The output is the vehicle control decision.
[0089] The execution layer includes the on-vehicle controllers of each vehicle in the heavy-haul train. The execution layer receives the decisions output by the agents of the last level in the decision layer and translates them into actuator instructions to control the train operation.
[0090] It should be noted that the number of agents in the last level is the same as the number of on-vehicle controllers of each vehicle in the execution layer.
[0091] For the splitting of each level of the decision layer I, the following principles are followed:
[0092] 1. As much as possible, each agent should control vehicles with complete train control functions, that is, with both traction and braking functions.
[0093] 2. As much as possible, the number of trains controlled by each agent should be similar.
[0094] 3. According to the computing power of the training / deployment hardware, determine the maximum number of decisions max{N n} of each agent.
[0095] As shown in the appendix Figure 2 Taking the 303-car heavy-haul train formation of (1 locomotive + 100 gondolas) * 3 as an example, in this embodiment, the decision layer has 3 levels, with 1, 3, and 6 agents respectively at each level, and N n are 3, 2, and 2 respectively. The execution layer is split into 12 car groups, and the specific splitting method is as follows:
[0096] For I1, which is the top-level decision and does not need to be split, it only has 1 agent.
[0097] For the splitting of I2, according to Principle 1, the locomotive needs to be included in the vehicle control of each agent to meet the traction function; according to Principle 2, the number of vehicles controlled by each agent is the same. Therefore, I2 = 3, and each agent controls a 1 + 100 formation.
[0098] For I3, when Principle 1 cannot be satisfied, according to Principle 2 and Principle 3, therefore I3 = 6, and each I3 agent is split into 2 agents, which control 1 locomotive + 50 gondolas and 50 gondolas in 2 cases, so that the entire I3 includes 6 agents.
[0099] Further, for the execution layer under the jurisdiction of I3, assuming that the computing power allows, each I3 agent is split into 2 execution layers, respectively controlling 1 locomotive + 25 gondolas or 25 gondolas.
[0100] In some embodiments of the present invention, each level of agent can be taken over in other ways instead of self-training. For example, agent A 1 can be taken over by the driver, who directly outputs instructions to the secondary agent. At this time, the algorithm automatically becomes an assisted driving algorithm, which minimizes the coupler force on the basis of satisfying the driver's decision.
[0101] Some embodiments of the present invention further include the following steps:
[0102] In cases where the agent is taken over by the driver, etc., when the number of decisions given by the upper-level agent cannot fully match the number of lower-level agents, an interpolation method is used for resampling to match the number of lower-level agents. For example, the B-spline interpolation method.
[0103] In some embodiments of the present invention, each agent is preset with a built-in basic decision, and the input of each level of agent is the decision of the upper-level agent or the built-in basic decision.
[0104] The built-in basic decision takes the current vehicle speed, the maximum speed limit, and the slope as inputs and outputs decisions at the required granularity.
[0105] For each agent, when the upper-level decision instruction exists, execute the upper-level decision instruction; when the upper-level decision instruction does not exist, use the built-in basic decision as the upper-level decision instruction.
[0106] In some embodiments of the present invention, each agent in the decision model can be replaced by an external interface. By allowing any degree of external intervention, from manual control to large model embedding, it can serve as a good base for external decisions and improve the performance of external decisions.
[0107] In some embodiments of the present invention, the built-in basic decision is obtained based on a mechanism-based model.
[0108] In some embodiments of the present invention, the method for establishing a reinforcement learning model for each agent in the decision model includes the following steps:
[0109] Define the decision space of each agent
[0110] where is the decision for each agent, k1 is a predetermined maximum braking threshold, k1 < 0, k2 is a predetermined maximum traction threshold, k2 > 0, N n$a_n$ is the number of decisions of the $n$-th layer agent, that is, the number of decisions used by the current agent to control the agents at the next level, which is the level command for the agents at the next level output by each agent. When $a = 0$, neither traction nor braking is applied, and it is in the coasting state at this time.
[0111] In this embodiment, a continuous decision space is adopted, that is, the decision range of the decision space includes continuous floating-point numbers from a predetermined maximum braking threshold to a predetermined maximum traction threshold.
[0112] Among them, the predetermined maximum braking threshold is set to -100, and the maximum traction threshold is set to 100. That is, the decision range is continuous floating-point numbers from -100 (maximum braking) to 100 (maximum traction).
[0113] It should be noted that when the agent is at the last level, the decision space directly represents the level range of the locomotive or freight car. For freight cars, since freight cars actually do not include a traction module, the levels of freight cars greater than 0 will be discarded.
[0114] In some embodiments of the present invention, except for the decision space of the last level, because it is necessary to directly command the braking and traction module and the form of the decision space needs to be fixed, the decision space of the upper-layer agents can have a more flexible form.
[0115] For example, when the communication bandwidth and computing power are sufficient, in order to give redundancy for accidental communication interruption, the decision space can:
[0116] Give the decisions for multiple subsequent time steps at one time, and at this time the decision space becomes where $m$ is the number of multiple time steps.
[0117] Give the speed curve to be followed, and at this time the decision space becomes where $m$ is the time step length of the given speed curve, and $v$ limit is the maximum speed limit of the current section.
[0118] When the communication bandwidth and computing power are limited, the decision space can be simplified to give the acceleration suggestion for the lower-layer agents, and at this time the decision space becomes That is, the suggestion for increasing or decreasing the level in the next time step.
[0119] Define the state space of each agent That is, the space describing the current state of the train in the environment simulation, which is the data that the agent can actually observe and serves as the basis for control decisions.
[0120] Among them, it includes the positions $s$ of all vehicles s , speeds $\dot{s}$ v , gradients $s$ trac , image preprocessing features $s$img The decisions of previous time steps and the vehicle range s controlled by the current agent mask . The vehicle range s controlled by the current agent mask is used to assist the agent in making decisions.
[0121] In some embodiments of the present invention, in order to ensure the robustness of poor communication states, the state space can be local information
[0122] For the agent directly controlling the vehicle the state space includes the position of the controlled vehicle speed gradient the decisions of previous time steps
[0123] For the agent indirectly controlling the vehicle
[0124] In some embodiments of the present invention, when communication conditions permit, the state space includes the state spaces of adjacent agents, that is
[0125] Design the reward for each agent which is the only feedback received by the agent for each operation and is used to evaluate the quality of the decision.
[0126] where μ k is the weight of different rewards, including the longitudinal impact reward speed reward upper layer instruction reward and safety limit reward
[0127] where, for each agent and the vehicle group it controls, the calculation formula for the longitudinal impact reward is:[[]]
[0128]
[0129] The calculation formula for the speed reward is:[[]]
[0130]
[0131] The calculation formula for the upper layer instruction reward is:[[]]
[0132]
[0133] The calculation formula for the safety limit reward is:[[]]
[0134]
[0135] Among them, is the maximum coupler force in the vehicle controlled by the agent and and are respectively the coupler forces before and after the vehicle group controlled by the agent and is the maximum coupler force received by each vehicle and v limit are respectively the average speed and the maximum limit speed of the vehicle group to be controlled is the decision made by the superior agent and is the decision response made by the agent and is the preset maximum safety limit of the coupler force.
[0136] In the analysis of the coupler force of heavy-haul trains, one of the main reasons for the increase in coupler force is that in small formation trains, improper operation causes the coupler force to continuously expand during transmission between vehicles. However, due to the insufficient length of small formation trains, the coupler force terminates before it reaches a destructive level. However, for 20,000-ton or even 30,000-ton heavy-haul trains, due to the increase in train scale, the gradually expanding coupler force caused by improper operation may exceed the safety limit. Therefore, in order to limit the gradual expansion of the coupler force between vehicles, the present invention adds an additional difference in coupler force between the front and rear couplers of the vehicle to the coupler force reward, which is used to prevent positive feedback on the coupler force in the decision-making process, thereby allowing for the expansion of multiple formations.
[0137] In some embodiments of the present invention, the reinforcement learning model of each agent adopts an Actor-Critic architecture, including two neural networks, namely the behavior network Act θ (s) and the critic network V θ (s). Among them, the behavior network is used to output a decision a according to the environmental observation results during training and prediction, and the critic network is used to evaluate the model decision and estimate the reward during training to assist in training the behavior network.
[0138] The method for establishing a reinforcement learning model for each agent in the decision-making model further includes the following steps:
[0139] Define the policy of each agent Among them, is a function that maps the environmental observation result to the execution decision and is the parameter of the neural network of the superior agent .
[0140] The hierarchical multi-agent decision-making model constructed by the present invention realizes independent and continuous control of hundreds of vehicles (including locomotives and freight cars) through hierarchical asynchronous execution of ultra-large action spaces, realizes flexible control of heavy-load train driving decisions, and maximizes the reduction of train longitudinal impulses. At the same time, it reduces the complexity of each agent, effectively reducing the action space that each agent needs to process from hundreds of dimensions to several dimensions, greatly improving training efficiency, effectively alleviating training divergence problems, and improving control accuracy.
[0141] S3: As attached Figure 3 As shown in the figure, the reinforcement learning model of each agent in the hierarchical multi-agent decision-making model is iteratively trained using the proximal policy optimization algorithm (PPO) based on real-time operation information.
[0142] In some embodiments of the present invention, in the asynchronous control scenario of heavy-loaded trains, a slight change in the control strategy may also lead to a drastic change in the coupler force, which in turn leads to a drastic change in the loss function, causing the control strategy to oscillate. Therefore, in step S3, during the iterative training of the reinforcement learning model of each intelligent agent, a stable loss function based on the proximal policy optimization algorithm is used to alleviate the training oscillation of each intelligent agent's strategy.
[0143] Among them, the expression of the stable loss function based on the proximal strategy optimization algorithm is:
[0144]
[0145] Among them, E t To calculate the expectation of the function for each time step t, π θ (a t |s t ) is the parameter θ and the environment s t Next, the strategy generates action a t The probability of θ (s) given, For the parameter θold and the environment s t Next, the strategy generates action a t The probability of θ (s) given, is the advantage function value at time step t, which measures the advantage of taking a certain action relative to the average level, is the importance weight, which indicates the ratio of the new strategy to the old strategy, and ∈ is the clipping range, which is used to limit the amplitude of the strategy update.
[0146] By using the loss function clipping of the proximal policy optimization algorithm, the training oscillation of the control strategy can be effectively alleviated, the stability of the strategy can be ensured, and the possibility of severe oscillation can be reduced.
[0147] In some embodiments of the present invention, in order to improve the training efficiency, the following steps are further included:
[0148] Establish a judgment criterion for similar agents, determine the agents that meet the judgment criterion as similar agents, and share network parameters among similar agents during the process of iteratively training the reinforcement learning models of each agent.
[0149] For example, taking whether it has a traction function as the judgment criterion, share network parameters for agents without locomotives, and share another set of network parameters for agents with locomotives.
[0150] In some embodiments of the present invention, in order to improve the training efficiency, the policy is allowed to be initialized as an artificial prior policy, including but not limited to driver policies, mechanism model policies, etc.
[0151] In some embodiments of the present invention, the method for iteratively training the reinforcement learning models of each agent in the hierarchical multi-agent decision-making model using the proximal policy optimization algorithm includes the following steps:
[0152] S31: Establish a network structure of behavior-critic, and the specific method includes:
[0153] Adopt a network structure composed of an encoder and a decoder. Use the encoder to map the values in the observation space to a hidden encoding, and use the decoder to map the hidden encoding to a decision, and construct the behavior network Act θ (s).
[0154] Adopt a time series feature extraction layer combined with a fully connected layer to construct the critic network V θ (s).
[0155] In some embodiments of the present invention, in order to achieve a balance between performance and computational efficiency, the encoder can use one-dimensional convolution, and the decoder can use a fully connected layer.
[0156] In some embodiments of the present invention, in order to effectively extract time series features, the critic network can adopt models such as LSTM, Transformer, SSM, etc. as the basic network architecture.
[0157] It should be noted that due to the limitation of the transmission speed of force on the coupler, the coupler force at each time step depends not only on the state of the previous time step, but also on the state of the train on a longer time scale. In order to achieve effective continuous estimation and real-time correction of the coupler force and realize effective evaluation of train operation, a time series neural network architecture that can effectively analyze long-term dependencies is adopted, including but not limited to models such as RNN, LSTM, Transformer, SSM, etc., as the critic model V for evaluating the reward including the coupler force.
[0158] S312: Set the behavior network Act θ (s) Optimization objective:
[0159] L Act = L CLIP (θ) + KL(π θ );
[0160] Where L CLIP (θ) is the stable loss function based on the proximal policy optimization algorithm, and KL(π θ ) is the KL divergence of the current action distribution, which is used to measure the decision diversity of the decision-making;
[0161] Set the optimization objective of the critic network
[0162] Where is the advantage function value calculated using the generalized advantage estimation method at time step t. In the hierarchical reinforcement learning framework, the generalized advantage estimation (GAE) balances the long-term and short-term rewards in the critic network by fusing the Monte Carlo (MC) and temporal difference (TD) methods, enabling the decisions output by the agent to adapt to the needs of agents at different levels.
[0163] The calculation formula of is:
[0164] Where
[0165] Where γ ∈ [0.9, 1] is the discount factor, which is used to consider the discount of future rewards when calculating the cumulative reward; λ ∈ [0, 1] is the hyperparameter of the generalized advantage estimation method, which is used to control the trade-off between bias and variance; is the temporal difference error at time step t; r t is the reward obtained at time step t, and V(s t ) and V(s t+1 ) are the state value estimates of states s t and s t+1 respectively.
[0166] Set the minimization of the total training objective L = -L Act + L cri , and iteratively train the model based on the minimization of the total training objective.
[0167] It should be noted that by adjusting the smoothing parameter λ, GAE can flexibly control the time span: it degenerates into MC (low bias, high variance) when λ → 1, and approximates TD (high bias, low variance) when λ → 0. In hierarchical decision-making, the top-level agent sets λ top→1 to capture long-term cumulative effects (such as whether the current operation will cause a sudden increase in the coupler force at a certain moment within the future short-term attention time period, and the comprehensive speed control within the time period), while the lower-layer agent adopts λ low →0, only focusing on short-term actions (such as the execution of the braking instruction in the next time step), thus saving computing resources.
[0168] For example, in the control of heavy-haul trains, the upper layer calculates the long-term advantages for the next 50 - 100 steps based on the complete trajectory, while the lower layer only needs the short-term advantages for the next 3 - 5 steps. This hierarchical design not only avoids cross-layer interference (such as the upper-layer agent setting = 0.99 and the lower-layer agent setting = 0.5), but also dynamically adjusts λ to adapt to different task requirements (such as temporarily increasing the lower-layer λ on steep slopes), ultimately achieving a balance between bias and variance and improving the training stability and decision-making efficiency of the AC architecture.
[0169] It should also be noted that when calculating the advantage function, both long-term rewards and short-term rewards need to be considered. The long-term reward means that the agent needs to consider the long-term impact of the decision, that is, the impact of the current decision on the rewards after multiple time steps, such as whether the current operation will cause a sudden increase in the coupler force after a long time. The short-term reward means that the agent needs to consider the short-term impact of the decision, such as how to comply with the instructions of the upper-level agent at the current time step and reduce the coupler force. In the asynchronous control of heavy-haul trains with hierarchical reinforcement learning, the decision-making focuses of different levels are different. For example, the higher-level agents need to consider the impact of long-term decisions more, while the lower-level agents do the opposite. Using the generalized advantage estimation method to calculate the advantage function value can facilitate the flexible adjustment of the long-term and short-term reward coefficients, enabling the decisions output by the agent to adapt to the needs of agents at different levels.
[0170] In some embodiments of the present invention, the method for iteratively training the model based on minimizing the total training objective specifically includes the following steps:
[0171] Each agent interacts with the interaction object, makes a decision based on the observation, and obtains a reward. Among them, the interaction object is the environment for real-time interaction or the cached experience saved in advance.
[0172] After each agent collects a predetermined amount of data, it packs the decision, observation, and reward and sends them to the proximal policy optimization algorithm, and trains the policy, Act θ (s) and V θ (s) based on the algorithm process to maximize the expected reward.
[0173] After one step of training, update the policy to the host and synchronize it to the agent, and repeat this cycle until the proximal policy optimization algorithm converges.
[0174] This training method allows for initial training from a pre-set basic decision, thereby reducing the time required for the algorithm to explore freely and improving the training efficiency.
[0175] S4: Deploy the trained hierarchical multi-agent decision-making model to the locomotive side to obtain the control decisions of the vehicles in the heavy-haul train and apply them to the traction and braking systems of the heavy-haul train.
[0176] In some embodiments of the present invention, the decision-making algorithm and the base model can be flexibly switched, allowing any degree of external intervention from training to deployment. By replacing any agent with external input from reinforcement learning pre-training, various upgrades or downgrades from manual control to large model embedding are compatible.
[0177] Some embodiments of the present invention further provide a heavy-haul train asynchronous control system for implementing the above-mentioned heavy-haul train asynchronous control method, as shown in the appendix Figure 4 and includes a locomotive side, a vehicle side, and a cloud side:
[0178] Among them, the locomotive side includes a decision-making module, and the decision-making module includes a hierarchical multi-agent decision-making model for giving real-time train operation decisions based on the real-time train operation information.
[0179] The vehicle side includes a measurement module, a communication module, and a power module; among them, the measurement module is used to obtain the real-time train operation information; the communication module is used to transmit the real-time train operation information and decision-making information between the vehicles of the train; the power module is used to output corresponding traction force and / or braking force based on the decision of the decision-making module to control the speed change of the heavy-haul train.
[0180] The cloud side includes a training module, which is used to pre-train the deep reinforcement learning models of each agent based on the initial data, continuously update the hierarchical multi-agent decision-making model based on the supplementary data uploaded by the train, self-train and optimize the hierarchical multi-agent decision-making model, and deploy the updated model to the locomotive side. In addition, the training module is also used to download information such as the line, terrain, and speed limit before each operation and cross-check with the measurement module.
[0181] Finally, it should be noted that the various embodiments in this specification are described in a progressive manner. The key points of each embodiment are the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other.
[0182] The above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them; although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that: still can modify the specific implementation manners of the present invention or perform equivalent replacements on some technical features; without departing from the spirit of the technical solutions of the present invention, they should all be covered within the scope of the technical solutions claimed by the present invention.
Claims
1. A heavy-load train asynchronous control method, characterized in that: The following steps are involved: Acquire real-time operation information of each vehicle in a heavy-load train; the real-time operation information includes speed information, position information, slope information and environmental information of each train; Constructing a hierarchical multi-agent decision model and establishing a reinforcement learning model for each agent in the decision model; the multi-agent decision model includes a decision layer and an execution layer; The decision layer includes multiple levels; the first level has one agent; except for the first level, the number of agents in each level is determined by the number of agents in the previous level and the number of decisions made by each agent in the previous level; the agents in the first level receive real-time operation information of the vehicle and output decisions to the agents in the second level; except for the agents in the first level, each agent receives the decisions output by the agents in the previous level and refines and splits them into decision inputs for the next level; The execution layer includes the onboard controller of each vehicle in the heavy-load train, and the execution layer receives the decision output by the last level of the intelligent agent in the decision layer and translates it into the execution mechanism command to control the train travel; Iteratively train the reinforcement learning model of each agent in the hierarchical multi-agent decision model using a proximal strategy optimization algorithm based on the real-time operation information; The trained hierarchical multi-agent decision-making model is deployed to obtain the control decisions of each vehicle in the heavy-haul train and act on the traction and braking system of the heavy-haul train.
2. The asynchronous control method for heavy-load trains according to claim 1, characterized in that: Further comprising the steps of: When the number of decisions given by the agents in the previous layer cannot completely match the number of agents in the next layer, the interpolation method is used to resample to match the number of agents in the next layer.
3. The asynchronous control method for heavy-load trains according to claim 1, characterized in that: Each agent has a preset built-in basic decision, which takes the vehicle's current speed, maximum speed limit and slope as input and outputs a decision of the required granularity; For each agent, when the upper-level decision instruction exists, it executes the upper-level decision instruction; when the upper-level decision instruction does not exist, the built-in basic decision is used as the upper-level decision instruction.
4. The asynchronous control method for heavy-load trains according to claim 1, characterized in that: The method of establishing a reinforcement learning model for each agent in the decision model comprises the following steps: Define the decision space for each agent in, For each agent’s decision, k1 is the predetermined maximum braking threshold, k2 is the predetermined maximum traction threshold, and N n is the number of decisions made by the current agent to control the next level of agents; Define the state space of each agent This includes the positions of all vehicles s s , speed v , slope s trac , image preprocessing features img , the decision at the previous time step and the vehicle range s controlled by the current agent mask ; Designing rewards for each agent Among them, μ k is the weight of different rewards; Includes vertical impact bonus Speed Reward Upper level command rewards and safety margin rewards The calculation formula for the vertical impact reward is: The speed bonus is calculated as: The calculation formula for the upper-level instruction reward is: The calculation formula for the safety margin reward is: in, For intelligent agents The maximum coupler force among the controlled vehicles, and Agent The coupler forces before and after the controlled vehicle group, is the maximum coupler force on each vehicle, and v limit are the average speed and maximum speed limit of the controlled vehicle group respectively, For the superior agent The decisions made, For intelligent agents Decision making response, It is the preset maximum safety limit of the coupler force.
5. The asynchronous control method for heavy-load trains according to claim 1 or 4, characterized in that: The reinforcement learning model of each agent adopts the Actor-Critic architecture, including the behavior network Act θ (s) and comment network V θ (s), the method for establishing a reinforcement learning model for each agent in the decision model further comprises the following steps: Define the strategy for each agent in, To analyze the environmental observation results Mapping to Execution Decisions The function of For the superior agent The parameters of the neural network.
6. The asynchronous control method for heavy-load trains according to claim 5, characterized in that: The following steps are involved: In the process of iteratively training the reinforcement learning model of each agent, a stable loss function based on the proximal policy optimization algorithm is used to alleviate the training oscillation of each agent's strategy; The expression of the stable loss function based on the proximal strategy optimization algorithm is: Among them, E t To calculate the expectation of the function for each time step t, π θ (a t |s t ) is the parameter θ and the environment s t Next, the strategy generates action a t The probability of θ (s) given, For the parameter θ old and Environments t Next, the strategy generates action a t The probability of θ (s) given, is the advantage function value at time step t, is the importance weight, which indicates the ratio of the new strategy to the old strategy, and ∈ is the clipping range, which is used to limit the amplitude of the strategy update.
7. The asynchronous control method for heavy-load trains according to claim 1 or 6, characterized in that: The following steps are involved: A judgment standard for similar intelligent agents is established, and intelligent agents that meet the judgment standard are determined as similar intelligent agents. In the process of iteratively training the reinforcement learning model of each intelligent agent, network parameters are shared between similar intelligent agents.
8. The asynchronous control method for heavy-load trains according to claim 6, characterized in that: The method for iteratively training the reinforcement learning model of each agent in the hierarchical multi-agent decision model using a proximal policy optimization algorithm comprises the following steps: The encoder-decoder composite network structure is used to map the value of the observation space to the hidden code, and the decoder is used to map the hidden code to the decision to build the behavior network Act. θ (s); The review network V is constructed by using the time series feature extraction layer and the composite fully connected layer. θ (s); Building a behavioral network Act θ (s) and comment network V θ (s); Set up behavior network Act θ Optimization goal of (s): L Act =L CLIP (θ)+KL(π θ ); Among them, L CLIP (θ) is the stable loss function based on the proximal policy optimization algorithm, KL(π θ ) is the KL divergence of the current action distribution, which is used to measure the decision diversity of the decision; Setting optimization goals for the review network in, is the value of the advantage function calculated using the generalized advantage estimation method at time step t, The calculation formula is: in Among them, γ∈[0.9,1] is a discount factor, which is used to consider the discount of future rewards when calculating cumulative rewards; λ∈[0,1] is a hyperparameter of the generalized advantage estimation method, which is used to control the trade-off between bias and variance; is the time difference error at time step t; r t is the reward obtained at time step t, V(s t ) and V(s t+1 ) are state s t and t+1 The state value estimate of Set the minimum total training target L = -L Act +L cri , the model is iteratively trained based on minimizing the total training objective.
9. The asynchronous control method for heavy-load trains according to claim 8, characterized in that: The method for iteratively training a model based on minimizing the total training objective comprises the following steps: Each agent interacts with the interaction object, makes decisions based on observations, and obtains rewards; After each agent collects a predetermined amount of data, it packages the decision, observation, and reward and sends them to the proximal strategy optimization algorithm, and trains the strategy, Act based on the algorithm process. θ (s) and V θ (s) until the proximal strategy optimization algorithm converges.
10. A heavy-load train asynchronous control system, used to implement the heavy-load train asynchronous control method according to any one of claims 1 to 9, characterized in that: include: The locomotive side includes a decision module, wherein the decision module includes a hierarchical multi-agent decision model, and is used to make real-time train operation decisions based on real-time train operation information; The vehicle side includes a measurement module, a communication module and a power module; the measurement module is used to obtain real-time operation information of the train; the communication module is used to transmit real-time operation information and decision information between vehicles in the train; the power module is used to output corresponding traction and / or braking force based on the decision of the decision module; The cloud includes a training module, which is used to pre-train the deep reinforcement learning model of each agent based on initial data, continuously update the hierarchical multi-agent decision model based on the supplementary data uploaded by the train, and perform self-training optimization on the hierarchical multi-agent decision model.
Citation Information
Cited By
Dynamic adjustment method and device for traction control of heavy haul train and electronic equipment
CN121291160A