Formation train risk decision control method and system in partially observable environment

By combining distributed partially observable Markov decision models and artificial potential field theory with multi-sensor fusion technology and the R-MADDPG algorithm, the problem of risk prediction and control in train formation operation was solved, and safe and efficient train operation in complex environments was achieved.

CN120069523BActive Publication Date: 2025-12-05BEIJING JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510055365.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-12-05
Estimated Expiration
2045-01-14

AI Technical Summary

Technical Problem

In train convoy operations, existing technologies struggle to effectively predict and control risk scenarios, leading to reduced operational safety, especially in high-speed and complex environments where there is a risk of accidents.

Method used

By employing a distributed partially observable Markov decision model and artificial potential field theory, combined with multi-sensor fusion technology and the R-MADDPG algorithm, the train status is perceived in real time. By setting a risk function through a risk potential field, the decision model is optimized to achieve risk situation perception and control.

Benefits of technology

It improves the safety and reliability of train platooning operations, enables timely prediction and assessment of potential risks, and ensures the safe and efficient operation of trains in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069523B_ABST
    Figure CN120069523B_ABST
Patent Text Reader

Abstract

The application provides a kind of platoon train risk decision control method and system under partially observable environment, belongs to train operation safety control technical field, obtains the multi-dimensional running state information of platoon train;Using the platoon train risk decision model trained in advance, the multi-dimensional running state information of platoon train obtained is processed, and the platoon train risk control decision is obtained.The application uses distributed partially observable Markov decision under the condition that the running state of train and environmental condition exist uncertainty, continuously monitors and evaluates the running environment risk of platoon train, quantitatively describes the real-time risk situation using artificial potential field theory;Adopt cyclic multi-agent deep deterministic policy gradient algorithm as the solver of model, iteratively update the optimal policy network and value network parameters, realize the risk decision control of platoon train, realize the risk decision control of platoon train under partially observable environment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of train operation safety control, and particularly relates to a train formation risk decision control method and system in a partially observable environment. BACKGROUND

[0002] With the development of urban rail transit network operation, the phenomenon of passenger flow-power mismatch caused by the change of passenger flow over time gradually becomes obvious, which brings a series of problems in investment, service level, operation and maintenance cost, etc.: (1) With the growth of passenger flow demand, the speed of capacity improvement of some lines is difficult to match the speed of passenger flow growth, and the problem of insufficient capacity gradually becomes prominent; (2) The distribution law of passenger flow presents the characteristics of space-time imbalance, and the phenomenon of passenger flow-power mismatch gradually becomes prominent; (3) The empty running of vehicles in the initial stage of the line and in the off-peak period will cause energy waste and high vehicle wear and tear, resulting in high operation and maintenance cost. In order to alleviate the above problems, the development of flexible marshalling technology is proposed, that is, multiple trains run in a formation mode without relying on physical connection and tracking, using independent traction and braking systems, and running cooperatively with a small interval through wireless communication, and according to the real-time change of passenger flow, the efficient use of vehicle resources and the accurate matching of passenger flow and vehicle flow are realized through online dynamic adjustment.

[0003] However, the formation running mode not only improves the running efficiency, but also brings new challenges to the train operation safety management. The train formation running process needs to deal with risks such as foreign matter intrusion in the external environment and interruption of traction power supply, and also faces the running risks caused by communication interruption between trains and performance degradation of train braking system, etc. The generation and accumulation of various risks may cause train facility and equipment failure or train collision, derailment and other accidents, any potential risk may lead to disaster, reduce the reliability of train operation control system and threaten the safety of train operation. Since the train formation running mode shortens the running interval and each moving body runs at the maximum value of the rated speed, the average speed of urban rail transit in high-speed operation can reach 80-120 kilometers per hour, which also puts higher requirements on the guarantee of safety.

[0004] At present, the research on train formation operation at home and abroad mostly stays in the stage of theoretical exploration and preliminary practice. The risk data types and sample sizes in actual operation are seriously insufficient. The train operation environment is increasingly complex. Meanwhile, due to the development of technology and the demand for system integration, the complexity and randomness of the operation risk of the formation train are also increasing, which greatly hinders the subsequent research on the risk situation prediction and timely decision control of the formation train. Therefore, how to extract and analyze the features of the risk scene elements of the formation train and how to optimize the decision control algorithm of the train in the risk scene are still to be further studied. Therefore, it is of great significance and application value to carry out the research on the risk decision control of the train formation in the virtual formation operation mode, which provides method support for early and timely prediction and evaluation of potential risks and provides theoretical and technical support for realizing the safe and stable operation of the train formation. SUMMARY

[0005] The purpose of the present application is to provide a formation train risk decision control method and system in a partially observable environment, which realizes more timely risk situation awareness and more correct risk control decision of the formation train in a risk scene, so as to solve at least one technical problem existing in the background technology.

[0006] In order to achieve the above-mentioned purpose, the present application adopts the following technical scheme:

[0007] In a first aspect, the present application provides a formation train risk decision control method in a partially observable environment, comprising:

[0008] Obtaining multi-dimensional running state information of the formation train;

[0009] Using a pre-trained formation train risk decision model to process the obtained multi-dimensional running state information of the formation train to obtain a formation train risk control decision; wherein, training the distributed partially observable Markov decision model comprises: combining a train longitudinal dynamics model to calculate the interval distance tracked by the formation train in the "relative braking mode" in the virtual reconnection tracking operation mode, establishing the minimum safety interval of the virtual reconnection train formation operation safety, determining the risk discrete area of the virtual formation train; constructing a distributed partially observable Markov decision model, setting a risk function based on the risk discrete area of the virtual formation train, determining a value function, and iteratively optimizing the optimal parameters of the model through a recurrent multi-agent deep deterministic policy gradient algorithm to obtain a trained formation train risk decision model.

[0010] As a further limitation of the first aspect of the present application, the train inside the virtual formation mode adopts the "relative braking mode" for tracking operation. Through the construction of the train longitudinal dynamics model, a formation train safety braking model is established, and the expected tracking interval is calculated, comprising:

[0011] The forces affecting the operation of the train during operation are: its own gravity, traction or braking force, resistance; through force analysis, the train movement process is determined;

[0012] Based on the train dynamics model, the train inside the virtual recombination train operation mode adopts the "relative braking mode" to track the operation, and the train formation operation safety braking model is established to ensure the train operation safety, the following train is also decelerated after the lead train decelerates by normal braking at a certain moment, and then the ideal braking distance of the train formation is calculated;

[0013] During dynamic marshalling, the most favorable braking is taken by the front vehicle, and the most unfavorable braking is taken by the rear vehicle, and the minimum braking distance of the virtual marshalling train is calculated.

[0014] As a further limitation of the first aspect of the application, a train formation risk decision model is constructed based on a distributed Markov decision method, assuming that there are N trains in the train formation, denoted as χ={1,2,…,N}, based on the state variable of each train i, the state space S of the train formation is modeled; The state of the entire train formation is the combination of all train states, and the multi-dimensional space formed by the environment state is the complete system state space;

[0015] The observation space of the train contains the key state information of itself and the environmental information within a certain range, the action space A of the train formation is modeled, and the dynamics behavior of the train is controlled by the continuous action of acceleration, and the train acceleration space is discretized into four discrete values.

[0016] Based on the probability distribution of the initial state of the model, at each time step, an action is taken, and the probability distribution on the state space is updated according to the transition function model, the transition model of the following train in the coordinate system established by the track is determined, and the new interval distance after a time step is determined.

[0017] As a further limitation of the first aspect of the application, the risk situation of the train formation is estimated by using the artificial potential field theory, assuming that each train runs in an artificial potential field, the front vehicle can be a target or an obstacle, and there is an attractive force and a repulsive force between the train and the front and rear vehicles to avoid collision between the train and the obstacle; The size of the system risk at this moment is represented by the size of the potential energy between the train formation; By setting a risk threshold, the risk state of the system is divided into safe, warning and emergency; By assigning a higher negative return to a state with higher risk, the train is encouraged to take more safe and effective actions to avoid collision.

[0018] As a further limitation of the first aspect of the application, the R-MADDPG algorithm is used as a solver for the model, and the policy network and the value network are initialized for each agent, the target policy network and the target value network are initialized, and the cumulative expected reward is determined; at each time step t, the trainee selects an action according to its current observation and the hidden state of the recurrent neural network The risk estimate value of taking different candidate actions in the current state is calculated, and the action that can maximize the long-term cumulative reward is selected through the policy network; after the trainee agent performs the action, the environment is transferred to the next state s t+1 and returns a local observation and a reward to each agent. action reward next observation observation of other agents observation of other agents at the next time step and the hidden state of the recurrent neural network before and after the action is performed and composing an experience tuple, which is stored in the experience replay buffer D i .

[0019] As a further limitation of the first aspect of the application, the policy network model and the value network model are updated, including:

[0020] A batch of experience samples is randomly extracted from the experience replay buffer D i of each agent, and the batch size is set to B to form a new state transition data;

[0021] The cumulative reward brought by the current action is calculated through the Bellman equation, which is used to estimate the expected return that can be obtained by acting according to the target policy in the future state;

[0022] The Q value output by the Critic network is calculated through the Critic network;

[0023] The loss function of the value network is defined, and the value network parameters are updated by minimizing the loss function, and the gradient descent method is used to update the weight parameters of the value network;

[0024] The weight parameters of the policy network are updated using the policy gradient of the Q value, and the cumulative expected reward function is used as the optimization objective function; the policy network parameters are updated through policy gradient ascent;

[0025] The parameters of the target policy network and the target value network are updated according to the soft update rule;

[0026] The algorithm is evaluated whether it converges by observing whether the average reward of the agent is stable or observing whether the parameter changes of the policy network and the value network are less than a threshold until convergence, and a platoon train risk decision model is obtained.

[0027] In a second aspect, the present application provides a platoon train risk decision control system in a partially observable environment, comprising:

[0028] An acquisition module is configured to acquire multi-dimensional running state information of the platoon train.

[0029] A decision module is configured to utilize a pre-trained platoon train risk decision model to process the acquired multi-dimensional running state information of the platoon train, and obtain a platoon train risk control decision.

[0030] In a third aspect, the present application provides a non-transitory computer readable storage medium for storing computer instructions, which, when executed by a processor, implement the platoon train risk decision control method in a partially observable environment according to the first aspect.

[0031] In a fourth aspect, the present application provides a computer device comprising a memory and a processor, wherein the processor and the memory are in communication with each other, the memory stores program instructions executable by the processor, and the processor invokes the program instructions to execute the platoon train risk decision control method in a partially observable environment according to the first aspect.

[0032] In a fifth aspect, the present application provides an electronic device comprising a processor, a memory and a computer program, wherein the processor is connected with the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to enable the electronic device to execute instructions for implementing the platoon train risk decision control method in a partially observable environment according to the first aspect.

[0033] The application has the advantages that: based on multi-sensor fusion technology, real-time perception of train multi-dimensional running state information is realized, based on vehicle-vehicle communication technology, real-time interaction of adjacent train state information is realized; a distributed partially observable Markov decision model of formation train operation is defined, based on observation of the formation train in an uncertain environment and ideal braking distance and minimum braking distance calculated by a train longitudinal dynamics model, a risk function is constructed by using an artificial potential field theory, and the state space is discretized and a reward function is constructed according to the risk function; centralized training is carried out by using an R-MADDPG algorithm, a historical observation experience pool is constructed, a batch of experience samples are randomly extracted to continuously interact with the environment and train the network, the network parameters are updated according to a soft update rule, and the iteration is stopped until the parameter change is less than a threshold value, and finally the strategy model and the value model network are obtained, so that the strategy of the formation train can be stably and reliably adjusted.

[0034] The advantages of the additional aspects of the application will be more apparent from the following description section or will be understood through the practice of the application. BRIEF DESCRIPTION OF DRAWINGS

[0035] In order to more clearly illustrate the technical solutions of the embodiments of the application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0036] Figure 1 The formation train communication network topology and risk decision level schematic diagram described in the embodiments of the application.

[0037] Figure 2 The train safety braking model and formation train emergency safety braking model schematic diagram established based on multi-dimensional perception information such as train position, speed and acceleration described in the embodiments of the application.

[0038] Figure 3 The formation train operation risk decision control method based on distributed partially observable Markov decision described in the embodiments of the application.

[0039] Figure 4 The formation train risk decision control algorithm specific process schematic diagram described in the embodiments of the application. DETAILED DESCRIPTION

[0040] The embodiments of the application will be described in detail below, and the examples of the embodiments are shown in the drawings, wherein the same or similar reference signs represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below through the drawings are exemplary and are only used to explain the application, and cannot be explained as a limitation on the application.

[0041] In order to make the present application more easily understood, the present application will be further explained and described hereinafter with reference to the drawings and specific embodiments, and the specific embodiments do not constitute limitations to the embodiments of the present application.

[0042] It should be understood by those skilled in the art that the drawings are only schematic views of the embodiments, and the components in the drawings are not necessarily essential for implementing the present application.

[0043] Embodiment 1

[0044] In this embodiment 1, a train formation risk decision control system in a partially observable environment is provided, comprising: an acquisition module for acquiring multi-dimensional running state information of the train formation; a decision module for processing the acquired multi-dimensional running state information of the train formation by using a pre-trained train formation risk decision model to obtain a train formation risk control decision; wherein training the distributed partially observable Markov decision model comprises: combining a train longitudinal dynamics model to calculate the interval distance tracked by the train formation in a virtual reconnection tracking running mode using a "relative braking mode", establishing a virtual reconnection train formation running safety minimum safety interval, and determining a virtual train formation risk discrete region; constructing a distributed partially observable Markov decision model, setting a risk function based on the virtual train formation risk discrete region through a risk potential field, determining a value function, and iteratively optimizing the optimal parameters of the model through a recurrent multi-agent deep deterministic policy gradient algorithm to obtain the trained train formation risk decision model.

[0045] In this embodiment, the above system can be used to implement a train formation risk decision control method in a partially observable environment, comprising: acquiring multi-dimensional running state information of the train formation; processing the acquired multi-dimensional running state information of the train formation by using a pre-trained train formation risk decision model to obtain a train formation risk control decision; wherein training the distributed partially observable Markov decision model comprises: combining a train longitudinal dynamics model to calculate the interval distance tracked by the train formation in a virtual reconnection tracking running mode using a "relative braking mode", establishing a virtual reconnection train formation running safety minimum safety interval, and determining a virtual train formation risk discrete region; constructing a distributed partially observable Markov decision model, setting a risk function based on the virtual train formation risk discrete region through a risk potential field, determining a value function, and iteratively optimizing the optimal parameters of the model through a recurrent multi-agent deep deterministic policy gradient algorithm to obtain the trained train formation risk decision model.

[0046] In this embodiment, multi-dimensional running state information of the high-speed train is sensed in real time based on a multi-sensor fusion technology, and car-to-car communication is established between virtual reconnection formation running trains, specifically comprising:

[0047] By using Beidou, inertial navigation, speed sensor, transponder, laser radar and camera and other multi-sensor data, the train operation state information is solved in real time based on multi-sensor information fusion train positioning and perception algorithm; through the train operation environment collaborative perception method based on information interaction, the train operation environment information is obtained, the high-precision positioning and continuity perception of the virtual reconnection formation train operation state are realized, the perception and communication technology support is provided for the train tracking interval control and train group operation planning adjustment in the virtual reconnection mode, and the safe and efficient collaborative operation control of high-speed trains in the virtual reconnection mode is ensured.

[0048] Since the train formation operation needs real-time and reliable communication, each train has a corresponding node in the information layer to store real-time speed and position information. In order to prevent data loss and data errors, each train communicates with its adjacent train, and algebraic graph theory can be used to describe the communication relationship between formation trains. The communication topology can be modeled as a graph G=(v,ε), where v represents nodes 1,…,n, and ε represents the edge set. Define the weighted adjacency matrix of the graph as A=[a ij ] n×n , which represents the communication relationship between trains. If node i can receive information from node j, then a ij =1, otherwise a ij =0. If for any a ij =1, a ij =1, the graph is called undirected. As described earlier, each train communicates with its adjacent train. A can be represented as shown in equation (1):

[0049]

[0050] The trains in the virtual formation mode adopt the "relative braking mode" for tracking operation. By constructing a train longitudinal dynamics model, a formation train safety braking model is established, and the expected tracking interval is calculated, which is specifically:

[0051] During the operation of the train, the forces that affect the operation are: its own gravity, traction or braking force, and resistance. Through force analysis, the train movement process can be described by equation (2):

[0052]

[0053] In the formula, M, γ, u t , u b are related to the performance of the train, M is the mass of the train, γ is the train rotation mass coefficient, u t , u b are the traction and braking coefficients respectively; p(t) and v(t) are the train position and speed at time t respectively; F t (v) and F b(v) are the tractive and braking forces of the train at the current speed, respectively, where the braking force is further divided into service braking F sb and emergency braking F eb ; the resistance of the train at the current state is composed of basic resistance R a (v) and additional resistances including slope additional resistance f g (p), curve additional resistance f r (p) and tunnel additional resistance f t (p).

[0054] Based on the train dynamics model, the train inside the virtual recombination train operation mode adopts the "relative braking mode" to track the operation, in order to ensure the safety of train operation, the train formation operation safety braking model is established, after the lead train performs service braking deceleration at a certain time, the following train also decelerates, and then the ideal braking distance of the train formation is calculated, the specific calculation formula is shown in the following formula (3):

[0055]

[0056] In the formula, and are the initial speed and target speed of the lead train, respectively; and are the initial speed and target speed of the following train, respectively, and d p is the safety margin.

[0057] In the dynamic marshalling process, the most unfavorable situation is that the front vehicle takes the most favorable braking, and the rear vehicle takes the most unfavorable braking, at this time, the calculation formula of the minimum braking distance d l of the virtual marshalling train is shown in the following formula (4-6):

[0058] d l = d fmax -d lmin +d p (4)

[0059]

[0060] In the formula, v0 is the initial speed, a d is the maximum acceleration of the train, a b is the emergency braking acceleration of the train, a bs is the service braking acceleration, t b is the communication delay in the running section, t c is the general equipment reaction time after receiving the command, t re is the traction control cut-off time, t e is the emergency braking establishment time, and t bu is the additional time for emergency braking establishment.

[0061] A train platoon risk decision-making model is constructed based on the distributed Markov decision method, treating each train in the platoon as an intelligent agent. Let there be N trains in the platoon, denoted as χ = {1, 2, ..., N}, and each train i has its own state variables, including its position x. i (distance relative to a reference point), velocity v i acceleration a i Braking system status b i (Normal or faulty, etc.), Communication connection status c i (Communication delays with other trains and control centers), etc. Therefore, the state space S of the train platoon can be modeled as shown in equation (7):

[0062] S i ={v i ,x i ,a i ,b i ,c i} (7)

[0063] The state of the entire formation is a combination of the states of all trains, which can be represented as shown in equation (8):

[0064] S = {S1,S2,…,S} N} (8)

[0065] Environmental conditions are also considered, such as track conditions r (e.g., curve curvature, gradient, track smoothness, etc.) and weather conditions ω (sunny, rainy, snowy, foggy, etc., which affect the train's braking distance and driving stability). Therefore, the complete system state space S is a multi-dimensional space composed of both the train state and the environmental state.

[0066] Each train can only obtain partial state information about itself and environmental information within a limited range as observations. For example, a train can accurately know its own speed and braking system status, but for certain states of other trains, it may only be able to obtain partial information through communication or through indirect estimation based on sensors. The train's observation space includes its own key state information and environmental information within a certain range, which can be defined by equation (9):

[0067]

[0068] in It is the train's estimate of the state of adjacent trains, and r and ω are local track and weather condition observations.

[0069] The action space A of the platoon is modeled, considering that the dynamic behavior of the train is mainly controlled by the sustained effect of acceleration. In order to simplify the model, the train acceleration space is discretized into four discrete values, as shown in the following equation (10):

[0070] A = {a0, a1, a2, a3} (10)

[0071] In the formula, a0, a1, a2, a3 respectively represent the control output of four situations of rated acceleration, maintaining train speed, rated braking and emergency braking.

[0072] The transition function describes the dynamic state of the train after each time step, based on the probability distribution of the initial (or current) state of the model, an action is taken at each time δt, and the probability distribution on the state space is updated according to the transition function model T (s, a, s') = P (s'|s, a). Let v i ,x i ,a i be the speed, position and acceleration of the ith train, and the time sample be δt. The following equation (11) represents the transition model of the following train in the coordinate system established by the track:

[0073]

[0074] According to the above formula, the transition model of the leading train in the track coordinate system can also be obtained, and the new interval distance after a time step is represented by the following equation (12):

[0075] d o (t+δt) = x T i-1 (t) - x T i (t) (12)

[0076] The main goal of the observation function Z (o, a, s, s') = P (o|a, s, s') is to calculate the distance traveled by the train in the coordinate system after a time step according to the selected action a, and the distance between the train and the front car at the new time can be calculated by equation (12).

[0077] By using the artificial potential field theory (APF) to estimate the risk situation of the platoon train, it is assumed that each train runs in an artificial potential field, and the front car can be a target or an obstacle, and there is an attractive force and a repulsive force between the train and the front and rear cars to avoid collision between the train and the obstacle. Therefore, the size of the potential energy between the platoon trains can be used to represent the size of the system risk at this moment, and its calculation formula is shown in the following equations (13-15):

[0078]

[0079] where U ix (d ij ) represents the distance risk between platooning vehicles, U iv (V ij ) represents the speed risk between platooning vehicles, U i APF represents the system comprehensive risk, ω i is the environmental weight parameter at the current time. By setting the risk threshold, the risk state of the system is divided into three states s1, s2, s3, which correspond to safe, warning and emergency states respectively. The corresponding model is shown in the following formula (16).

[0080]

[0081] By assigning a higher negative reward to a state with higher risk, the train can be encouraged to take more safe and effective actions to avoid collision. When there is new environmental information, it can be updated in real time, so that the system can constantly adapt to changing conditions and maintain safe operation. In addition, the reward function assigns a value to the state-action pair to encourage desired actions. The designed reward function is shown in the following formula (17):

[0082]

[0083] The R-MADDPG algorithm is selected as the solver of the model. This algorithm can simultaneously learn the action strategy and communication strategy of each agent, using local observations and communication information of other agents as inputs to the policy network. Then, according to the output, it takes action and sends communication messages. During the training process, the state transition data of all agents at the same time is taken from the experience pool to guide the value network training.

[0084] Initialize the policy network and the value network for each agent, where θ i and φ i are the parameters of the policy network and the value network respectively. The policy function is fitted by the Actor network with weight parameter θ i , and the value function is fitted by the Critic network with weight parameter φ i . and are the hidden states of the recurrent neural network (such as LSTM) before and after action execution. At the same time, the target policy network and the target value network are initialized, and let where the policy function of the i-th agent is denoted as μ i , and the cumulative expected reward can be represented as shown in formula (18):

[0085]

[0086] where r i (t) represents the reward obtained by the ith agent at time t. The specific algorithm execution steps include:

[0087] (1) At each time step t, the train agent calculates the risk estimation value of taking different candidate actions in the current state according to its current observation and the recurrent neural network hidden state , and selects the action that can maximize the long-term cumulative reward through the policy network

[0088] (2) After the train agent performs the action, the environment is transferred to the next state s t+1 according to the actions of all agents, and returns a local observation and a reward to each agent

[0089] (3) Each agent will store an experience tuple composed of the current observation action reward next observation observations of other agents observations of other agents at the next time step and the hidden states of the recurrent neural network (such as LSTM) before and after action execution and into the experience replay buffer D i .

[0090] The policy network model and the value network model are updated, and the specific steps include:

[0091] (1) A batch of experience samples are randomly extracted from the experience replay buffer D i of each agent, and the batch size is set as B to form a new state transition data, denoted as

[0092] (2) The cumulative reward brought by the current action is calculated through the Bellman equation, which is used to estimate the expected return that can be obtained by acting according to the target policy in the future state, and the calculation is shown in the following formula (19):

[0093]

[0094] where represents the immediate reward obtained by agent i in time step t, sample k, and γ is the discount factor, For the target value network, the input includes the observation of the next time step t+1 The action generated by the target policy network according to the observation and hidden state of the next time step And the observation of other agents in the next time step And the hidden state of the recurrent neural network

[0095] (3) Calculate the Q value of the value network output by the Critic network, denoted as

[0096] (4) Define the loss function of the value network as the mean square error as shown in equation (20),

[0097]

[0098] And update the value network parameters by minimizing the loss function, using gradient descent to update the weight parameters φ of the value network i The update formula of the value network parameters is as follows:

[0099]

[0100] (5) Update the weight parameters θ of the policy network using the policy gradient of the Q value i , and the cumulative expected reward function J(θ i ) as the optimization objective function, the gradient can be expressed as:

[0101]

[0102] Update the policy network parameters θ i by policy gradient ascent, and the update formula is as follows:

[0103]

[0104] Where a π is the learning rate of the policy network.

[0105] (6) Update the parameters of the target policy network and the target value network according to the soft update rule, as shown in equation (24):

[0106]

[0107] Where k represents the number of training iterations, and τ∈(0,1) is the soft update coefficient.

[0108] The above-mentioned agent interaction, experience collection and network training process is repeated, and after multiple time steps, whether the algorithm converges is evaluated by observing whether the average reward (considering the risk-adjusted reward) of the agent is stable or whether the changes of the policy network and the value network parameters are less than a threshold value ∈, such as and

[0109] When the algorithm converges, the policy network of each train agent can be used to make risk-controllable and efficient decisions according to the observation information in actual operation, so as to realize the safe and efficient operation of the train formation.

[0110] Embodiment 2

[0111] In this embodiment 2, a train formation risk control decision method based on an interactive partially observable Markov decision process is provided to realize more timely risk situation awareness and more correct risk control decision of the train formation in a risk scenario.

[0112] In this embodiment, the train formation risk control decision method based on a distributed partially observable Markov decision process includes: establishing train- train communication between virtual re-connection formation running trains, and continuously monitoring and evaluating the environmental collision risk based on the train- train communication and multi-sensor fusion technology to realize real-time perception of the multi-dimensional running state information of the train formation. By using the collected multi-dimensional state information, the probability and severity of the partially observable Markov decision process are used for continuous risk estimation, the real-time risk situation of the train formation is estimated by setting the reward function and the risk function, and a timely control method is used to keep the risk level acceptable at all times.

[0113] Since the current mobile block method cannot fundamentally meet the demand for improving the transport capacity, in order to further shorten the running interval, based on the relative braking distance running method, the concept of "virtual train formation" is proposed to achieve the rational use of resources. The following train can dynamically complete the marshalling operation in the running process by using wireless communication technology and establishing train- train communication with the adjacent train, and sharing the running state between trains.

[0114] Since each train formation communicates with its adjacent train. Algebraic graph theory can be used to describe the communication relationship between the train formations. The communication topology can be modeled as a graph G = (v, e), where v represents the nodes 1…, n, and e represents the edge set. The weighted adjacency matrix of the graph is defined as A = [a ij ] n×n , which is used to represent the communication relationship between trains, as shown in the following formula:

[0115]

[0116] Based on the train dynamics model, the train inside the virtual formation mode adopts the "relative braking mode" to track the operation, in order to ensure the safety of train operation, the train formation operation safety braking model is established, the following train is also decelerated after the lead train decelerates by normal braking at a certain time, and then the ideal braking distance of the train formation is calculated, and the specific calculation formula is as follows:

[0117]

[0118] The calculation formula of the minimum braking distance d of the virtual formation train in the most unfavorable situation (i.e. the most favorable braking for the front vehicle, and the most unfavorable braking for the rear vehicle) is as follows: l The calculation formula of the minimum braking distance d of the virtual formation train in the most unfavorable situation (i.e. the most favorable braking for the front vehicle, and the most unfavorable braking for the rear vehicle) is as follows:

[0119] d l =d fmax -d lmin +d p

[0120]

[0121] By using the distributed partially observable Markov decision process (Dec-POMDPs) to continuously monitor and evaluate the risk of train formation operation scene, the risk is continuously estimated from the occurrence probability and severity, and the distributed partially observable Markov decision model of train formation is constructed, which is always controlled within the acceptable risk level. This method supports the autonomous decision-making ability in train formation, and can realize timely and safe decision control. The distributed partially observable Markov decision model of train formation is constructed, and its state space, observation space, action space, transition function and the like are defined.

[0122] By using the artificial potential field theory (APF) to estimate the risk situation of train formation, it is assumed that each train runs in an artificial potential field, and the front vehicle can be a target or an obstacle. There is an attractive force and a repulsive force between the trains and the front and rear vehicles to avoid collision between the trains and the obstacles. Therefore, the size of the potential energy between the trains can be used to represent the size of the system risk at this moment, and the calculation formula is as follows:

[0123]

[0124] By setting the risk threshold, the risk state of the system is divided into three states s1, s2, s3, which correspond to safe, warning and emergency respectively, and the corresponding model is represented as follows.

[0125]

[0126] In addition, the reward function is designed as follows by assigning a value to the state-action pair to encourage the desired action:

[0127]

[0128] In order to solve the optimal strategy, the R-MADDPG algorithm is selected as the solver of the model, using local observations and communication information of other agents as input of the policy network, and then taking action and sending communication message according to the output. The state transition data of all agents at the same time is taken out from the experience pool to guide the value network training.

[0129] During the training process, a batch of experience samples is randomly selected from the experience replay buffer D i of each agent, and the batch size is B, to form new state transition data, denoted as The value network is trained by calculating the state transition data at the same time. The cumulative reward of performing the current action is calculated by Bellman equation, which is used to estimate the expected return obtained by acting according to the target policy in the future state, and the calculation is shown as follows:

[0130]

[0131] The Q value output by the Critic network is calculated, the loss function of the value network is defined as the mean square error, and the value network parameters are updated by minimizing the loss function. The gradient descent method is used to update the weight parameters φ i of the value network, and the update formula of the value network parameters is shown as follows:

[0132]

[0133] The weight parameters θ i of the policy network are updated by the policy gradient of the Q value, and the cumulative expected reward function J(θ i ) is used as the optimization objective function. The policy network parameters θ i are updated by policy gradient ascent, and the expression is shown as follows:

[0134]

[0135] The parameters of the target policy network and the target value network are updated according to the soft update rule, as shown in the following formula:

[0136]

[0137] The above-mentioned agent interaction, experience collection and network training process is repeated, and after multiple time steps, whether the average reward (considering the risk-adjusted reward) of the observed agent is stable or whether the changes of the policy network and value network parameters are less than a threshold ∈ are observed to evaluate whether the algorithm converges, such as and

[0138] Embodiment 3

[0139] In this embodiment 3, a high-speed train formation anti-collision control method based on deep learning and model prediction algorithm is provided, as shown in the figure, which includes the following processing steps: Figure 1

[0140] Step 1: Based on the multi-sensor real-time perception of the multi-dimensional running state information of the formation train, a car-car communication link is established between the virtual reconnection formation running trains, and a car-car communication topology structure is constructed.

[0141] Based on the current most advanced multi-sensor fusion technology, the multi-dimensional state information of high-speed trains in complex running environment is perceived in all directions and multiple angles in real time and accurately. Through the establishment of an efficient, stable and strong anti-interference car-car communication link between each train in the virtual reconnection formation running mode, the reliability and stability of the communication are ensured by using advanced communication technology and equipment, and then a scientific and reasonable, stable and reliable communication topology structure that can adapt to different running scenarios is constructed. Such a communication topology structure not only ensures the rapid and accurate transmission and sharing of information between trains, but also maintains good communication performance when facing various emergencies and complex running conditions, providing solid technical support for the safe, stable and efficient operation of high-speed train virtual reconnection formation.

[0142] Since each formation train communicates with its adjacent trains, algebraic graph theory can be used to describe the communication relationship between the formation trains. The communication topology can be modeled as a graph G=(ν,ε), where ν represents the nodes 1…,n and ε represents the edge set. The weighted adjacency matrix of the graph is defined as A=[a ij ] n×n , which is used to represent the communication relationship between cars, and can be represented as formula (1) shown:

[0143]

[0144] Step 2: Combine the train longitudinal dynamics model to calculate the ideal brake distance using the rated brake deceleration and the minimum brake distance under the most unfavorable conditions for the formation train in the relative braking mode.

[0145] On the basis of the above scheme, the specific operation of step 2 includes:

[0146] ​Step 2-1: The forces acting on the train during operation include its own gravity, traction or braking force, and resistance. Through force analysis, the train movement process can be described by equation (2):

[0147]

[0148] In the formula, M, γ, u t , u b are all related to the performance of the train, M is the mass of the train, γ is the train rotational mass coefficient, u t , u b are the traction and braking coefficients respectively; p(t) and v(t) are the train position and speed at time t respectively; F t (v) and F b (v) are the traction and braking forces of the train at the current speed, where the braking force is divided into normal braking F sb and emergency braking F eb ; the resistance of the train in the current state is composed of basic resistance R a (v) and additional resistance including slope additional resistance f g (p), curve additional resistance f r (p), and tunnel additional resistance f t (p).

[0149] Step 2-2: Based on the train dynamics model, the train inside the virtual formation mode adopts the "relative braking mode" to track operation. To ensure train operation safety, a train formation operation safety braking model is established. After the lead train performs normal braking deceleration at a certain time, the following train also decelerates, and then the ideal braking distance of the formation train is calculated. The specific calculation formula is shown in equation (3) as follows:

[0150]

[0151] In the formula, and are the initial speed and target speed of the lead train respectively; and are the initial speed and target speed of the following train respectively, and d p is the safety margin.

[0152] Step 2-3: Calculate the minimum braking distance d l of the virtual formation train in the most unfavorable situation (i.e. the most favorable braking for the front vehicle, and the most unfavorable braking for the rear vehicle). The calculation formula of the minimum braking distance d l of the virtual formation train is shown in equations (4-6) as follows:

[0153] d l = d fmax -d lmin +d p(4)

[0154]

[0155] where v0is the initial speed, a d is the maximum acceleration of the train, a b is the emergency braking acceleration of the train, a bs is the normal braking acceleration, t b is the communication delay within the operation section, t c is the general device reaction time after receiving the command, t re is the traction control cut-off time, t e is the emergency braking establishment time, t bu is the additional time for emergency braking establishment.

[0156] The risk control system takes the internal information of the formation train state and the external information of the environment as input. The internal input includes sensor information about the position and speed of each train (usually provided by positioning and speed measurement modules), as well as emergency braking capability, which can be converted into standard and emergency stopping distances of the train. By selecting the safest operation strategy under different risk scenarios, decision control of risk is achieved.

[0157] Step 3: Due to the uncertainty of the running state of the formation train and the environmental conditions, the Decentralized Partially Observable Markov Decision Process (Dec-POMDP) is a probabilistic method for modeling continuous processes of systems under uncertain conditions, which is an extension of Markov Decision Process under partial unknown system state. By using Dec-POMDP to continuously monitor and evaluate the risk of formation train operation scenarios, the risk is continuously estimated from the probability of occurrence and severity, and always controlled within an acceptable risk level. This method supports autonomous decision-making capabilities in formation trains and enables timely and safe decision control.

[0158] Step 3-1: Construct a Decentralized Partially Observable Markov Decision Model for Formation Trains.

[0159] Each train in the formation is considered as an intelligent agent. Let there be N trains in the formation, denoted as χ = {1, 2, …, N}, each train i has its own state variables, including position x i (distance relative to a reference point), speed v i , acceleration a i , brake system status b i (normal or failure, etc.), communication connection status c i(Communication delays with other trains and control centers), etc. Therefore, the state space S of the train platoon can be modeled as shown in equation (7):

[0160] S i ={v i ,x i ,a i ,b i ,c i} (7)

[0161] The state of the entire formation is a combination of the states of all trains, which can be represented as shown in equation (8):

[0162] S = {S1,S2,…,S} N} (8)

[0163] Environmental conditions are also considered, such as track conditions r (e.g., curve curvature, gradient, track smoothness, etc.) and weather conditions ω (sunny, rainy, snowy, foggy, etc., which affect the train's braking distance and driving stability). Therefore, the complete system state space S is a multi-dimensional space composed of both the train state and the environmental state.

[0164] Each train can only obtain partial state information about itself and environmental information within a limited range as observations. For example, a train can accurately know its own speed and braking system status, but for certain states of other trains, it may only be able to obtain partial information through communication or through indirect estimation based on sensors. The train's observation space includes its own key state information and environmental information within a certain range, which can be defined as shown in equation (9):

[0165]

[0166] in It is the train's estimate of the state of adjacent trains, and r and ω are local track and weather condition observations.

[0167] The risk decision control system takes internal information about the train's status and external information about the environment as inputs. Internal inputs include information about the train's own status, while external inputs include information about preceding and following trains and the surrounding environment, such as their position, speed, and acceleration. These can be translated into the rated safe distance and minimum braking distance for stopping the train. The system's output is the appropriate control action taken to avoid collisions with detected preceding and following trains and obstacles.

[0168] The motion space A of the trains in the platoon is modeled, considering that the dynamic behavior of the train is mainly controlled by the continuous action of acceleration. To simplify the model, we consider discretizing the train acceleration space into four discrete values, as shown in Equation (10):

[0169] A = {a0, a1, a2, a3} (10)

[0170] a0, a1, a2, a3 in the formula represent the control output of the four cases of rated acceleration, maintaining speed, rated braking and emergency braking respectively.

[0171] The conversion function describes the dynamic state of the train after each time step, based on the probability distribution of the initial (or current) state of the model, an action is taken at each time step δt, and the probability distribution on the state space is updated according to the transition function model T (s, a, s') = P (s' | s, a). Let v i ,x i ,a i be the speed, position and acceleration of the ith train respectively, and the time sample is δt. The following formula (11) represents the transition model of the following train in the coordinate system established by the track:

[0172]

[0173] According to the above formula, the transition model of the leading train in the track coordinate system can also be obtained, and the new interval distance after a time step is shown in the following formula (12):

[0174] d o (t+δt) = x T i-1 (t) - x T i (t) (12)

[0175] The main goal of the observation function Z (o, a, s, s') = P (o | a, s, s') is to calculate the distance traveled by the train in the coordinate system after a time step according to the selected action a, and the distance between the train and the front car at the new time can be calculated by formula (11).

[0176] Step 3-2: Define the risk function of the formation train, and set the reward function of the model according to the size of the risk value.

[0177] By using the artificial potential field theory (APF) to estimate the risk situation of the formation train, it is assumed that each train runs in an artificial potential field, and the front car can be a target or an obstacle, and there is an attractive force and a repulsive force between the train and the front and rear cars to avoid collision between the train and the obstacle. Therefore, the size of the potential energy between the formation trains can be used to represent the size of the system risk at this moment, and its calculation formula is shown in the following formulas (13-15):

[0178]

[0179] U i APF = ω i [U ix (x ij )+ U iv (v i )] (15)

[0180] where U ix (d ij ) represents the distance risk between the platoon trains, U iv (V ij ) represents the speed risk between the platoon trains, U i APF represents the comprehensive risk of the system, and ω i is the environmental weight parameter at the current time. By setting the risk threshold, the risk state of the system is divided into three states s1, s2, and s3, corresponding to the safe, warning, and emergency states, respectively. The corresponding model is shown in the following formula (16).

[0181]

[0182] The reward function is in the form of cost (or negative reward) and is assigned to each decision (action) made by the model in a given state. The role of the reward function is to encourage decisions that are close to the system goal, while penalizing those that are far from the system goal. According to this goal, negative rewards are assigned to states that are considered unsafe, such as those with a high probability of collision with the preceding vehicle or obstacles. By assigning higher negative rewards to states with higher risks, the train can be encouraged to take safer and more effective actions to avoid collisions. When new environmental information is available, it can be updated in real time, allowing the system to continuously adapt to changing conditions and maintain safe operation. In addition, the reward function assigns numerical values to state-action pairs to encourage desired actions. The designed reward function is shown in the following formula (17):

[0183]

[0184] Step 4: To solve the optimal strategy, the R-MADDPG algorithm is selected as the solver of the model. This algorithm can simultaneously learn the action strategy and communication strategy of each agent, allowing agents to obtain other observations through communication during distributed execution to improve information modeling capabilities. Local observations and communication information from other agents are used as inputs to the policy network, and actions and communication messages are taken based on the output of the policy network during training. The state transition data of all agents at the same time is taken from the experience pool to guide the value network training.

[0185] Step 4-1: Initialize the policy network and value network where θ i and φ i are the parameters of the policy network and the value network, respectively, and the actor network with weight parameters θ i is used to fit the policy function, and the critic network with weight parameters φ i is used to fit the value function, and are the hidden states of the recurrent neural network (e.g., LSTM) before and after the action execution. The target policy network and the target value network are initialized simultaneously and let and where the policy function of the i-th agent is denoted as μ i , and the cumulative expected reward can be expressed as shown in equation (18):

[0186]

[0187] where r i (t) represents the reward obtained by the i-th agent at time t.

[0188] Step 4-2: First, the action is selected according to the observation and the historical hidden state, and the environment is transferred to the next state according to the actions of all agents, and the observation and reward are returned to the agents, and then the environment feedback and experience storage are performed. The specific algorithm execution steps include:

[0189] (1) At each time step t, the train agent calculates the risk estimate value of taking different candidate actions in the current state according to its current observation and the recurrent neural network hidden state , and selects the action that can maximize the long-term cumulative reward through the policy network

[0190] (2) After the train agent executes the action, the environment is transferred to the next state s t+1 according to the actions of all agents, and a local observation and a reward are returned to each agent

[0191] (3) Each agent forms an experience tuple consisting of the current observation action reward next observation observations of other agents observations of other agents at the next time step and the hidden states of the recurrent neural network (e.g., LSTM) before and after the action execution and and stores it in the experience replay buffer Di In.

[0192] Step 4-3: Update the parameters of the policy network model and the value network model, the specific steps include:

[0193] (1) Randomly sample a batch of experience samples from the experience replay buffer D i of each agent, set the batch size as B, and form a new state transition data, denoted as

[0194] (2) Calculate the cumulative reward brought by the current action through the Bellman equation, which is used to estimate the expected return obtained by acting according to the target policy in the future state, calculated as shown in equation (19):

[0195]

[0196] wherein, represents the immediate reward obtained by agent i in time step t, sample k, and γ is the discount factor, is the target value network, which is used to calculate the target Q value, its input includes the observation at the next time step t+1, the action generated by the target policy network according to the observation and hidden state of the next time step, and the observation and the hidden state of the recurrent neural network of other agents at the next time step

[0197] (3) Calculate the Q value output by the Critic network, denoted as

[0198] (4) Define the loss function of the value network as the mean square error as shown in equation (20),

[0199]

[0200] and update the value network parameters by minimizing the loss function, use gradient descent method to update the weight parameters φ i of the value network, the update formula of the value network parameters is shown in equation (21):

[0201]

[0202] (5) Update the weight parameters θ i of the policy network using the policy gradient of the Q value, the cumulative expected reward function J(θ i ) is used as the optimization objective function, and its gradient can be expressed as shown in equation (22):

[0203]

[0204] Update the policy network parameters θ by policy gradient ascent i , the update formula is shown in the following formula (23):

[0205]

[0206] Wherein a π is the learning rate of the policy network.

[0207] Step 4-4: Perform target network update and iterative convergence.

[0208] Update the parameters of the target policy network and the target value network according to the soft update rule, as shown in formula (24):

[0209]

[0210] Wherein k represents the number of training iterations, and τ∈(0, 1) is a soft update coefficient.

[0211] Repeat the above agent interaction, experience collection and network training process, and after a plurality of time steps, whether the average reward (considering the risk-adjusted reward) of the agent is stable or whether the change of the policy network and the value network parameters is less than a threshold ∈ is observed to evaluate whether the algorithm converges, such as and

[0212] When the algorithm converges, the policy network of each train agent can be used to make risk-controllable and efficient decisions according to the observation information in actual operation, so as to realize the safe and efficient operation of the train formation.

[0213] As can be seen from the technical solutions provided by the above embodiment, the embodiment realizes real-time perception of train multi-dimensional running state information based on multi-sensor fusion technology, realizes real-time interaction of adjacent train state information based on train-to-train communication technology, defines a distributed partially observable Markov decision model for train formation running, adopts an artificial potential field theory to construct a risk function according to the risk function, and constructs a reward function by discretizing the state space; the R-MADDPG algorithm is used for centralized training, a historical observation experience pool is constructed, a batch of experience samples are randomly extracted for continuous environment interaction and network training, the network parameters are updated according to the soft update rule, and the iteration is stopped until the parameter change is less than a threshold, so that the final policy model and value model network are obtained, and the train formation policy can be adjusted stably and reliably.

[0214] Embodiment 4

[0215] The embodiment 4 provides a non-transitory computer readable storage medium for storing computer instructions, which, when executed by a processor, implement the train formation risk decision control method in a partially observable environment as described above.

[0216] Embodiment 5

[0217] The embodiment 5 provides a computer device, comprising a memory and a processor, the processor and the memory are in communication with each other, the memory stores program instructions executable by the processor, and the processor invokes the program instructions to execute the train formation risk decision control method in a partially observable environment as described above.

[0218] Embodiment 6

[0219] The embodiment 6 provides an electronic device, comprising a processor, a memory and a computer program, wherein the processor is connected with the memory, and the computer program is stored in the memory; when the electronic device is running, the processor executes the computer program stored in the memory, so that the electronic device executes instructions for implementing the train formation risk decision control method in a partially observable environment as described above.

[0220] In summary, the train formation risk decision control method in a partially observable environment according to the embodiments of the present application. The method comprises the following steps: real-time sensing of multi-dimensional running state information of the train based on a multi-sensor fusion technology, real-time interaction of state information of adjacent trains based on a train-to-train communication technology, and construction of a train formation communication network topology; calculation of an ideal braking interval of adjacent trains and a minimum safety interval in the most unfavorable case based on a dynamics model; in the case that there is uncertainty in the running state and environmental conditions of the train, continuous monitoring and evaluation of the running environment risk of the train formation by using a distributed partially observable Markov decision, and quantitative description of the real-time risk situation by using an artificial potential field theory; adoption of a recurrent multi-agent deep deterministic policy gradient algorithm as a solver of the model, iterative updating of optimal policy network and value network parameters, and realization of risk decision control of the train formation; and the present application can realize risk decision control of the train formation in a partially observable environment.

[0221] Although the specific embodiments of the present application are described above in combination with the drawings, the description is not a limitation on the scope of protection of the present application, and those skilled in the art should understand that various modifications or changes made on the basis of the technical solutions disclosed in the present application without creative labor should be covered within the scope of protection of the present application.

Claims

1. A method for risk decision control of platoon trains in a partially observable environment, characterized in that, The method comprises the following steps: Obtain multi-dimensional running state information of the train formation; The pre-trained platoon train risk decision model is used to process the obtained multi-dimensional running state information of the platoon train, and a platoon train risk control decision is obtained. The platoon train risk decision model is trained, including: combining a train longitudinal dynamics model, calculating the interval distance of the platoon train using the "relative braking mode" in the virtual reconnection tracking mode, establishing the minimum safety interval of the virtual reconnection train platoon running safety, and determining the virtual platoon train risk dispersion area; a distributed partially observable Markov decision model is constructed, a risk function is set through a risk potential field based on the virtual platoon train risk dispersion area, a value function is determined, the optimal parameters of the model are iteratively optimized through a recurrent multi-agent deep deterministic policy gradient algorithm, and a trained platoon train risk decision model is obtained; wherein, the R-MADDPG algorithm is used as a model solver, a strategy network and a value network are initialized for each agent, a target strategy network and a target value network are initialized, and the cumulative expected reward is determined; at each time step , the train calculates the risk estimate value of different candidate actions under the current state according to its current observation and the recurrent neural network hidden state , and selects the action that can maximize the long-term cumulative reward through the strategy network; after the train agent performs the action, the environment is transferred to the next state according to the actions of all agents, and a local observation and a reward are returned to each agent; each agent stores an experience tuple composed of the current observation , action , reward , next observation , observation of other agents , observation of other agents at the next time step , and hidden states of the recurrent neural network before and after action execution and into an experience replay buffer .

2. The method of claim 1, wherein, In the virtual formation mode, the train inside adopts a "relative braking mode" for tracking operation; a train formation safety braking model is constructed through the construction of a train longitudinal dynamics model, and an expected tracking interval is calculated, including: During the operation of the train, the forces affecting the operation include the gravity of the train itself, traction or braking force, and resistance; through force analysis, the train motion process is determined; Based on the train dynamics model, in the virtual re-formation operation mode, the train inside adopts a "relative braking mode" for tracking operation; a train formation operation safety braking model is established to ensure train operation safety; after the lead train decelerates by normal braking at a certain time, the following train also decelerates, and then the ideal braking distance of the train formation is calculated; During the dynamic marshalling process, the most favorable braking is taken by the front train, while the most unfavorable braking is taken by the rear train, and the minimum braking distance of the virtual marshalling train is calculated.

3. The method of claim 1, wherein, A risk decision-making model for vehicle queuing is constructed based on the distributed Markov decision method, assuming that there are vehicles in the queuing. A train, denoted as Based on each train It has its own state variables, and the state space of the queuing vehicles is... Modeling is performed; the state of the entire formation is a combination of the states of all trains, which, together with the environmental state, constitute a multi-dimensional space as the complete system state space. The observation space of the train contains its own key state information and the environmental information within a certain range, which is used to model the action space of the train The dynamics of the train is controlled by the continuous acceleration, and the acceleration space of the train is discretized into four discrete values. Based on the probability distribution of the initial state of the model, at each time step, an action is taken, and the probability distribution on the state space is updated according to the transition function model; the transition model of the following train in the coordinate system established by the track is determined, and the new interval distance after a time step is determined.

4. The method of claim 1, wherein, The risk situation of the train formation is estimated by using the artificial potential field theory; it is assumed that each train runs in an artificial potential field; the front train can be a target or an obstacle; there is an attractive force and a repulsive force between the train and the front and rear trains to avoid collision between the train and the obstacle; the size of the potential energy between the train formation is used to represent the size of the system risk at this moment; The risk state of the system is divided into safe, warning and emergency states by setting a risk threshold; by assigning a higher negative reward to a state with higher risk, the train is encouraged to take more safe and effective actions to avoid collision.

5. The method of claim 1, wherein, The strategy network model and the value network model are updated, including: a batch of experience samples is randomly sampled from the experience replay buffer of each agent , and a new state transition data is composed , and a new state transition data is composed The cumulative reward brought by the current action is calculated by the Bellman equation, which is used to estimate the expected return that can be obtained in the future state by following the target strategy; calculating the value network output by the Critic network values; Define the loss function of the value network, update the parameters of the value network by minimizing the loss function, and use gradient descent method to update the weight parameters of the value network; With The weight parameters of the policy network are updated by using the policy gradient of the value, and the expected reward function is accumulated as an optimization objective function; the policy network parameters are updated by policy gradient ascent. According to the soft update rule, the parameters of the target strategy network and the target value network are updated; Whether the algorithm converges is evaluated by observing whether the average reward of the agent is stable or observing whether the change of the parameters of the strategy network and the value network is less than a threshold, until convergence, and the risk decision model of the train formation is obtained.

6. A platoon train risk decision control system in a partially observable environment, characterized in that, The method comprises the following steps: An acquisition module is used to acquire multi-dimensional running state information of the train formation; The decision module is used for processing the obtained multi-dimensional running state information of the train consist by using a pre-trained consist train risk decision model to obtain a consist train risk control decision; wherein, the consist train risk decision model is trained, including: combining a train longitudinal dynamics model, calculating a "relative braking mode” tracking interval distance inside the consist train in a virtual reconnection tracking running mode, establishing a virtual reconnection train consist running safety minimum safety interval, determining a virtual consist train risk dispersion area; a distributed partially observable Markov decision model is constructed, a risk function is set through a risk potential field based on the virtual consist train risk dispersion area, a value function is determined, the optimal parameters of the model are iteratively optimized through a recurrent multi-agent deep deterministic policy gradient algorithm to obtain the trained consist train risk decision model; wherein, the R-MADDPG algorithm is used as a model solver, a strategy network and a value network are initialized for each agent, a target strategy network and a target value network are initialized, and the cumulative expected reward is determined; at each time step , the train calculates the risk estimate value of different candidate actions under the current state according to its current observation and the recurrent neural network hidden state , and selects the action that can maximize the long-term cumulative reward through the strategy network; after the train agent performs the action, the environment is transferred to the next state according to the actions of all agents, and returns a local observation and a reward to each agent; each agent stores an experience tuple composed of the current observation , action , reward , next observation , observation of other agents , observation of other agents at the next time step , and hidden states of the recurrent neural network before and after action execution and in an experience replay buffer .

7. A non-transitory computer-readable storage medium, comprising: The non-transitory computer readable storage medium is used to store computer instructions, which are executed by the processor to implement the train formation risk decision control method in the partially observable environment according to any one of claims 1-5.

8. A computer device, comprising: The processor and the memory are in communication with each other, the memory stores program instructions executable by the processor, and the processor invokes the program instructions to execute the train formation risk decision control method in the partially observable environment according to any one of claims 1-5.

9. An electronic device, comprising: The method comprises the following steps: A processor, a memory and a computer program; wherein the processor is connected with the memory, the computer program is stored in the memory, when the electronic device is running, the processor executes the computer program stored in the memory, so that the electronic device executes the instructions of the platoon train risk decision control method in the partially observable environment as claimed in any one of claims 1-5.

Citation Information

Patent Citations

  • Multi-unmanned aerial vehicle formation cluster control method based on multi-agent deep reinforcement learning

    CN115755949A

  • Saturation attack method and device for patrolling ammunition and storage medium

    CN116068889A