Formation train risk decision control method and system in partially observable environment
Through the combination of distributed partial observable Markov decision-making model and train longitudinal dynamic model, the challenge of risk decision-making control in train formation operation is solved, more timely risk perception and more correct risk control decisions are achieved, and the safety and reliability of train formation operation is improved.
Patent Information
- Application Number
- CN202510055365.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-01-14
AI Technical Summary
While improving operational efficiency, the train formation operation mode has increased the challenges of train operation safety management, including external environmental risks and internal vehicle communication interruptions, resulting in an increase in the possibility of operational risks and accidents.
The distributed partial observable Markov decision model is adopted, combined with the train longitudinal dynamic model, and the relative braking mode tracking interval distance inside the queue train is calculated, the minimum safe safety interval for the virtual reconnected train fleet operation is established, the discrete area of the virtual queue train risk is determined, and the risk function is set through the risk potential field to determine the value function. The model is iteratively optimized using the cyclic multi-agent depth deterministic strategy gradient algorithm to obtain the train risk decision model of the train.
It realizes more timely risk situation awareness and more correct risk control decisions in risk scenarios, improves the safety and reliability of train formation operations, and reduces the risk of accidents.
Smart Images

Figure CN120069523A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of train operation safety control, and particularly relates to a risk decision-making control method and system for formation trains in a partially observable environment. Background Art
[0002] With the development of the networked operation of urban rail transit, the phenomenon of passenger flow-capacity mismatch caused by the change of passenger flow over time has gradually become obvious, bringing a series of problems in terms of investment, service level, operation and maintenance cost, etc.: (1) With the growth of passenger flow demand, it is difficult for the capacity of some lines to match the growth rate of passenger flow, and the problem of insufficient capacity has gradually become prominent; (2) The distribution law of passenger flow travel shows the characteristics of spatio-temporal imbalance, and the phenomenon of passenger flow-capacity mismatch has gradually become prominent; (3) The empty running of vehicles during the initial stage and off-peak periods of the line will cause energy consumption waste and excessive vehicle wear, resulting in high operation and maintenance costs. To alleviate the above problems, the development of flexible formation technology is proposed, that is, multiple trains run in formation without relying on physical connection, use independent traction and braking systems, and operate in coordination with a very small interval by wireless communication. According to the real-time change of passenger flow, efficient utilization of resources such as vehicles and precise matching of passenger flow and vehicle flow are achieved through online dynamic adjustment.
[0003] However, while the formation operation mode improves the operation efficiency, it also brings new challenges to the train operation safety management. During the formation operation process of trains, it is necessary to cope with risks such as foreign object intrusion and traction power supply interruption in the external environment, and also face operation risks caused by communication interruption between vehicles and performance degradation of the train braking system within the formation. The generation and accumulation of various risks may cause train facility failures or accidents such as train rear-end collisions and derailments. Any potential risk may lead to disasters, reduce the reliability of the train operation control system, and threaten the train operation safety. Since the formation operation mode of trains shortens the operation interval and each moving body runs at the maximum value of the rated speed, and the average speed of urban rail transit can reach 80-120 kilometers per hour during high-speed operation, this also puts forward higher requirements for ensuring safety.
[0004] At present, the research on train formation operation at home and abroad mostly stays in the stage of theoretical exploration and preliminary practice. The types and sample sizes of risk data that occur in its actual operation are seriously insufficient, and the train operation environment is becoming increasingly complex. At the same time, due to the needs of technological development and system integration, the complexity and randomness of the operation risks of formation trains have increased accordingly, which has greatly hindered the subsequent research on the prediction of the operation risk situation of formation trains and timely decision-making control. Therefore, issues such as how to extract and analyze the characteristic elements of the risk scenarios of formation trains and how to optimize the decision-making control algorithm of trains under risk scenarios still need to be further studied. Therefore, it is of great significance and application value to carry out research on the risk decision-making control of train formations for the virtual formation operation mode of trains, which provides method support for predicting and evaluating potential risks in advance and in a timely manner, and provides theoretical and technical support for the safe and stable operation of train formations. Summary of the Invention
[0005] The purpose of the present invention is to provide a risk decision-making control method and system for formation trains in a partially observable environment, so as to achieve more timely risk situation perception and more correct risk control decisions for formation trains in risk scenarios, and to solve at least one of the technical problems existing in the above background technology.
[0006] In order to achieve the above purpose, the present invention adopts the following technical solutions:
[0007] In the first aspect, the present invention provides a risk decision-making control method for formation trains in a partially observable environment, including:
[0008] Obtain multi-dimensional operation state information of formation trains;
[0009] Use the pre-trained risk decision-making model for formation trains to process the obtained multi-dimensional operation state information of formation trains to obtain risk control decisions for formation trains; among them, training the distributed partially observable Markov decision-making model includes: combining the longitudinal train dynamics model, calculating the tracking interval distance of formation trains inside using the "relative braking mode" in the virtual coupled train tracking operation mode, establishing the minimum safe interval for the safe operation of virtual coupled train formations, and determining the risk discrete area of virtual formation trains; constructing a distributed partially observable Markov decision-making model, based on the risk discrete area of virtual formation trains, setting a risk function through a risk potential field, determining a value function, and iteratively optimizing the optimal parameters of the model through a cyclic multi-agent deep deterministic policy gradient algorithm to obtain a trained risk decision-making model for formation trains.
[0010] As a further limitation of the first aspect of the present invention, trains inside the virtual formation mode track and run using the "relative braking mode". By constructing the longitudinal train dynamics model, a safe braking model for formation trains is established, and the expected tracking interval is calculated, including:
[0011] During the operation of the train, the forces affecting the operation are: its own gravity, traction force or braking force, and resistance; through force analysis, the train movement process is determined;
[0012] Based on the train dynamics model, in the virtual coupled formation operation mode, the "relative braking mode" is adopted for the trains inside the formation to track and run. To ensure train operation safety, a safety braking model for train formation operation is established. At a certain moment, after the leading train performs normal braking and decelerates, the following train also decelerates accordingly, and then the ideal braking distance of the formation trains is calculated;
[0013] During the dynamic formation process, according to the most favorable braking of the leading train and the most unfavorable braking of the following train, the minimum braking distance of the virtual formation trains is calculated.
[0014] As a further limitation of the first aspect of the present invention, a risk decision-making model for formation trains is constructed based on the distributed Markov decision method. Suppose there are N trains in the formation, denoted as χ = {1, 2,..., N}. Based on each train i having its own state variables, the state space S of the formation trains is modeled; the state of the entire formation is the combination of the states of all trains, and the multi-dimensional space jointly constituted by the environmental state serves as the complete system state space;
[0015] The observation space of the train includes its own key state information and the environmental information within a certain range around it. The action space A of the formation trains is modeled. The dynamic behavior of the train is controlled by the continuous action of acceleration, and the train acceleration space is discretized into four discrete values;
[0016] Based on the probability distribution of the initial state of the model, at each time step, an action is taken, and the probability distribution on the state space is updated according to the transfer function model. The transition model of the following train in the coordinate system established by the track is determined, and the new interval distance after a time step is determined.
[0017] As a further limitation of the first aspect of the present invention, the risk situation of the formation trains is estimated by using the artificial potential field theory. Suppose each train runs in an artificial potential field. The leading train can be a target or an obstacle. There are attractive and repulsive forces between the train and the leading and following trains to avoid collisions between the train and obstacles; the magnitude of the potential energy between the formation trains is used to characterize the risk of the system at this moment; the risk state of the system is divided into three states: safe, warning, and emergency by setting a risk threshold; by assigning a higher negative reward to the state with higher risk, the train is encouraged to take safer and more effective actions to avoid collisions.
[0018] As a further limitation of the first aspect of the present invention, the R-MADDPG algorithm is used as the solver of the model. The policy network and the value network are initialized for each agent, the target policy network and the target value network are initialized, and the cumulative expected reward is determined; at each time step t, the train calculates the risk estimation values of taking different candidate actions in the current state according to its current observation and the hidden state of the recurrent neural network and selects the action that can maximize the long-term cumulative reward through the policy network; after the train agent executes the action, the environment transfers to the next state s t+1 according to the actions of all agents and returns a local observation and a reward to each agent. Each agent combines the current observation action reward next observation observations of other agents observations of other agents at the next time step and the hidden states of the recurrent neural network before and after the action execution and to form an experience tuple and store it in the experience replay buffer D i .
[0019] As a further limitation of the first aspect of the present invention, the policy network model and the value network model are updated, including:
[0020] Randomly extract a batch of experience samples from the experience replay buffer D i of each agent, and set the batch size to B to form new state transition data;
[0021] Calculate the cumulative reward brought by executing the current action through the Bellman equation, which is used to estimate the expected return that can be obtained by acting according to the target policy in the future state;
[0022] Calculate the Q value output by the value network through the Critic network;
[0023] Define the loss function of the value network, update the parameters of the value network by minimizing the loss function, and use the gradient descent method to update the weight parameters of the value network;
[0024] Update the weight parameters of the policy network with the policy gradient of the Q value, and use the cumulative expected reward function as the optimization objective function; update the parameters of the policy network through policy gradient ascent;
[0025] Update the parameters of the target policy network and the target value network according to the soft update rule;
[0026] The algorithm is evaluated for convergence by observing whether the average reward of the agent is stable or whether the changes in the parameters of the policy network and the value network are less than a threshold. The formation train risk decision-making model is obtained until convergence.
[0027] In a second aspect, the present invention provides a formation train risk decision and control system in a partially observable environment, including:
[0028] An acquisition module for acquiring multi-dimensional operation state information of the formation train;
[0029] A decision-making module for processing the acquired multi-dimensional operation state information of the formation train by using a pre-trained formation train risk decision-making model to obtain a formation train risk control decision; wherein, training the distributed partially observable Markov decision-making model includes: combining the longitudinal dynamics model of the train, calculating the tracking interval distance of the formation train adopting the "relative braking mode" inside the formation train in the virtual coupled train tracking operation mode, establishing the minimum safe interval for the safe operation of the virtual coupled train formation, and determining the risk discrete area of the virtual formation train; constructing a distributed partially observable Markov decision-making model, setting a risk function based on the risk discrete area of the virtual formation train to determine the value function, and iteratively optimizing the optimal parameters of the model through the cyclic multi-agent deep deterministic policy gradient algorithm to obtain a trained formation train risk decision-making model.
[0030] In a third aspect, the present invention provides a non-transitory computer-readable storage medium for storing computer instructions, which when executed by a processor, implement the formation train risk decision and control method in a partially observable environment as described in the first aspect.
[0031] In a fourth aspect, the present invention provides a computer device including a memory and a processor, where the processor and the memory communicate with each other, the memory stores program instructions executable by the processor, and the processor calls the program instructions to execute the formation train risk decision and control method in a partially observable environment as described in the first aspect.
[0032] In a fifth aspect, the present invention provides an electronic device including: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory to enable the electronic device to execute the instructions for implementing the formation train risk decision and control method in a partially observable environment as described in the first aspect.
[0033] Advantages of the present invention: Based on multi-sensor fusion technology, the multi-dimensional operation state information of the train is perceived in real time, and the real-time interaction of the state information of adjacent trains is realized based on vehicle-to-vehicle communication technology; a distributed partially observable Markov decision model for the operation of formation trains is defined, and based on the observations of formation trains in an uncertain environment and the ideal braking distance and minimum braking distance calculated by the longitudinal dynamics model of the train, an artificial potential field theory is used to construct a risk function, and the state space is discretized according to the risk function and a reward function is constructed; the R-MADDPG algorithm is used for centralized training, a historical observation experience pool is constructed, a batch of experience samples are randomly selected to continuously perform environment interaction and network training, and the network parameters are updated according to the soft update rule until the parameter change is less than the threshold to stop iteration, and the final policy model and value model network are obtained, which can stably and reliably adjust the strategy of formation trains.
[0034] The advantages of the additional aspects of the present invention will be more clearly given in the following description part, or can be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0036] Figure 1 It is a schematic diagram of the communication network topology and risk decision hierarchy of the formation train described in the embodiment of the present invention.
[0037] Figure 2 It is a schematic diagram of the train safety braking model and the formation train emergency safety braking model established based on multi-dimensional perception information such as train position, speed, and acceleration described in the embodiment of the present invention.
[0038] Figure 3 It is a schematic diagram of the specific process of the formation train operation risk decision control method based on distributed partially observable Markov decision described in the embodiment of the present invention.
[0039] Figure 4 It is a schematic diagram of the specific process of the formation train risk decision control algorithm described in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals indicate the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described through the drawings are exemplary and are only used to explain the present invention, and cannot be construed as a limitation of the present invention.
[0041] For ease of understanding the present invention, the following further explains the present invention with specific embodiments in conjunction with the accompanying drawings, and the specific embodiments do not constitute a limitation on the embodiments of the present invention.
[0042] Those skilled in the art should understand that the drawings are only schematic diagrams of the embodiments, and the components in the drawings are not necessarily essential for implementing the present invention.
[0043] Embodiment 1
[0044] In this Embodiment 1, first, a formation train risk decision control system in a partially observable environment is provided, including: an acquisition module for acquiring multi-dimensional operation state information of the formation train; a decision module for processing the acquired multi-dimensional operation state information of the formation train by using a pre-trained formation train risk decision model to obtain a formation train risk control decision; wherein, training the distributed partially observable Markov decision model includes: combining the longitudinal dynamics model of the train, calculating the tracking interval distance in the "relative braking mode" inside the formation train in the virtual coupled train tracking operation mode, establishing the minimum safe interval for the safe operation of the virtual coupled train formation, determining the risk discrete area of the virtual formation train; constructing a distributed partially observable Markov decision model, setting a risk function through a risk potential field based on the risk discrete area of the virtual formation train, determining a value function, and iteratively optimizing the optimal parameters of the model through a cyclic multi-agent deep deterministic policy gradient algorithm to obtain a trained formation train risk decision model.
[0045] In this embodiment, using the above system, a formation train risk decision control method in a partially observable environment can be realized, including: acquiring multi-dimensional operation state information of the formation train; processing the acquired multi-dimensional operation state information of the formation train by using a pre-trained formation train risk decision model to obtain a formation train risk control decision; wherein, training the distributed partially observable Markov decision model includes: combining the longitudinal dynamics model of the train, calculating the tracking interval distance in the "relative braking mode" inside the formation train in the virtual coupled train tracking operation mode, establishing the minimum safe interval for the safe operation of the virtual coupled train formation, determining the risk discrete area of the virtual formation train; constructing a distributed partially observable Markov decision model, setting a risk function through a risk potential field based on the risk discrete area of the virtual formation train, determining a value function, and iteratively optimizing the optimal parameters of the model through a cyclic multi-agent deep deterministic policy gradient algorithm to obtain a trained formation train risk decision model.
[0046] In this embodiment, based on multi-sensor fusion technology, the multi-dimensional operation state information of high-speed trains is sensed in real time, and vehicle-to-vehicle communication is established between the trains operating in a virtual coupled formation, specifically including:
[0047] By utilizing multi-sensor data such as Beidou, inertial navigation, speed sensors, transponders, lidar, and cameras, the real-time operating status information of the train is calculated in real-time based on the multi-sensor information fusion train positioning and perception algorithm; the train operating environment information is obtained through the train operating environment collaborative perception method based on information interaction, realizing high-precision positioning and continuous perception of the operating status of the virtual coupled formation train, providing perception communication technology support for train tracking interval control and train group operation planning adjustment in the virtual coupled mode, and ensuring the safe and efficient collaborative operation control of high-speed trains in the virtual coupled mode.
[0048] Since the formation operation of trains requires real-time and reliable communication, each train has a corresponding node to store real-time speed and position information in the information layer. To prevent data loss and data errors, each train communicates with its adjacent trains, and algebraic graph theory can be used to describe the communication relationship between the formation trains. The communication topology can be modeled as a graph G = (ν, ε), where ν represents nodes 1…, n, and ε represents the set of edges. Define the weighted adjacency matrix of the graph as A = [a ij n×n , which is used to characterize the communication relationship between trains. If node i can receive information from node j, then a ij = 1, otherwise a ij = 0. If for any a ij = 1 and a ij = 1, then the graph is called an undirected graph. As mentioned above, each train communicates with its adjacent trains. A can be expressed as shown in Equation (1):
[0049]
[0050] In the virtual formation mode, the trains inside the formation operate in a "relative braking mode" for tracking. By constructing the longitudinal dynamics model of the train, a safety braking model for the formation train is established, and the expected tracking interval is calculated. Specifically:
[0051] During the operation of the train, the forces affecting the operation are: its own gravity, traction or braking force, and resistance. Through force analysis, the train movement process can be described by Equation (2):
[0052]
[0053] In the formula, M, γ, u t , u b are all related to the train's own performance. M is the train mass, γ is the train's gyratory mass coefficient, u t , u b are the traction and braking coefficients respectively; p(t) and v(t) are the train position and speed at time t; F t (v) and F b (v) are the traction force and braking force of the train at the current speed, where the braking force is further divided into service braking force F sb and emergency braking force F eb ; the resistance of the train in the current state consists of the basic resistance R a (v) and additional resistances including the additional resistance f g (p) due to slope, the additional resistance f r (p) due to curve, and the additional resistance f t (p) due to tunnel.
[0054] Based on the train dynamics model, in the virtual coupled formation operation mode, the trains inside the formation adopt the "relative braking mode" to track and run. To ensure the safety of train operation, a safety braking model for train formation operation is established. When the leading train decelerates by service braking at a certain moment, the following train also decelerates accordingly, and then the ideal braking distance of the formation trains is calculated. The specific calculation formula is shown as formula (3) below:
[0055]
[0056] In the formula, and are the initial braking speed and the target speed of the leading train respectively; and are the initial braking speed and the target speed of the following train respectively, and d p is the safety margin.
[0057] During the dynamic formation process, the most unfavorable situation is that the leading train takes the most favorable braking, while the following train takes the most unfavorable braking. At this time, the minimum braking distance d l of the virtual formation trains is calculated as shown in formula (4 - 6) below:
[0058] d l = d fmax - d lmin + d p (4)
[0059]
[0060] In the formula, v 0 is the starting speed, a d is the maximum acceleration of train traction, a b is the emergency braking acceleration of the train, a bs is the service braking acceleration, t b is the communication delay within the operation section, t c is the general equipment response time after receiving the command, t re is the traction control cut-off time, t e is the emergency braking establishment time, t bu is the additional time for emergency braking establishment.
[0061] Based on the distributed Markov decision method, a risk decision-making model for formation trains is constructed, and each train in the formation is regarded as an agent. Suppose there are N trains in the formation, denoted as χ = {1, 2, …, N}, and each train i has its own state variables, including the position x i (distance relative to a certain reference point), speed v i , acceleration a i , braking system state b i (normal or faulty, etc.), communication connection state c i (communication delay with other trains and the control center), etc. Therefore, the state space S of the formation trains can be modeled as shown in Equation (7):
[0062] S i = {v i , x i , a i , b i , c i} (7)
[0063] The state of the entire formation is the combination of the states of all trains, which can be expressed as shown in Equation (8):
[0064] S = {S 1 , S 2 , …, S N} (8)
[0065] The state related to the environment is also considered, such as the track condition r (e.g., curve curvature, slope, track flatness, etc.), weather condition ω (sunny, rainy, snowy, foggy, etc., which will affect the braking distance and driving stability of the train), etc. Therefore, the complete system state space S is a multi-dimensional space jointly composed of the train state and the environment state.
[0066] Each train can only obtain partial state information of itself and environmental information within a limited range around it as observations. For example, a train can accurately know its own speed and braking system state, but for some states of other trains, it may only obtain partial information through communication or indirect estimation based on sensors. The observation space of the train contains its own key state information and environmental information within a certain range around it, and Equation (9) can be defined as:
[0067]
[0068] where is the train's estimate of the state of adjacent trains, and r and ω are local observations of the track and weather conditions.
[0069] Model the action space A of the formation train, considering that the dynamic behavior of the train is mainly controlled by the continuous action of acceleration. To simplify the model, consider discretizing the train acceleration space into four discrete values, as shown in the following equation (10):
[0070] A = {a 0 , a 1 , a 2 , a 3} (10)
[0071] The a 0 , a 1 , a 2 , a 3 in the formula respectively represent the control outputs for the four cases of rated acceleration, maintaining vehicle speed, rated braking, and emergency braking.
[0072] The transition function describes the dynamic state of the train after each time step. Based on the probability distribution of the initial (or current) state of the model, at each time step δt, an action is taken, and the probability distribution on the state space is updated according to the transition function model T(s, a, s′) = P(s′|s, a). Let v i , x i , a i be the speed, position, and acceleration of the i-th train respectively, and the time sample is δt. The following equation (11) represents the transition model of the following train in the coordinate system established by the track:
[0073]
[0074] Similarly, the transition model of the leading train in the track coordinate system can be obtained according to the above formula. The new headway after one time step is represented by the following equation (12):
[0075] d o (t + δt) = x T i-1 (t) - x T i (t) (12)
[0076] The main objective of the observation function Z(o, a, s, s′) = P(o|a, s, s′) is to calculate the distance traveled by the train according to the selected action a after one time step in the coordinate system. The distance between the train and the leading vehicle at the new moment can be calculated by equation (12).
[0077] By using the Artificial Potential Field (APF) theory to estimate the risk situation of formation trains, it is assumed that each train operates in an artificial potential field. The leading train can be either a target or an obstacle. There are both attractive and repulsive forces between trains, so as to avoid collisions between trains and obstacles. Therefore, the magnitude of the potential energy between formation trains can be used to characterize the magnitude of the system risk at this moment, and its calculation formula is shown as follows in equations (13 - 15):
[0078]
[0079] Where U ix (d ij ) represents the distance risk between formation trains, U iv (V ij ) represents the speed risk between formation trains, U i APF represents the system comprehensive risk, ω i is the environmental weight parameter at the current moment. By setting a risk threshold, the risk state of the system is divided into three states s 1 , s 2 , s 3 , corresponding to the three states of safety, warning, and emergency respectively, and the corresponding model is shown as follows in equation (16).
[0080]
[0081] By assigning higher negative rewards to states with higher risks, trains can be encouraged to take safer and more effective actions to avoid collisions. When there is new environmental information, it can be updated in real time, enabling the system to continuously adapt to changing conditions and maintain safe operation. In addition, the reward function assigns numerical values in the form of state - action pairs to encourage desired actions, and the designed reward function is shown as follows in equation (17):
[0082]
[0083] The R - MADDPG algorithm is selected as the solver for the model. This algorithm can simultaneously learn the action strategies and communication strategies of each agent, use local observations and the communication information of other agents as the input of the policy network, and then take actions and send communication messages according to its output. During the training process, the state transition data of all agents at the same moment are taken out from the experience pool to guide the training of the value network.
[0084] Initialize the policy network and the value network for each agent, where θ i and φ iThey are the parameters of the policy network and the value network, and the Actor network with weight parameter θ i is used to fit the policy function, and the Critic network with weight parameter φ i is used to fit the value function. and are the hidden states before and after action execution of the recurrent neural network (such as LSTM). At the same time, initialize the target policy network and the target value network and let Among them, the policy function of the i-th agent is denoted as μ i , and its cumulative expected reward can be expressed as shown in Equation (18):
[0085]
[0086] Among them, r i (t) represents the reward obtained by the i-th agent at time step t. The specific algorithm execution steps include:
[0087] (1) At each time step t, the train agent calculates the risk estimation values of different candidate actions in the current state according to its current observation and the hidden state of the recurrent neural network , and selects the action that can maximize the long-term cumulative reward through the policy network
[0088] (2) After the train agent executes the action, the environment transfers to the next state s t+1 according to the actions of all agents, and returns a local observation and a reward
[0089] (3) Each agent combines the current observation action reward next observation observations of other agents observations of other agents at the next time step and the hidden states before and after action execution of the recurrent neural network (such as LSTM) and to form an experience tuple and store it in the experience replay buffer D i .
[0090] Update the policy network model and the value network model. The specific steps include:
[0091] (1) From the experience replay buffer D of each agent iRandomly select a batch of experience samples, with the batch size being B, to form new state transition data, denoted as
[0092] (2) Calculate the cumulative reward brought by executing the current action through the Bellman equation, which is used to estimate the expected return that can be obtained by acting according to the target policy in the future state. The calculation is as shown in the following formula (19):
[0093]
[0094] where represents the immediate reward obtained by agent i at time step t in sample k, γ is the discount factor, is the target value network, which is used to calculate the target Q value. Its inputs include the observation at the next time step t + 1 the action generated by the target policy network based on the observation and hidden state at the next time step and the observations of other agents at the next time step and the hidden state of the recurrent neural network
[0095] (3) Calculate the Q value output by the value network through the Critic network, denoted as
[0096] (4) Define the loss function of the value network as the mean square error as shown in formula (20),
[0097]
[0098] and update the value network parameters by minimizing the loss function. Use the gradient descent method to update the weight parameters φ of the value network i . The update formula for the value network parameters is as shown in the following formula (22):
[0099]
[0100] (5) Update the weight parameters θ of the policy network using the policy gradient of the Q value i , and the cumulative expected reward function J(θ i ) is used as the optimization objective function. Its gradient can be expressed as shown in formula (22):
[0101]
[0102] Update the policy network parameters θ through policy gradient ascent i , and the update formula is as shown in the following formula (23):
[0103]
[0104] where aπ is the learning rate of the policy network.
[0105] (6) Update the parameters of the target policy network and the target value network according to the soft update rule, as shown in Equation (24):
[0106]
[0107] where k represents the number of training iterations, and τ ∈ (0, 1) is the soft update coefficient.
[0108] Repeat the above processes of agent interaction, experience collection, and network training. After multiple time steps, evaluate whether the algorithm converges by observing whether the average reward of the agent (the reward after considering risk adjustment) is stable, or observing whether the changes in the parameters of the policy network and the value network are less than a threshold ∈, as and
[0109] After the algorithm converges, the policy network of each train agent can be used to make risk - controllable and efficient decisions based on the observed information during actual operation, so as to achieve the safe and efficient operation of the formation trains.
[0110] Embodiment 2
[0111] In this Embodiment 2, a risk control decision - making method for formation trains based on an interactive partially observable Markov decision process is provided to achieve more timely risk situation perception and more correct risk control decisions for formation trains in risk scenarios.
[0112] In this embodiment, the risk control decision - making method for formation trains based on a distributed partially observable Markov decision process includes: establishing vehicle - to - vehicle communication between the trains operating in virtual coupled formation, and continuously monitoring and evaluating the environmental collision risk based on vehicle - to - vehicle communication and multi - sensor fusion technology to real - time perceive the multi - dimensional operation state information of the formation trains. By using the collected multi - dimensional state information, continuous risk estimation is carried out using the occurrence probability and severity of the partially observable Markov decision process. The real - time risk situation of the formation trains is estimated by setting a reward function and a risk function, and a timely control method is adopted to keep it at an acceptable risk level all the time.
[0113] Since the current moving block mode cannot fundamentally meet the demand for improving transport capacity, in order to further shorten the headway, based on the relative braking distance driving mode, the concept of "virtual train formation" is proposed to achieve the reasonable utilization of resources. The trains use wireless communication technology to establish vehicle - to - vehicle communication with adjacent trains, and the vehicles share the operation status through information interaction. Combining the ground reference information and the vehicle - to - vehicle communication data, the follower vehicle can dynamically complete the formation operation during the running process.
[0114] Since each platooning train communicates with its adjacent trains, algebraic graph theory can be used to describe the communication relationship between platooning trains. The communication topology can be modeled as a graph G=(ν,ε), where ν represents nodes 1…,n and ε represents the set of edges. Define the weighted adjacency matrix of the graph as A=[a ij n×n , which is used to characterize the communication relationship between trains, as shown in the following formula:
[0115]
[0116] Based on the train dynamics model, in the virtual platooning mode, the trains inside the platoon track and run in the "relative braking mode". To ensure train operation safety, a safety braking model for platoon train operation is established. At a certain moment, after the leading train performs normal braking and decelerates, the following train also decelerates accordingly, and then the ideal braking distance of the platoon train is calculated. The specific calculation formula is as follows:
[0117]
[0118] Calculate the platoon train in the most unfavorable situation (that is, the leading train takes the most favorable braking and the following train takes the most unfavorable braking). At this time, the minimum braking distance d l of the virtual formation train is calculated as follows:
[0119] d l =d fmax -d lmin +d p
[0120]
[0121] By using the Decentralized Partially Observable Markov Decision Process (Dec-POMDPs) to continuously monitor and evaluate the risk of the platoon train operation scenario, continuously estimate the risk from the occurrence probability and severity, construct a distributed partially observable Markov decision model for the platoon train, and always control it within an acceptable risk level. This method supports the autonomous decision-making ability in the platoon train and can achieve timely and safe decision control. Construct a distributed partially observable Markov decision model for the platoon train and define its state space, observation space, action space, transition function, etc.
[0122] By using the Artificial Potential Field (APF) theory to estimate the risk situation of formation trains, it is assumed that each train operates in an artificial potential field. The leading train can be either a target or an obstacle. There are both attractive and repulsive forces between trains, to avoid collisions between trains and obstacles. Therefore, the magnitude of the potential energy between formation trains can be used to characterize the magnitude of the system risk at this moment, and its calculation formula is shown as follows:
[0123]
[0124] By setting a risk threshold, the risk state of the system is divided into three states s 1 , s 2 , s 3 , corresponding to the three states of safety, warning, and emergency respectively, and the corresponding model is shown as follows:
[0125]
[0126] In addition, the reward function assigns a value in the form of a state-action pair to encourage the desired action, and the designed reward function is shown as follows:
[0127]
[0128] To solve the optimal strategy, the R-MADDPG algorithm is selected as the model solver. The local observation and the communication information of other agents are used as the input of the policy network, and then actions are taken and communication messages are sent according to its output. The state transition data of all agents at the same moment is taken out from the experience pool to guide the training of the value network.
[0129] During the training process, a batch of experience samples is randomly drawn from the experience replay buffer D i of each agent. Let the batch size be B, and a new state transition data is formed, denoted as The training of the value network is guided by calculating the state transition data at the same moment. The cumulative reward brought by executing the current action is calculated through the Bellman equation, which is used to estimate the expected return that can be obtained by acting according to the target policy in the future state, and the calculation is shown as follows:
[0130]
[0131] The Q value output by the value network is calculated through the Critic network. The loss function of the value network is defined as the mean square error, and the parameters of the value network are updated by minimizing the loss function. The gradient descent method is used to update the weight parameters φ i of the value network, and the update formula of the value network parameters is shown as follows:
[0132]
[0133] Update the weight parameter θ of the policy network using the policy gradient of Q value i ,accumulate the expected reward function J(θ i ) as the optimization objective function, and update the policy network parameter θ by policy gradient ascent i ,and the representation form is as shown in the following formula:
[0134]
[0135] Update the parameters of the target policy network and the target value network according to the soft update rule, as shown in the following formula:
[0136]
[0137] Repeat the above agent interaction, experience collection and network training processes. After multiple time steps, evaluate whether the algorithm converges by observing whether the average reward of the agent (the reward after considering risk adjustment) is stable, or observing whether the changes in the parameters of the policy network and the value network are less than a threshold ∈, as and
[0138] Embodiment 3
[0139] In this Embodiment 3, a collision prevention control method for high-speed train formations based on deep learning and model prediction algorithms is provided, as Figure 1 shown, and it includes the following processing steps:
[0140] Step 1: Based on multi-sensor real-time perception of the multi-dimensional operation state information of the formation trains, establish a vehicle-to-vehicle communication link between the trains operating in virtual coupled formation, and construct a vehicle-to-vehicle communication topology structure.
[0141] Based on the current most advanced multi-sensor fusion technology, comprehensively and multi-angularly perceive the multi-dimensional state information of high-speed trains in complex operating environments in real time and accurately. By establishing an efficient, stable and highly anti-interference vehicle-to-vehicle communication link between each train in the virtual coupled formation operation mode. Using advanced communication technologies and equipment to ensure the reliability and stability of communication, and then construct a scientific, reasonable, stable and reliable communication topology structure that can adapt to different operating scenarios. Such a communication topology structure can not only ensure the fast and accurate transmission and sharing of information between trains, but also maintain good communication performance in the face of various emergencies and complex operating conditions, providing strong technical support for the safe, stable and efficient operation of high-speed train virtual coupled formations.
[0142] Since each formation train communicates with its adjacent trains, algebraic graph theory can be used to describe the communication relationship between formation trains. The communication topology can be modeled as a graph G=(ν,ε), where ν represents nodes 1…,n and ε represents the set of edges. Define the weighted adjacency matrix of the graph as A=[a ij n×n , which is used to characterize the communication relationship between trains and can be expressed as shown in Equation (1):
[0143]
[0144] Step 2: Combine the longitudinal train dynamics model to calculate the ideal braking distance of the formation train using the rated braking deceleration and the minimum braking distance under the most unfavorable conditions in the relative braking mode.
[0145] Based on the above scheme, the specific operations of Step 2 include:
[0146] Step 2-1: The forces acting on the train during operation are: its own gravity, traction or braking force, and resistance. Through force analysis, the train movement process can be described by Equation (2):
[0147]
[0148] In the formula, M, γ, u t , u b are all related to the train's own performance. M is the train mass, γ is the train's rotary mass coefficient, u t , u b are the traction and braking coefficients respectively; p(t) and v(t) are the train position and speed at time t; F t (v) and F b (v) are the traction and braking forces of the train at the current speed respectively, where the braking force is further divided into service braking F sb and emergency braking F eb ; the resistance of the train in the current state consists of the basic resistance R a (v) and the additional resistance including the grade additional resistance f g (p), the curve additional resistance f r (p) and the tunnel additional resistance f t (p).
[0149] Step 2-2: Based on the train dynamics model, in the virtual formation mode, the trains inside the formation track and run in the "relative braking mode". To ensure train operation safety, a safety braking model for train formation operation is established. When the leading train decelerates by service braking at a certain moment, the following train also decelerates accordingly, and then calculate the ideal braking distance of the formation train. The specific calculation formula is as shown in Equation (3) below:
[0150]
[0151] In the formula, and are respectively the initial braking speed and the target speed of the leading train; and are respectively the initial braking speed and the target speed of the following train, and d p is the safety margin.
[0152] Step 2-3: Calculate the minimum braking distance d l of the formation train under the most unfavorable condition (i.e., the leading train takes the most favorable braking while the following train takes the most unfavorable braking). The calculation formula of d
[0153] d l = d fmax - d lmin + d p (4)
[0154]
[0155] In the formula, v 0 is the starting speed, a d is the maximum acceleration of train traction, a b is the emergency braking acceleration of the train, a bs is the service braking acceleration, t b is the communication delay within the running section, t c is the general equipment response time after receiving the command, t re is the traction control cut-off time, t e is the emergency braking establishment time, t bu is the additional time for emergency braking establishment.
[0156] The risk control system takes the internal information of the formation train state and the external information of the environment as inputs. The internal inputs include sensor information about the positions and speeds of individual trains (generally provided by the positioning and speed measurement modules), as well as the emergency braking capabilities, which can be converted into the criteria for train stopping and the emergency distances. By selecting the safest operation strategy under different risk scenarios, the decision-making control of risks is achieved.
[0157] Step 3: Since there are uncertain parts in the operating state and environmental conditions of the formation trains, the Decentralized Partially Observable Markov Decision Process (Dec-POMDPs) is a probabilistic method for modeling the continuous process of a system under uncertain conditions. It is a generalization of the Markov decision process when part of the system state is unknown. The risk of the formation train operation scenario can be continuously monitored and evaluated by using Dec-POMDP. By continuously estimating the risk from the occurrence probability and severity, it is always controlled within an acceptable risk level. This method supports the autonomous decision-making ability in the formation trains and can achieve timely and safe decision control.
[0158] Step 3-1: Construct a decentralized partially observable Markov decision model for the formation trains.
[0159] Regard each train in the formation as an agent. Suppose there are N trains in the formation, denoted as χ = {1, 2, …, N}. Each train i has its own state variables, including the position x i (distance relative to a certain reference point), speed v i , acceleration a i , braking system state b i (normal or faulty, etc.), communication connection state c i (communication delay with other trains and the control center), etc. Therefore, the state space S of the formation trains can be modeled as shown in Equation (7):
[0160] S i = {v i , x i , a i , b i , c i} (7)
[0161] The state of the entire formation is the combination of the states of all trains, which can be expressed as shown in Equation (8):
[0162] S = {S 1 , S 2 , …, S N} (8)
[0163] The state related to the environment is also considered, such as the track condition r (e.g., curve curvature, slope, track flatness, etc.), weather condition ω (sunny, rainy, snowy, foggy, etc., which will affect the braking distance and driving stability of the train), etc. Therefore, the complete system state space S is a multi-dimensional space jointly composed of the train state and the environmental state.
[0164] Each train can only obtain its own partial state information and the environmental information within a limited surrounding range as observations. For example, a train can accurately know its own speed and the status of the braking system, but for some states of other trains, it may only obtain partial information through communication or make an indirect estimation based on sensors. The observation space of a train includes its own key state information and the environmental information within a certain surrounding range, which can be defined as shown in Equation (9):
[0165]
[0166] where is the estimation of the state of adjacent trains, and r and ω are the observations of the local track and weather conditions.
[0167] The risk decision control system takes the internal information of the train state and the external information of the environment as inputs. The internal inputs include the information about the state of its own train, and the external inputs include the information about the trains in front and behind and the surrounding environment, including their positions, speeds, and accelerations, etc. They can be converted into the rated safety distance and the minimum braking distance for stopping the train. The output of the system is the appropriate control actions taken to avoid collisions with the detected trains in front and behind and obstacles.
[0168] Model the action space A of the formation trains. Considering that the dynamic behavior of the train is mainly controlled by the continuous action of acceleration. To simplify the model, we consider discretizing the train acceleration space into four discrete values, as shown in Equation (10):
[0169] A = {a 0 , a 1 , a 2 , a 3} (10)
[0170] The a 0 , a 1 , a 2 , a 3 in the equation represent the control outputs for the four cases of rated acceleration, maintaining vehicle speed, rated braking, and emergency braking respectively.
[0171] The transition function depicts the dynamic state of the train after each time step. Based on the probability distribution of the initial (or current) state of the model, at each time step δt, an action is taken, and the probability distribution on the state space is updated according to the transition function model T(s, a, s′) = P(s′|s, a). Let v i , x i , a i be the speed, position, and acceleration of the i-th train respectively, and the time sample is δt. The following Equation (11) represents the transition model of the following train in the coordinate system established by the track:
[0172]
[0173] According to the above formula, the transition model of the leading train in the track coordinate system can also be obtained. The new headway after a time step is shown in the following formula (12):
[0174] d o (t + δt) = x T i-1 (t) - x T i (t) (12)
[0175] The main objective of the observation function Z(o, a, s, s′) = P(o|a, s, s′) is to calculate the distance traveled by the train according to the selected action a after a time step in the coordinate system. The distance between the train and the leading train at the new moment can be calculated by formula (11).
[0176] Step 3-2: Define the risk function of the formation train and set the reward function of the model according to the magnitude of the risk value.
[0177] By using the artificial potential field theory (Artificial potential field, APF) to estimate the risk situation of the formation train, it is assumed that each train operates in an artificial potential field. The leading train can be either a target or an obstacle. There are both attractive and repulsive forces between the train and the leading and trailing trains to avoid collisions between the train and obstacles. Therefore, the magnitude of the potential energy between the formation trains can be used to characterize the magnitude of the system risk at this moment, and its calculation formula is shown in the following formulas (13-15):
[0178]
[0179] U i APF = ω i [U ix (x ij ) + U iv (v i )] (15)
[0180] Among them, U ix (d ij ) represents the distance risk between the formation trains, U iv (V ij ) represents the speed risk between the formation trains, U i APF represents the system comprehensive risk, and ω i is the environmental weight parameter at the current moment. By setting the risk threshold, the risk state of the system is divided into three states s 1 , s 2,s 3 , corresponding to the three states of safety, warning, and emergency respectively, and the corresponding model is represented as shown in the following formula (16).
[0181]
[0182] The reward function takes the form of cost (or negative reward) and is assigned to each decision (action) made by the model in a specified state. The role of the reward function is to encourage those decisions that are close to the system goal and at the same time punish those decisions that are far from the system goal. According to this goal, negative rewards are assigned to those states that are considered unsafe, such as those states with a high probability of collision with the vehicle in front or obstacles. By assigning higher negative rewards to states with higher risks, the train can be encouraged to take safer and more effective actions to avoid collisions. When there is new environmental information, it can be updated in real time, enabling the system to continuously adapt to changing conditions and maintain safe operation. In addition, the reward function assigns numerical values in the form of state-action pairs to encourage desired actions, and the designed reward function is shown in the following formula (17):
[0183]
[0184] Step 4: To solve the optimal policy, the R-MADDPG algorithm is selected as the solver for the model. This algorithm can simultaneously learn the action policies and communication policies of each agent, and allows agents to obtain other observations through communication during the distributed execution process to improve the information modeling ability. Use the local observation and the communication information of other agents as the input of the policy network, and then take actions and send communication messages according to its output. During the training process, the state transition data of all agents at the same moment are taken out from the experience pool to guide the training of the value network.
[0185] Step 4-1: Initialize the policy network for each agent and the value network where θ i and φ i are the parameters of the policy network and the value network respectively. Use the Actor network with weight parameter θ i to fit the policy function, and use the Critic network with weight parameter φ i to fit the value function. and are the hidden states before and after the action execution of the recurrent neural network (such as LSTM). At the same time, initialize the target policy network and the target value network and let and where the policy function of the i-th agent is represented as μ i , and its cumulative expected reward can be represented as shown in formula (18):
[0186]
[0187] where r i (t) represents the reward obtained by the i-th agent at time t.
[0188] Step 4-2: First, select an action based on the observation and historical hidden state. The environment transfers to the next state according to the actions of all agents and returns the observation and reward to the agents. Subsequently, environmental feedback and experience storage are performed. The specific algorithm execution steps include:
[0189] (1) At each time step t, the train agent calculates the risk estimation values of taking different candidate actions in the current state based on its current observation and the recurrent neural network hidden state and selects the action that maximizes the long-term cumulative reward through the policy network
[0190] (2) After the train agent executes the action, the environment transfers to the next state s t+1 and returns a local observation and a reward
[0191] (3) Each agent combines the current observation action reward next observation observations of other agents observations of other agents at the next time step and the hidden states before and after the action execution of the recurrent neural network (such as LSTM) and to form an experience tuple and store it in the experience replay buffer D i .
[0192] Step 4-3: Update the parameters of the policy network model and the value network model. The specific steps include:
[0193] (1) Randomly sample a batch of experience samples from the experience replay buffer D i of each agent. Let the batch size be B and form new state transition data, denoted as
[0194] (2) Calculate the cumulative reward brought by executing the current action through the Bellman equation to estimate the expected return that can be obtained by acting according to the target policy in the future state. The calculation is as shown in the following formula (19):
[0195]
[0196] Among them, represents the immediate reward obtained by agent i at time step t in sample k, and γ is the discount factor. is the target value network, which is used to calculate the target Q value. Its inputs include the observation at the next time step t + 1 the action generated by the target policy network based on the observation and hidden state at the next time step and the observations of other agents at the next time step and the hidden state of the recurrent neural network
[0197] (3) Calculate the Q value output by the value network through the Critic network, denoted as
[0198] (4) Define the loss function of the value network as the mean squared error as shown in Equation (20).
[0199]
[0200] And update the parameters of the value network by minimizing the loss function, and use the gradient descent method to update the weight parameters φ of the value network i , the update formula of the value network parameters is as shown in the following Equation (21):
[0201]
[0202] (5) Update the weight parameters θ of the policy network using the policy gradient of the Q value i , and the cumulative expected reward function J(θ i ) is used as the optimization objective function, and its gradient can be expressed as in Equation (22):
[0203]
[0204] Update the policy network parameters θ through policy gradient ascent i , and the update formula is as shown in the following Equation (23):
[0205]
[0206] where a π is the learning rate of the policy network.
[0207] Step 4-4: Perform target network update and iterative convergence.
[0208] Update the parameters of the target policy network and the target value network according to the soft update rule, as shown in Equation (24):
[0209]
[0210] Where k represents the number of training iterations, and τ ∈ (0, 1) is the soft update coefficient.
[0211] Repeat the above processes of agent interaction, experience collection, and network training. After multiple time steps, evaluate whether the algorithm converges by observing whether the average reward of the agent (the reward considering risk adjustment) is stable, or observing whether the changes in the parameters of the policy network and the value network are less than a threshold ∈, as and
[0212] When the algorithm converges, the policy network of each train agent can be used to make risk - controllable and efficient decisions based on the observed information during actual operation, so as to achieve the safe and efficient operation of the formation trains.
[0213] It can be seen from the technical solutions provided in this embodiment above that this embodiment is based on multi - sensor fusion technology to real - time sense the multi - dimensional operation state information of trains, and based on vehicle - to - vehicle communication technology to realize the real - time interaction of adjacent train state information; define a distributed partially observable Markov decision model for the operation of formation trains, and based on the observed information of formation trains in an uncertain environment and the ideal braking distance and minimum braking distance calculated by the train longitudinal dynamics model, construct a risk function using the artificial potential field theory, discretize the state space according to the risk function and construct a reward function; use the R - MADDPG algorithm for centralized training, construct a historical observation experience pool, randomly extract a batch of experience samples to continuously perform environment interaction and network training, update the network parameters according to the soft update rule until the parameter changes are less than the threshold to stop the iteration, and obtain the final policy model and value model network, which can stably and reliably adjust the formation train strategy.
[0214] Embodiment 4
[0215] This Embodiment 4 provides a non - transient computer - readable storage medium, which is used to store computer instructions. When the computer instructions are executed by a processor, the above - mentioned formation train risk decision - making control method in a partially observable environment is realized.
[0216] Embodiment 5
[0217] This Embodiment 5 provides a computer device, including a memory and a processor. The processor and the memory communicate with each other. The memory stores program instructions executable by the processor, and the processor calls the program instructions to execute the above - mentioned formation train risk decision - making control method in a partially observable environment.
[0218] Embodiment 6
[0219] Embodiment 6 of the present invention provides an electronic device, including: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory, so that the electronic device executes the instructions for implementing the formation train risk decision control method in the partially observable environment as described above.
[0220] In summary, the formation train risk decision control method in the partially observable environment described in the embodiments of the present invention. The method includes: real-time sensing of multi-dimensional train operation state information based on multi-sensor fusion technology, real-time interaction of adjacent train state information based on vehicle-to-vehicle communication technology, and construction of a formation train communication network topology; calculation of the ideal braking interval between adjacent trains and the minimum safety interval under the most unfavorable conditions based on a dynamic model; continuous monitoring and evaluation of the formation train operation environment risk using distributed partially observable Markov decision-making under the condition of uncertainty in the train operation state and environmental conditions, and quantitative description of the real-time risk situation using the artificial potential field theory; using the cyclic multi-agent deep deterministic policy gradient algorithm as the solver of the model, iteratively updating the parameters of the optimal policy network and the value network, and realizing the risk decision control of the formation train; the present invention can realize the risk decision control of the formation train in the partially observable environment.
[0221] Although the specific implementation manners of the present invention are described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that based on the technical solutions disclosed in the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts should be covered within the protection scope of the present invention.
Claims
1. A risk decision control method for platooning vehicles in a partially observable environment, characterized in that: include: Obtain multi-dimensional operation status information of platoon vehicles; The pre-trained platoon risk decision model is used to process the acquired platoon multi-dimensional operation status information to obtain the platoon risk control decision; wherein, the distributed partially observable Markov decision model is trained, including: combining the longitudinal dynamics model of the train, calculating the tracking interval distance of the "relative braking mode" inside the platoon under the virtual reconnection tracking operation mode, establishing the minimum safe interval for the virtual reconnection train platoon operation safety, and determining the virtual platoon risk discrete area; constructing a distributed partially observable Markov decision model, based on the virtual platoon risk discrete area, setting the risk function through the risk potential field, determining the value function, and iteratively optimizing the optimal parameters of the model through the cyclic multi-agent deep deterministic policy gradient algorithm to obtain the trained platoon risk decision model.
2. The risk decision control method for platooning vehicles in a partially observable environment according to claim 1 is characterized in that: In the virtual formation mode, the train adopts the "relative braking mode" to track and run. By constructing the longitudinal dynamics model of the train, the formation train safety braking model is established, and the expected tracking interval is calculated, including: During the operation of the train, the forces that affect its operation include: its own gravity, traction or braking force, and resistance; through force analysis, the train movement process is determined; Based on the train dynamics model, the train adopts the "relative braking mode" to track and run in the virtual multiple-unit formation operation mode. To ensure the driving safety, a safe braking model for train formation operation is established. At a certain moment, after the leading train performs common braking deceleration, the following train also decelerates accordingly, and then the ideal braking distance of the formation train is calculated. In the dynamic marshaling process, the minimum braking distance of the virtual marshaling train is calculated based on the most favorable braking of the front car and the most unfavorable braking of the rear car.
3. The risk decision control method for platooning vehicles in a partially observable environment according to claim 1, characterized in that: A risk decision model for train formation is constructed based on the distributed Markov decision method. Assume that there are N trains in the formation, denoted as χ = {1, 2, ..., N}. Based on the fact that each train i has its own state variables, the state space S of the train formation is modeled; the state of the entire formation is the combination of all train states, and the multidimensional space formed by the environmental state is used as the complete system state space. The observation space of the train contains its own key state information and the surrounding environment information within a certain range. The action space A of the train formation is modeled. The dynamic behavior of the train is controlled by the continuous effect of acceleration, and the train acceleration space is discretized into four discrete values. Based on the probability distribution of the initial state of the model, an action is taken at each step, and the probability distribution in the state space is updated according to the transfer function model to determine the transition model of the following train in the coordinate system established by the track and determine the new interval distance after a time step.
4. The risk decision control method for platooning vehicles in a partially observable environment according to claim 1, characterized in that: The risk situation of trains in formation is estimated by using the theory of artificial potential energy field. It is assumed that each train runs in an artificial potential energy field. The front train can be a target or an obstacle. There is attraction and repulsion between the train and the front and rear trains to avoid collision between the train and the obstacles. The size of the system risk at this moment is characterized by the size of the potential energy of the train. By setting risk thresholds, the system's risk state is divided into three states: safe, warning, and emergency. By assigning higher negative rewards to higher-risk states, trains are encouraged to take safer and more effective actions to avoid collisions.
5. The risk decision control method for platooning vehicles in a partially observable environment according to claim 1, characterized in that: The R-MADDPG algorithm is used as the solver of the model. The policy network and value network are initialized for each agent, the target policy network and target value network are initialized, and the cumulative expected reward is determined. At each time step t, the train is trained according to its current observation and the hidden state of the recurrent neural network Calculate the risk estimate of taking different candidate actions in the current state, and select the action that maximizes the long-term cumulative reward through the policy network; after the train agent performs the action, the environment transfers to the next state s according to the actions of all agents t+1 , and returns a local observation to each agent and rewards Each agent will observe the current action award Next observation Observations by other agents Observations of other agents at the next time step and the hidden states of the recurrent neural network before and after the action is executed and Form an experience tuple and store it in the experience playback buffer D i middle.
6. The risk decision control method for platooning vehicles in a partially observable environment according to claim 5 is characterized in that: Updates to the policy network model and value network model include: Replay the experience buffer D from each agent i Randomly select a batch of experience samples from , set the batch size to B, and form new state transition data; The Bellman equation is used to calculate the cumulative reward of executing the current action, which is used to estimate the expected return that can be obtained by acting according to the target strategy in the future state; Calculate the Q value output by the value network through the Critic network; Define the loss function of the value network, update the value network parameters by minimizing the loss function, and use the gradient descent method to update the weight parameters of the value network; Use the policy gradient of the Q value to update the weight parameters of the policy network, and use the accumulated expected reward function as the optimization objective function; update the policy network parameters through policy gradient ascent; Update the parameters of the target policy network and the target value network according to the soft update rule; Whether the algorithm has converged is evaluated by observing whether the average reward of the agent is stable or whether the change in the parameters of the strategy network and the value network is less than a threshold. Once convergence occurs, a risk decision model for platooning vehicles is obtained.
7. A risk decision control system for platooning vehicles in a partially observable environment, characterized in that: include: An acquisition module is used to obtain multi-dimensional operation status information of platoon vehicles; The decision module is used to use the pre-trained platoon risk decision model to process the acquired platoon multi-dimensional operation status information to obtain the platoon risk control decision; wherein, the distributed partially observable Markov decision model is trained, including: combining the train longitudinal dynamics model, calculating the tracking interval distance of the platoon adopting the "relative braking mode" in the virtual reconnection tracking operation mode, establishing the minimum safe interval for the virtual reconnection train platoon operation safety, and determining the virtual platoon risk discrete area; constructing a distributed partially observable Markov decision model, based on the virtual platoon risk discrete area, setting the risk function through the risk potential field, determining the value function, and iteratively optimizing the optimal parameters of the model through the cyclic multi-agent deep deterministic policy gradient algorithm to obtain a trained platoon risk decision model.
8. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by the processor, the risk decision control method for platooning vehicles in a partially observable environment as described in any one of claims 1 to 6 is implemented.
9. A computer device, characterized in that: It includes a memory and a processor, the processor and the memory communicate with each other, the memory stores program instructions that can be executed by the processor, and the processor calls the program instructions to execute the risk decision control method for platoon vehicles in a partially observable environment as described in any one of claims 1 to 6.
10. An electronic device, characterized in that: include: A processor, a memory and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory so that the electronic device executes instructions for implementing the risk decision control method for platoon vehicles in a partially observable environment as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Multi-unmanned aerial vehicle formation cluster control method based on multi-agent deep reinforcement learning
CN115755949A
Saturation attack method and device for patrolling ammunition and storage medium
CN116068889A
Multi-agent self-organizing synergistic hunting method in non-convex environment
CN117574950A
Unmanned aerial vehicle formation coordination control method
CN118244799A
Multi-agent path planning method based on deep reinforcement learning
CN118536684A