Multi-vehicle joint longitudinal control method based on improved MADDPG

By improving the MADDPG algorithm, combined with the perception radius and reward attenuation value, the state and reward communication between multiple agents is optimized, which solves the traffic safety problem in the stop-and-go wave traffic scenario, effectively suppresses the spread of stop-and-go waves and improves traffic safety.

CN116935669BActive Publication Date: 2025-10-10SOUTHEAST UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202310893941.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-20
Publication Date
2025-10-10
Estimated Expiration
2043-07-20

AI Technical Summary

Technical Problem

Existing technologies pay little attention to micro-driving strategies for traffic safety in stop-and-go traffic scenarios and are unable to effectively suppress the occurrence of stop-and-go traffic phenomena.

Method used

The improved MADDPG algorithm is adopted, combined with the perception radius and reward decay value, to design the state communication method and reward communication method among multiple agents. Traffic safety is optimized through deep reinforcement learning. The perception radius D and reward decay value are introduced to construct the SRM-MADDPG algorithm to optimize the state and reward communication among multiple agents.

Benefits of technology

It can effectively suppress the spread of stop-and-go traffic waves, improve traffic safety, enhance cooperation among multiple intelligent agents, reduce algorithm computational complexity, and solve the "lazy intelligent agent" problem, which has certain foresight and practical significance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116935669B_ABST
    Figure CN116935669B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on the multi-vehicle joint longitudinal control method of improved MADDPG, including steps: S1, using simulation software SUMO to model traffic stop-and-go wave, model each vehicle and design the single lane travel route of vehicle;S2, using the traffic control interface Traci built-in simulation software SUMO and the Python of realizing deep reinforcement learning interact;S3, calibration and verification are carried out to simulation model;S4, according to multi-vehicle longitudinal control task, the joint state space of multiple agents, action space and joint reward are designed;SRM-MADDPG algorithm is constructed to improve the state communication mode and reward communication mode between multiple agents;S5, set policy objective function and value function.The application can enhance the cooperation between multiple agents, reduce the complexity of calculation, improve the performance of multi-agent system, effectively optimize traffic safety and suppress the occurrence of traffic stop-and-go wave phenomenon.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of multi-agent reinforcement learning, and in particular to a multi-vehicle joint longitudinal control method based on an improved MADDPG. Background Art

[0002] With the development of deep reinforcement learning, some studies have used deep Q networks (DQN) or deep recurrent Q networks (DRQN) to train independent learners and achieved good performance. Recently, the parameter sharing structure has been combined with the Trust Region Policy Optimization (TRPO) algorithm and applied to continuous action scenarios, and has performed well in environments with continuous action spaces.

[0003] The use of multi-agent algorithms to coordinate traffic flow control is a feasible and effective technology. In existing research, Chinese patent CN116150639A discloses a multi-agent variable speed limit control method based on behavioral trend clustering and feature mapping. By performing lateral feature mapping between source agents and shared agents in the same group, the learning process is accelerated, and finally the road speed limit is controlled. Chinese patent CN115100850A discloses a hybrid traffic flow control method based on a multi-performer critic framework, integrating importance sampling mechanism, long short-term memory network and attention mechanism, with road traffic rate as the optimization goal. However, the above studies pay less attention to the micro-driving strategy to improve traffic safety in stop-and-go traffic scenarios. Some studies use reinforcement learning to design reward functions to punish unsafe driving behaviors, but these methods cannot guarantee safety in the initial stages of training. Summary of the Invention

[0004] Purpose of the invention: The purpose of the present invention is to provide a multi-vehicle joint longitudinal control method based on improved MADDPG that can effectively optimize traffic safety and suppress the occurrence of traffic stop-go wave phenomenon.

[0005] Technical solution: The multi-vehicle joint longitudinal control method of the present invention includes the following steps:

[0006] S1, the simulation software SUMO was used to model the traffic stop-go wave, model each vehicle and design the vehicle's single-lane travel route;

[0007] S2 uses the built-in traffic control interface Traci in the simulation software SUMO to interact with Python for deep reinforcement learning to obtain real-time traffic environment information and control the longitudinal behavior of the vehicle;

[0008] S3, calibrate and verify the simulation model, using the data in the UTE trajectory database to calibrate the parameters of the IDM model;

[0009] S4, based on the multi-vehicle longitudinal control task, designs the joint state space, action space, and joint reward function of the multi-agents. It also introduces the perception radius D and reward decay value, and constructs the SRM-MADDPG algorithm to improve the state communication and reward communication methods among the multi-agents.

[0010] S5, set the policy objective function and value function.

[0011] Furthermore, in step S2, the TensorFlow deep learning framework is selected and a mature control strategy is trained. The implementation steps are as follows:

[0012] S21, abstracts the complex traffic simulation environment into a representative state feature vector as the input of the policy network and the value network. Different longitudinal control strategies extract different state feature vectors;

[0013] S22, define the structure of the policy network and value network according to the requirements, and define the mapping method from the state feature vector to the action feature vector;

[0014] S23, during the training of the reinforcement learning agent, the parameters of the policy network and value network represented by the TensorFlow method are adjusted through the training data generated by the traffic environment until the trained parameter structure is converged and saved;

[0015] S24, by calling the saved neural network parameters and,testing the performance of the trained control strategy in different scenarios.

[0016] Furthermore, in step S3, the detailed steps for calibrating the parameters of the IDM model using the data in the UTE trajectory database are as follows:

[0017] S31, extracting the trajectory data of the following vehicle, performing data cleaning and preprocessing to ensure data quality and consistency;

[0018] S32 defines the fitness function for evaluating the degree of IDM model parameter fitting and prediction accuracy, using the error between the model prediction and the actual observed trajectory as the evaluation indicator. The formula is as follows:

[0019]

[0020] Where n represents the total number of simulation steps, They represent the actual headway and simulated headway of the following vehicle at the kth step respectively;

[0021] S33, determining the IDM model parameters that need to be adjusted and defining the range of each parameter;

[0022] S34, setting the parameters of the genetic algorithm, including population size, number of iterations, crossover rate and mutation rate, and convergence tolerance;

[0023] S35, calibrate the IDM model using the UTE dataset.

[0024] Furthermore, in step S31, the pre-processed data needs to be screened, and the screening principles are as follows:

[0025] A1) The selected target vehicle always follows the preceding vehicle and its speed is no less than 10 km / h;

[0026] A2) The front and rear vehicles remain in the same lane, without changing lanes, and the distance between them is less than 120m;

[0027] A3) The front and rear vehicles are in the same lane and the distance between their heads is less than 120m.

[0028] Furthermore, in step S4, the design of the local joint state space, action space and decaying joint reward function of the multi-agent is as follows:

[0029] B1) Multi-vehicle local joint state space

[0030] The state information of the agents within the perception radius D is broadcast to the perception source agent. For agent n, the local joint state space with a perception radius of D is expressed as:

[0031]

[0032] in, represents the state information of the nth agent at time t, Represents the state information of the n-1th agent at time t, Represents the state information of the nDth agent at time t;

[0033] The state information of each agent contains 4 variables. Taking the nth agent as an example, its state information include:

[0034]

[0035] in, represents the speed of the nth agent at time t, represents the acceleration of the nth agent at time t, represents the distance between the nth agent and the preceding vehicle at time t;

[0036] When the number of agents within the perception radius is less than D, the input of state information at position 0 to position n-D is the state information of the leading vehicle, and the state information at position less than 0 is all recorded as 0;

[0037] B2) Multi-vehicle action space

[0038] The acceleration of each agent is controlled to be an arbitrary continuous variable between -3 m / s 2 and 3 m / s 2 ; each agent in the multi-agent system is subjected to a safety constraint:

[0039]

[0040] wherein L min is the minimum safety distance between two vehicles, and Δt is the duration of the strategy taken;

[0041] The action space of the multi-agent is represented as:

[0042]

[0043] wherein represents the acceleration taken by agent n at time t, j = 1, 2,..., n;

[0044] B3) Multi-vehicle decay joint reward

[0045] The decay joint reward is represented as:

[0046]

[0047] wherein R n represents the decay joint reward obtained by agent n, N represents the number of vehicles in the vehicle platoon, and θ is the discount factor; r i represents the reward value obtained by each agent only considering itself, and the reward value of agent n is represented as:

[0048] r n = αReward l + βReward v + λ(Reward a + Reward f )

[0049] wherein Reward l is the penalty value obtained after the vehicle exits the following mode; Reward v is the vehicle passing efficiency reward value; Reward a is the negative value of the square of the target vehicle acceleration; and Reward fis the negative value of the fuel consumption of the intelligent connected vehicle, and α, β, and λ are the weights of different components of the reward function.

[0050] Furthermore, in step S5, the policy objective function and value function are set as follows:

[0051] C1) Strategy objective function

[0052] The MADDPG approximate gradient formula is:

[0053]

[0054] in, represents the policy gradient of agent i; θ i represents the policy network parameters of agent i; J i (θ i ) represents the strategic performance of agent i; E represents the expected operation of the behavioral strategy of agent i; Represents the action a for agent i i In state s i The action-value function Q under i The gradient of i represents the input of the environment state observed by agent i;

[0055] Clipping MADDPG limits the policy update to a reasonable range, as shown in the following formula:

[0056]

[0057] Among them, ∈ is a hyperparameter; f is a function that represents the i (θ i ) to perform cropping operations, ensuring that J i (θ i ) will not exceed the specified range [1-∈, 1+∈];

[0058] By constraining the policy update and clipping operation clip(·), the policy optimization objective function is rewritten as:

[0059]

[0060] Among them, L CLIP (π θ ) represents the policy loss function calculated using the clipping operation in the MADDPG algorithm; π θ represents the policy network;

[0061] When the advantage function value A π When it is positive, the maximum value of the objective function does not exceed (1+∈)A π ;

[0062] When the advantage function value A π When it is negative, the value of the objective function is (1-∈)A π and A π The smaller value of

[0063] C2) Value Function

[0064] Define the action-value function Q π (s t , a t ) is used to evaluate the agent in state s according to the strategy π t Perform the given operation a t The degree of goodness is expressed as:

[0065]

[0066] Where γ represents the discount factor, which controls the importance of future returns; V π (s t+1 ) means that under the strategy π, in state s t+1 The state value function under R(s t , s t+1 ) indicates that in the slave state s t Transfer to state s t+1 When you receive instant rewards,

[0067] Based on the Markov decision process, the autonomous agent controlled by reinforcement learning is in state s at time t t When, according to the strategy π(a t |s t ) Select action a t , expressed as π(als)=P(A=a|S=s); the environment receives the action made by the agent and generates a reward r t , and according to the state transition probability function P(s t+1 |s t , a t )Transfer to the next state s t+1 .

[0068] Compared with the prior art, the present invention has the following significant effects:

[0069] 1. For a single-vehicle longitudinal control strategy based on deep reinforcement learning, the DDPG algorithm is used to combine the information of two downstream vehicles into the state space and impose safety constraints on the action space to ensure traffic safety. The design of a multi-objective reward function balances the weights of different aspects such as following distance and efficiency to further optimize the model's longitudinal control performance. It can effectively suppress the propagation of stop-and-go traffic waves, has certain practical significance, and lays the foundation for subsequent multi-agent parameter design. By analyzing the impact of different agent design parameters and reinforcement learning training parameters on algorithm sensitivity, the performance of the agent system is further improved.

[0070] 2. By combining the DDPG algorithm with a parameter-sharing architecture, a multi-agent algorithm, SRM-DDPG, was designed to enhance collaboration among multiple agents and reduce the algorithm's computational complexity. SRM-DDPG improves the state and reward communication methods among multiple agents. First, it introduces a perception radius, broadcasting the state information of agents within the perception radius to the perception source, enabling agents to proactively perceive changes in the driving state of the preceding vehicle, providing a degree of foresight. Second, it introduces a reward decay value that accumulates the decayed reward values ​​of each agent behind the target vehicle, allocating more specific and effective rewards to each agent, effectively addressing the "lazy agent" problem. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] Figure 1 It is the overall flow chart of the present invention;

[0072] Figure 2 It is a multi-vehicle longitudinal control framework based on SRM-MADDPG;

[0073] Figure 3 (a) is a schematic diagram of the perception radius D = 1, (b) is a schematic diagram of the perception radius D = 2, and (c) is a schematic diagram of the perception radius D = 3;

[0074] Figure 4 This is a diagram of the multi-agent reward communication method;

[0075] Figure 5 This is the flow chart of SUMO traffic simulation;

[0076] Figure 6 Flowchart of the reinforcement learning vehicle control experimental platform;

[0077] Figure 7 (a) is the speed curve of each vehicle in the fleet under the SRM-DDPG control strategy, and (b) is the speed curve of each vehicle in the fleet under the IDM baseline control strategy;

[0078] Figure 8(a) shows the acceleration curves of the first and last vehicles in the convoy under the SRM-DDPG control strategy, and (b) shows the acceleration curves of the first and last vehicles in the convoy under the IDM baseline control strategy.

[0079] Figure 9 (a) is the spatiotemporal trajectory of the vehicle under the control of the IDM baseline model, and (b) is the spatiotemporal trajectory of the vehicle under the control of the SRM-DDPG model. DETAILED DESCRIPTION

[0080] The present invention will be described in further detail below with reference to the accompanying drawings and specific implementations.

[0081] like Figure 1 The figure shows the overall process diagram of the present invention. The detailed implementation steps of the present invention are as follows:

[0082] Step 1: Use the open source traffic simulation software SUMO to model the traffic stop-go wave, model each vehicle and design the vehicle's single-lane route. Figure 5 As shown in the figure, according to the operation process of SUMO simulation software, configure Net.xml—road file, Rou.xml—vehicle routing file, and Sumocfg suffix file—simulation configuration file in sequence.

[0083] 11) For the Net.xml road file, use the Netedit program that comes with the SUMO simulation software to directly define and draw the single-lane road segment.

[0084] 12) In the Rou.xml file, the vehicle routing file uses the IDM (Intelligent Driver Model) to control the human-driven vehicle. Furthermore, the lead vehicle is defined in the routing file to create a stop-and-go traffic wave scenario.

[0085] 13) For the Sumocfg suffix file - simulation configuration file, combined with the road file and vehicle routing file, the vehicle can run in a predefined manner on the road, and the simulation time and simulation step are set at the same time. The parameters of the simulation configuration file are simulation start time 0 and end time 3600, that is, a total simulation time of 3600s and a simulation time step of 1s.

[0086] In the second step, SUMO is used to interact with Python, which implements deep reinforcement learning, through the built-in traffic control interface Traci (Traffic Control Interface) to obtain traffic environment information in real time and control the longitudinal behavior of the vehicle.

[0087] 21) The TensorFlow 1.14 deep learning framework was selected, and Python 3.8 was used as the programming language. Furthermore, TensorFlow served as the Python interface between the SUMO simulation software and the deep reinforcement learning algorithm. TensorFlow can implement the determined neural network structure and train a mature control strategy, which can be divided into the following four steps:

[0088] 211) The complex traffic simulation environment is abstracted into representative state feature vectors as the input of the strategy network and the value network. Different longitudinal control strategies extract different state feature vectors.

[0089] 212) Define the structure of the strategy network and value network according to the requirements and define the mapping method from the state feature vector to the action feature vector.

[0090] 213) During the training of the reinforcement learning agent, the parameters of the policy network and value network represented by the TensorFlow method can be adjusted by the training data generated by the traffic environment until convergence and the trained parameter structure is saved.

[0091] 214) By calling the saved neural network parameters, the performance of the trained control strategy can be tested in different scenarios.

[0092] 22) Train the control objects in the external control system. The traffic control interface Traci of the SUMO simulation software allows accessing and retrieving the object data in the running simulation and manipulating the behavior of these control objects online in real time.

[0093] Importing the Traci module in Python allows users to retrieve information about the current state of the target vehicle and issue precise commands to set the speed and position of the vehicle in the next time step. The flow chart of the reinforcement learning vehicle control experimental platform is as follows: Figure 6 shown.

[0094] 221) Since only the vehicle speed can be set in Traci but the vehicle acceleration cannot be set directly, the acceleration can be converted into the instantaneous speed δv = a*dt, where a is the acceleration and dt is the length of each time step in the simulation.

[0095] Step 3: Calibrate and verify the simulation model. Use data from the UTE (Ubiqutious Traffic Eyes) trajectory database to calibrate the parameters of the IDM model. The detailed steps are as follows:

[0096] 31) Extract the trajectory data of the following vehicle, perform data cleaning and preprocessing to ensure data quality and consistency. After processing the data set, it needs to be further screened:

[0097] 311) The selected target vehicle always follows the preceding vehicle during the moving process, and the target vehicle speed is not less than 10 km / h.

[0098] 312) The front and rear vehicles remain in the same lane, there is no lane changing, and the distance between vehicles is less than 120m.

[0099] 313) The front and rear vehicles are in the same lane and the distance between the front of the vehicles is less than 120m.

[0100] 32) Define the fitness function to evaluate the degree of IDM model parameter fitting and prediction accuracy, and use the error between the model prediction and the actual observation trajectory as the evaluation indicator. The formula is as follows:

[0101]

[0102] Where n represents the total number of simulation steps, and They represent the actual headway and simulated headway of the following vehicle at the kth step respectively.

[0103] 33) Identify the IDM model parameters that need to be adjusted and define the range of each parameter.

[0104] 34) Set the parameters of the genetic algorithm, including population size, number of iterations, crossover rate and mutation rate, convergence tolerance, etc.

[0105] 35) Use the UTE dataset to calibrate the IDM model.

[0106] Step 4: Propose and describe the multi-vehicle longitudinal control task, and improve the multi-agent reinforcement learning algorithm based on the task.

[0107] like Figure 2 As shown in the figure, a multi-vehicle longitudinal control framework based on SRM-MADDPG is combined with the parameter sharing structure and the MADDPG algorithm (Multi-Agent Deep Deterministic Policy Gradient, multi-agent reinforcement learning algorithm), making it applicable to multi-vehicle continuous control scenarios. The introduction of the sensing radius D and the reward decay value innovatively improves the state communication and reward communication methods among multi-agents to form SRM-MADDPG (Sensing radius-Reward attenuation-Multi agent MADDPG). Its working process is described in detail with pseudo code, see Table 1.

[0108] Table 1 Detailed implementation process of the SRM-MADDPG algorithm

[0109]

[0110] 41) Based on the SRM-MADDPG algorithm, the local joint state space, action space and attenuated joint reward of multiple agents are designed.

[0111] 411) Definition of multi-vehicle local joint state space

[0112] Based on V2X communication technology and information sharing between vehicles, the communication protocol of the leading vehicle following model widely used in convoy following is improved. The state information of the intelligent agent within the perception radius D is broadcast to the perception source intelligent agent. The perception radius is as follows: Figure 3 As shown in (a), (b), and (c) in Figure 2. For agent n, the local joint state space with a perception radius of D is expressed as:

[0113]

[0114] in, represents the state information of the nth agent at time t, Represents the state information of the n-1th agent at time t, Represents the state information of the nDth agent at time t. The state information of each agent contains 4 variables. Taking the nth agent as an example, its state information include:

[0115]

[0116] in, represents the speed of the nth agent at time t, represents the acceleration of the nth agent at time t, represents the distance between the nth agent and the preceding vehicle at time t. Specifically, when the number of agents within the perception radius is less than D, the state information at positions n-1 to nD that are 0 is input as the state information of the preceding vehicle, while the state information at positions less than 0 is recorded as 0.

[0117] 412) Multi-vehicle action space definition

[0118] Each agent controls the vehicle's acceleration to -3m / s 2 to 3m / s 2 To prevent vehicle collisions from interrupting training when the multi-agent system randomly selects actions to control the vehicle platoon before a mature strategy is learned, which would hinder the agents from learning the optimal strategy, safety constraints are imposed on each agent in the multi-agent system. The action space of the multi-agent system can be expressed as:

[0119]

[0120] wherein, denotes the acceleration of agent n at time t, j = 1, 2,..., n.

[0121] 413) Multi-vehicle decay joint reward definition

[0122] For a vehicle group composed of multiple vehicles, whether the reward obtained by each agent in the system is reasonable directly affects whether the multi-agent system trained can cooperate to achieve the expected purpose. The actions of the front vehicle will have a series of effects on the following vehicle, therefore, a multi-vehicle joint reward definition method is proposed to strengthen the connection and communication among vehicles in the vehicle group. The reward communication mode between vehicles is as shown in Figure 4 . The decay joint reward definition method is as follows:

[0123]

[0124] wherein, R n denotes the decay joint reward obtained by agent n, N refers to the number of vehicles in the vehicle group, and θ is the discount factor, which represents that the influence of the front vehicle on the rear vehicle becomes smaller as the front and rear vehicles are farther apart in the vehicle group.

[0125] wherein, r h denotes the reward value obtained by each agent only considering itself, and the behavior reward of each agent in the system is represented by a multi-objective reward function, that is, for agent n, its reward value can be represented as:

[0126] r n = aReward l + bReward v + l(Reward a + Reward f ) (5)

[0127] wherein, Reward l is the penalty value obtained after the vehicle leaves the following mode; Reward v is the reward value of vehicle passing efficiency; Reward a is the negative value of the square of the target vehicle acceleration; Reward f is the negative value of the fuel consumption of the intelligent connected vehicle; and a, b and l are the weights of different components of the reward function. After careful parameter adjustment, the following values are used: a = 2, b = 1, and l = 2.

[0128] 42) Setting reinforcement learning parameters

[0129] 421) In SRM-MADDPG, the linear attenuation learning rate method is used to reduce the learning rate during the training process, so that the learning rate drops from the set initial value to 0, which can enhance the stability of the later stage of training to a certain extent and improve the convergence level and convergence effect.

[0130] 422) Apply the tanh activation function to the SRM-MADDPG algorithm.

[0131] 423) Disable the experience replay mechanism during SRM-MADDPG algorithm training.

[0132] Step 5: Set the strategy objective function and value function.

[0133] Strategy objective function:

[0134] The MADDPG approximate gradient formula is:

[0135]

[0136] in, represents the policy gradient of agent i; θ i represents the policy network parameters of agent i; J i (θ i ) represents the strategic performance of agent i; E represents the expected operation of the behavioral strategy of agent i; Represents the action a for agent i i In state s i The action-value function Q under i The gradient of i represents the input of the environment state observed by agent i.

[0137] In order to avoid drastic changes in policy updates, MADDPG is clipped to limit the policy updates to a reasonable range, as shown in the following formula:

[0138]

[0139] Among them, ∈ is a hyperparameter that usually takes 0.1 or 0.2; f is a function that represents the i (θ i ) to perform cropping operations, ensuring that J i (θ i ) will not exceed the specified range [1-∈,1+∈].

[0140] By constraining the policy update and clipping operation clip(·), the policy optimization objective function shown in Equation (7) can be rewritten as:

[0141] L CLIP (π θ )=E[min(Ji (θ i )A π ,clip(J i (θ i )),1-∈,1+∈)A π ](8)

[0142] Among them, L CLIP (π θ ) represents the policy loss function calculated using the clipping operation in the MADDPG algorithm; π A represents the policy network.

[0143] When the advantage function value A π When it is positive, the maximum value of the objective function does not exceed (1+∈)A π ; When the advantage function value A π When it is negative, the value of the objective function is (1-∈)A π and A π The smaller value of

[0144] Value function: action value function Q π (s t ,a t ) is proposed to evaluate the agent in state s according to the strategy π t Perform the given operation a t The degree of goodness can be expressed as:

[0145]

[0146] Where γ represents the discount factor, which controls the importance of future returns; V π (s t+1 ) means that under the strategy π, in state s t+1 The state value function under R(s t ,s t+1 ) indicates that in the slave state s t Transfer to state s t+1 Receive instant rewards.

[0147] Based on the Markov decision process, the autonomous agent controlled by reinforcement learning is in state s at time t t According to the strategy π(a t |s t ) Select action a t , which can be expressed as π(a|s)=P(A=a|S=s). The environment receives the action made by the agent and generates a reward r t , and according to the state transition probability function P(s t+1 |s t ,a t )Transfer to the next state st+1 .

[0148] Step six, the speed and acceleration of each vehicle under the control of the SUMO-based IDM baseline model and the SRM-MADDPG multi-vehicle longitudinal control model are compared in detail, respectively, to obtain the superiority of the SRM-MADDPG multi-vehicle longitudinal control model compared with the IDM baseline model.

[0149] In the SRM-MADDPG algorithm with the experience replay mechanism disabled, each agent uses two neural networks, the actor network for generating the policy and the critic network for improving the policy. One of the two critic networks is used to calculate the policy advantage function, and the other is used to clip the overestimated action value. Two actor networks are used to generate the new and old policies before and after parameter update, respectively. The two critic networks are three-layer fully connected neural networks, and the middle layer has 100 neurons. The input layer consists of 10 neurons to receive the state vector, and the output layer has only one neuron, whose output is the Q value of the input state. The two actor networks are two three-layer fully connected neural networks, and the middle layer has 100 neurons. The input layer of the two actor networks consists of 10 neurons to receive the state vector, and the output of the output layer of one neuron is the average value and standard deviation of the action distribution, respectively.

[0150] The reinforcement learning hyperparameters are shown in Table 2. The intelligent connected vehicle fleet in this embodiment consists of 8 vehicles. It should be noted that CAV (Connected and Automated Vehicle) represents an intelligent connected vehicle controlled by the multi-agent algorithm SRM-MADDPG, and HDV (Human Driver Vehicle) represents a human-driven vehicle. The leading vehicle that generates the traffic stop-and-go wave starts to decelerate at the 50th second and experiences a deceleration-then-acceleration process lasting 34 seconds.

[0151] Table 2 Reinforcement learning hyperparameters

[0152]

[0153] 51) Speed comparison. The speed comparison of the two control modes is as follows Figure 7As shown in Figures (a) and (b) of the platoon. The IDM model controls each vehicle in the platoon to closely follow the preceding vehicle from the outset and maintain a nearly identical speed. However, with the SRM-MADDPG control approach, each vehicle in the platoon dynamically adjusts its speed based on the status of the preceding vehicles, thanks to inter-vehicle communication. Therefore, to minimize the propagation of stop-and-go waves within the platoon, each vehicle slightly slows down from the outset and gradually distances itself from the preceding vehicle, preemptively preparing for potential stop-and-go waves. This provides a foresight advantage that the IDM model lacks. More importantly, while each vehicle under SRM-MADDPG control experiences a slight deceleration in advance, the magnitude of the deceleration is minimal, and each vehicle is increasingly less affected by the stop-and-go waves. This is most intuitively reflected in the increasing minimum speed of each vehicle.

[0154] 52) Acceleration comparison. The acceleration of the first and last vehicles in the convoy under the two control modes is as follows: Figure 8 As shown in Figures (a) and (b) of the platoon, the acceleration curve of the vehicle at the front of the platoon is similar to that of the leading vehicle, and the absolute value of the maximum acceleration is also very close to the 1.2m / s of the leading vehicle. 2 The acceleration and deceleration tendency is more obvious. The further back the vehicle is in the convoy, the less obvious the emergency acceleration and deceleration trend of the acceleration curve is, and the absolute value of the maximum acceleration also shows a very obvious downward trend. Especially starting from the fifth vehicle in the convoy, the acceleration of each vehicle is basically concentrated at -0.5m / s 2 -0.5m / s 2 There is no tendency for emergency acceleration or deceleration, that is, it is almost unaffected by traffic stop-and-go waves.

[0155] 53) Finally, in order to more intuitively show the superiority of the SRM-MADDPG model-controlled fleet compared to the IDM baseline model in eliminating traffic stop-go waves, the time-space diagram is drawn as follows Figure 9 This is shown in Figures (a) and (b) of the platoon. This figure shows the changes in position and speed of the vehicles as they move. Vehicles are black at low speeds and light gray at high speeds. As can be seen from the figure, the vehicles controlled by the IDM baseline model are unable to eliminate stop-and-go waves, as evidenced by the black portion of the figure showing no signs of dissipation. The platoon controlled by SRM-MADDPG, on the other hand, virtually suppresses the propagation of stop-and-go waves within the platoon, as evidenced by the purple-black portion of the figure spreading only at the front of the platoon and not throughout the entire platoon, demonstrating excellent control performance.

Claims

1. A multi-vehicle joint longitudinal control method based on improved MADDPG, characterized in that: The steps are as follows: S1, the simulation software SUMO was used to model the traffic stop-go wave, model each vehicle and design the vehicle's single-lane travel route; S2 uses the built-in traffic control interface Traci in the simulation software SUMO to interact with Python for deep reinforcement learning to obtain real-time traffic environment information and control the longitudinal behavior of the vehicle; S3, calibrate and verify the simulation model, using the data in the UTE trajectory database to calibrate the parameters of the IDM model; S4, based on the multi-vehicle longitudinal control task, designs the joint state space, action space and joint reward function of multiple agents; And introduce the perception radius D and reward decay value, construct the SRM-MADDPG algorithm to improve the state communication and reward communication methods among multiple agents; S5, set the policy objective function and value function; In step S3, the detailed steps for calibrating the parameters of the IDM model using the data in the UTE trajectory database are as follows: S31, extracting the trajectory data of the following vehicle, performing data cleaning and preprocessing to ensure data quality and consistency; S32 defines the fitness function for evaluating the degree of IDM model parameter fitting and prediction accuracy, using the error between the model prediction and the actual observed trajectory as the evaluation indicator. The formula is as follows: , Where n represents the total number of simulation steps, 、 They represent the actual headway and simulated headway of the following vehicle at the kth step respectively; S33, determining the IDM model parameters that need to be adjusted and defining the range of each parameter; S34, setting the parameters of the genetic algorithm, including population size, number of iterations, crossover rate and mutation rate, and convergence tolerance; S35, calibrate the IDM model using the UTE dataset; In step S4, the design of the local joint state space, action space and decaying joint reward function of the multi-agent is as follows: B1) Multi-vehicle local joint state space The state information of the agents within the perception radius D is broadcast to the perception source agent. For agent n, the local joint state space with a perception radius of D is expressed as: , in, represents the state information of the nth agent at time t, Represents the state information of the n-1th agent at time t, Represents the state information of the nDth agent at time t; The state information of each agent contains 4 variables. The state information of the nth agent include: , in, represents the speed of the nth agent at time t, represents the acceleration of the nth agent at time t, represents the distance between the nth agent and the preceding vehicle at time t; When the number of agents within the perception radius is less than D, the state information at positions 0 from n-1 to nD is input as the state information of the leading vehicle, and the state information at positions less than 0 is recorded as 0; B2) Multi-car action space Control the vehicle acceleration to -3m / s for each agent 2 to 3 m / s 2 An arbitrary continuous variable between ; imposes safety constraints on each agent in the multi-agent system: , in, is the minimum safe distance between two vehicles, The duration of the strategy; The action space of the multi-agent is represented as: , in, represents the acceleration taken by agent n at time t, j=1,2,…,n; B3) Multi-vehicle joint reward The decaying joint reward is expressed as: , in, represents the decaying joint reward obtained by agent n, N is the number of vehicles in the fleet, is the discount factor; Indicates that each agent only considers the reward value obtained by itself. The reward value for agent n is expressed as: , in, is the penalty value obtained after the vehicle leaves the following mode; Vehicle traffic efficiency reward value; It is the negative value of the square of the target vehicle's acceleration; is the negative value of the fuel consumption of intelligent connected vehicles, 、 and are the weights of the different components of the reward function.

2. The multi-vehicle joint longitudinal control method based on the improved MADDPG according to claim 1 is characterized in that: In step S2, the TensorFlow deep learning framework is selected and a mature control strategy is trained. The implementation steps are as follows: S21, abstracts the complex traffic simulation environment into a representative state feature vector as the input of the policy network and the value network. Different longitudinal control strategies extract different state feature vectors; S22, define the structure of the policy network and value network according to the requirements, and define the mapping method from the state feature vector to the action feature vector; S23, during the training of the reinforcement learning agent, the parameters of the policy network and value network represented by the TensorFlow method are adjusted through the training data generated by the traffic environment until the trained parameter structure is converged and saved; S24, by calling the saved neural network parameters and,testing the performance of the trained control strategy in different scenarios.

3. The multi-vehicle joint longitudinal control method based on the improved MADDPG according to claim 1 is characterized in that: In step S31, the pre-processed data needs to be screened. The screening principles are as follows: A1) The selected target vehicle always follows the preceding vehicle and its speed is no less than 10 km / h; A2) The front and rear vehicles remain in the same lane, without changing lanes, and the distance between them is less than 120m; A3) The front and rear vehicles are in the same lane and the distance between their heads is less than 120m.

4. The multi-vehicle joint longitudinal control method based on the improved MADDPG according to claim 1 is characterized in that: In step S5, the policy objective function and value function are set as follows: C1) Strategy objective function The MADDPG approximate gradient formula is: , in, represents the policy gradient of agent i; represents the policy network parameters of agent i; represents the strategic performance of agent i; E represents the expected operation of the behavioral strategy of agent i; Represents the action a for agent i i In state s i The action-value function Q under i The gradient of i represents the input of the environment state observed by agent i; Clipping MADDPG limits the policy update to a reasonable range, as shown in the following formula: , in, is a hyperparameter; Is a function that represents Perform the cropping operation to ensure Will not exceed the specified range ; Update and prune operations via constraint strategies , then the strategy optimization objective function is rewritten as: , in, represents the policy loss function calculated using the clipping operation in the MADDPG algorithm; represents the policy network; When the advantage function value When it is positive, the maximum value of the objective function does not exceed ; When the advantage function value When it is negative, the objective function takes the value and The smaller value of C2) Value Function Defining the action-value function Used to evaluate the agent's strategy In state Perform the given operation The degree of goodness is expressed as: , in, represents the discount factor, controlling for the importance of future returns; Indicates that under the strategy π, in state The state value function under ; Indicates that the slave state Transfer to state When you are in the game, you will get instant rewards; Based on the Markov decision process, the autonomous agent controlled by reinforcement learning is in state When, according to the strategy Select Action , expressed as ; The environment receives the action made by the agent and generates rewards , and according to the state transition probability function Transition to the next state .

Citation Information

Patent Citations

  • Mixed traffic flow control method based on deep reinforcement learning, medium and equipment

    CN115100850A

  • Multi-agent variable speed limit control method based on behavior trend clustering and feature mapping

    CN116150639A

  • Automatic driving behavior integrated decision-making method based on deep reinforcement learning

    CN115320640A

  • Single-point traffic signal control method based on intersection holographic data

    CN115691167A