Multi-agent-based multi-lane ramp confluence area vehicle control method and system

By introducing the Actor network of quadratic neurons into the multi-agent depth deterministic strategic gradient MADDPG algorithm, a multi-agent depth strategic gradient BQ-MADDPG algorithm based on quadratic neurons was constructed, which solved the problem of vehicle inlet decision control in the multi-lane ramp confluence area scenario, and achieved a more efficient and stable vehicle control effect.

CN119975359APending Publication Date: 2025-05-13CHANGAN UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510326979.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing multi-agent deep reinforcement learning algorithm is difficult to reach a stable state during the training process, or the training progresses slowly, resulting in difficulty in converging the model and unable to efficiently and stably solve the problem of vehicle inlet decision-making control in the multi-lane ramp confluence area scenario under mixed traffic.

Method used

Using the multi-agent depth deterministic strategic gradient MADDPG algorithm and the Actor network based on the quadratic neuron, a multi-agent depth strategic gradient BQ-MADDPG algorithm based on the quadratic neuron is constructed to control the motion state of the intelligent connected vehicle CAV in the multi-lane ramp confluence area until it is completely departed.

Benefits of technology

The nonlinear expression capability, strategy optimization efficiency, ability to adapt to complex interactive scenarios and generalization capabilities of the multi-agent deep strategic gradient BQ-MADDPG algorithm of secondary neurons has been significantly improved, and the decision-making accuracy and adaptability of intelligent connected vehicle CAV in complex traffic scenarios has been improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119975359A_ABST
    Figure CN119975359A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-agent-based vehicle control method and system for a multi-lane ramp confluence area. The vehicle control method comprises the following steps: constructing a simulation system of the multi-lane mixed traffic flow ramp confluence area; the simulation system based on the multi-lane mixed traffic flow ramp confluence area obtains the motion state of the intelligent agent; based on a multi-agent depth deterministic strategy gradient MADDPG algorithm, adjusting the motion state of the agents in the mixed traffic ramp confluence area; the method comprises the following steps of: constructing a multi-agent depth strategic gradient BQ-MADDPG algorithm based on a secondary neuron on the basis of an Actor network of the secondary neuron; in a multi-agent deep strategic gradient BQ-MADDPG algorithm based on secondary neurons, the motion state of an agent in a multi-lane ramp confluence area is adjusted until the agent completely drives away from the multi-lane ramp confluence area, and the method is used for solving the problems that an existing multi-agent deep reinforcement learning MADRL algorithm is difficult to reach a stable state in training, so that model convergence is difficult, and the efficiency is high. The problem of vehicle confluence decision control in a multi-lane ramp confluence area scene under mixed traffic cannot be efficiently solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent transportation technology, and relates to a vehicle control method and system for a multi-lane ramp merging area based on multiple intelligent agents. Background Art

[0002] Autonomous driving vehicles that integrate intelligence and networking are an important development direction for future vehicles. Intelligent transportation systems (ITS) improve traffic safety and comfort. Intelligent connected vehicles (CAVs) are an important part of intelligent transportation systems (ITS). By integrating artificial intelligence and automation technologies, intelligent connected vehicles (CAVs) are expected to improve road safety, alleviate traffic congestion, and fundamentally change traditional travel modes and traffic management models. However, it is still a huge challenge to develop reliable control strategies for intelligent connected vehicles (CAVs) to cope with the complexity of actual driving, especially in mixed traffic environments where intelligent connected vehicles (CAVs) and human-driven vehicles (HDVs) coexist for a long time. Intelligent connected vehicles (CAVs) not only need to respond to road objects, but also need to pay attention to the behavior of human-driven vehicles (HDVs). Among many challenging driving scenarios, the entrance ramp merging problem is one of the most difficult tasks. Nearly 300,000 accidents occur in merging areas every year, with nearly 50,000 deaths. In order to further improve traffic efficiency and safety, the decision-making control of vehicles in the merging process has become a research hotspot in recent years.

[0003] The multi-agent deep reinforcement learning (MADRL) method has shown great potential for the problem of vehicle cooperative control in ramp merging area scenarios with highly complex and highly dynamic solution spaces. By continuously learning and adapting to various environmental inputs, uncertainties and interferences during the training process, effective management of agents can be achieved. However, the multi-agent deep reinforcement learning (MADRL) algorithm often has difficulty reaching a stable state during the training process, or the training progress is slow, resulting in model convergence difficulties and inability to efficiently and stably solve the vehicle merging decision control problem in multi-lane ramp merging area scenarios under mixed traffic. Summary of the invention

[0004] In view of the shortcomings of the prior art, the present invention aims to provide a vehicle control method and system for a multi-lane ramp merging area based on a multi-agent, and input the intelligent connected vehicle CAV as an agent into the multi-agent deep deterministic policy gradient MADDPG algorithm. Based on the Actor network of quadratic neurons, a multi-agent deep policy gradient BQ-MADDPG algorithm based on quadratic neurons is proposed to control the movement state of each agent in the multi-lane ramp merging area until each agent completely leaves the multi-lane ramp merging area.

[0005] In order to achieve the above object, the present invention adopts the following technical solution: The present invention provides a vehicle control method for a multi-lane ramp merging area based on a multi-agent, comprising the following steps: constructing a simulation system for a multi-lane mixed traffic flow ramp merging area; obtaining the motion state of the agent based on the simulation system for the multi-lane mixed traffic flow ramp merging area; adjusting the motion state of the agent in the mixed traffic ramp merging area based on a multi-agent deep deterministic policy gradient MADDPG algorithm; constructing a multi-agent deep strategic gradient BQ-MADDPG algorithm based on a quadratic neuron based on an Actor network; adjusting the motion state of the agent in the multi-lane ramp merging area in the multi-agent deep strategic gradient BQ-MADDPG algorithm based on a quadratic neuron until the agent completely leaves the multi-lane ramp merging area.

[0006] Furthermore, the motion state of the agent includes: the state space corresponding to each agent , Action Space And the reward function .

[0007] Furthermore, the state space corresponding to each agent Includes: the current vehicle and the four vehicles closest to the current vehicle within the perception range;

[0008] in, is the current vehicle status information, and is the status information of the two vehicles closest to the current vehicle. and It is the status information of the two vehicles behind the current vehicle.

[0009] Furthermore, the action space corresponding to each agent include:

[0010] in, is the current vehicle's longitudinal acceleration, -4.5 3.5 ; is the current vehicle's lateral acceleration, -1.2 1.2 .

[0011] Furthermore, the reward function corresponding to each agent is include:

[0012] in, , , , , is a constant term, is the collision reward function of the vehicle, is the vehicle's speed reward function, is the vehicle’s comfort reward function, is the vehicle’s merging distance reward function, is the vehicle's choreographed reward function.

[0013] Furthermore, the multi-agent deep deterministic policy gradient MADDPG algorithm adjusts the movement state of the agent in the mixed traffic ramp merging area, including: each agent adjusts the movement state in the mixed traffic ramp merging area according to the Actor network and the Critic network; the Actor network includes an Actor training network and an Actor target network, and the Critic network includes a Critic training network and a Critic target network.

[0014] Furthermore, when the critic training network is updated, the objective function is first calculated. , using the objective function Calculating the loss function , and then use gradient descent to update the parameters of the agent Critic training network.

[0015] Furthermore, when the Actor training network is updated, the loss function is first calculated , and then use gradient ascent to update the parameters of the agent Actor training network.

[0016] Furthermore, the multi-lane mixed traffic flow ramp merging area includes a two-lane main road and a single-lane ramp.

[0017] The present invention also provides a multi-lane ramp merging area vehicle control system based on multi-agents, including: a simulation module: used to construct a simulation system for a multi-lane mixed traffic flow ramp merging area; an acquisition module: used to obtain the motion state of the agent based on the simulation system of the multi-lane mixed traffic flow ramp merging area; a first execution module: used to adjust the motion state of the agent in the mixed traffic ramp merging area based on the multi-agent deep deterministic policy gradient MADDPG algorithm; a construction module: used to construct a multi-agent deep strategic gradient BQ-MADDPG algorithm based on secondary neurons based on an Actor network; a second execution module: used in the multi-agent deep strategic gradient BQ-MADDPG algorithm based on secondary neurons to adjust the motion state of the agent in the multi-lane ramp merging area until the agent completely leaves the multi-lane ramp merging area.

[0018] Compared with the prior art, the present invention has the following beneficial technical effects: The present invention is based on a multi-agent based vehicle control method for a multi-lane ramp merging area. In the multi-agent deep strategic gradient BQ-MADDPG algorithm of secondary neurons, the intelligent connected vehicle CAV is controlled as an agent, and secondary neurons are added to the Actor network. This can significantly improve the nonlinear expression ability, strategy optimization efficiency, ability to adapt to complex interactive scenarios, and generalization ability of the multi-agent deep strategic gradient BQ-MADDPG algorithm of secondary neurons.

[0019] The present invention is based on a multi-agent vehicle control method for a multi-lane ramp merging area. In a ramp merging area scenario under multi-lane mixed traffic, the decision-making control problem of the intelligent connected vehicle CAV is modeled as a Markov process, and the corresponding state space, action space and reward function are designed. This can significantly improve the decision-making accuracy and adaptability of the intelligent connected vehicle CAV in complex traffic scenarios, optimize traffic flow and enhance the level of intelligence. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 It is a flow chart of the vehicle control method of a multi-lane ramp merging area based on multi-agent of the present invention; Figure 2 Schematic diagram of a ramp merging area for multi-lane mixed traffic flow in an embodiment of the present invention; Figure 3 A scene diagram of adjusting an intelligent agent at an entrance ramp of a mixed traffic ramp merging area in an embodiment of the present invention; Figure 4 A schematic diagram of adjusting the motion state of an intelligent body in a mixed traffic ramp merging area in an embodiment of the present invention; Figure 5 This is a diagram of the Actor network architecture based on secondary neurons in an embodiment of the present invention. DETAILED DESCRIPTION

[0021] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0022] Example 1 The present invention provides a multi-lane ramp merging area vehicle control method based on multi-agent, comprising the following steps: constructing a simulation system for a multi-lane mixed traffic flow ramp merging area; obtaining the motion state of the agent based on the simulation system for the multi-lane mixed traffic flow ramp merging area; adjusting the motion state of the agent in the mixed traffic ramp merging area based on a multi-agent deep deterministic policy gradient MADDPG algorithm; constructing a multi-agent deep strategic gradient BQ-MADDPG algorithm based on a quadratic neuron Actor network; adjusting the motion state of the agent in the multi-lane ramp merging area in the multi-agent deep strategic gradient BQ-MADDPG algorithm based on a quadratic neuron until the agent completely leaves the multi-lane ramp merging area, such as Figure 1 In this implementation, the intelligent agent is a CAV.

[0023] Construct a simulation system for the merging area of ​​a multi-lane mixed traffic flow ramp, specifically: Figure 2 As shown, a traffic scene consisting of a two-lane main road and a single-lane ramp is constructed. Compared with the tapered ramp merging scene that only needs to consider longitudinal control, the parallel ramp merging scene is more common in the real world. The present invention has an acceleration lane, which requires comprehensive consideration of the vehicle's lateral and longitudinal control, and is more complex.

[0024] The simulation system based on the ramp merging area of ​​multi-lane mixed traffic flow obtains the motion state of the intelligent agent, which includes the state space corresponding to each intelligent agent. , Action Space And the reward function ,like Figure 3 shown.

[0025] The state space corresponding to each agent Includes: the current vehicle and the four vehicles closest to the current vehicle within the perception range;

[0026] in, is the current vehicle status information, and is the status information of the two vehicles closest to the current vehicle. and It is the status information of the two vehicles behind the current vehicle.

[0027] The status information of each vehicle includes 6 status features, namely .in Real-time coordinate, Real-time coordinate, is the real-time lateral speed of the vehicle, is the real-time longitudinal velocity of the vehicle, is the real-time lateral acceleration of the vehicle, is the real-time longitudinal acceleration of the vehicle.

[0028] If there is no vehicle at the corresponding position within the radar range of the vehicle, the status information is 0.

[0029] The action space corresponding to each agent is for:

[0030] in, is the current vehicle's longitudinal acceleration, -4.5 3.5 ; is the current vehicle's lateral acceleration, -1.2 1.2 ;It should be noted that all vehicles follow the sports bike model.

[0031] The reward function corresponding to each agent is: for:

[0032] in, , , , , is a constant term, is the collision reward function of the vehicle, is the vehicle's speed reward function, is the vehicle’s comfort reward function, is the vehicle’s merging distance reward function, is the vehicle's choreographed reward function.

[0033] In this embodiment, after a trial and error process, is 1, is 1, is 3, is 0.5, is 2.

[0034] First consider the vehicle's collision reward function from the perspective of safety, comfort and efficiency .

[0035]

[0036] During driving, if the car collides, a penalty value of -100 will be given to guide the vehicle to drive safely. By giving a large negative reward for collision, the agent can be strongly motivated to avoid any behavior that may lead to a collision, thereby ensuring driving safety.

[0037] Vehicle speed reward function for:

[0038] in, is the real-time lateral speed of the vehicle; is the real-time longitudinal speed of the vehicle; The maximum speed of the vehicle is defined as 30 m / s. As the vehicle speed approaches the maximum speed, the penalty value will gradually decrease. This reward function allows the vehicle to learn efficient driving behavior.

[0039] Through the speed reward function , can guide the intelligent agent to maintain a reasonable speed, neither too fast nor too slow, thereby improving road traffic efficiency and driving safety.

[0040] Vehicle comfort reward function for:

[0041] in, is the instantaneous acceleration value of the vehicle, that is, the derivative of acceleration, which is usually used as an indicator of passenger comfort. High acceleration values ​​indicate discomfort, and low acceleration values ​​indicate comfort.

[0042] in, is the time interval of each vehicle control, is the real-time lateral acceleration of the vehicle, is the real-time longitudinal acceleration of the vehicle, is the lateral acceleration of the vehicle at the last moment, is the longitudinal acceleration of the vehicle at the previous moment.

[0043] Through the comfort reward function It can motivate the intelligent agent to maintain a stable driving state, reduce uncomfortable behaviors such as sudden acceleration and braking, and improve the passengers' riding experience.

[0044] Vehicle merging distance reward function for:

[0045] in, is the distance the vehicle travels in the acceleration lane, The length of the acceleration lane.

[0046] When a vehicle is traveling in the acceleration lane, it will receive a penalty value that increases with the distance to guide the vehicle to merge into the main road as quickly as possible.

[0047] Through the merging distance reward function It can guide the intelligent agent to merge smoothly into the main road at the right time and location, reducing traffic congestion and collision risks caused by improper merging.

[0048] The vehicle’s director reward function for:

[0049] in, The real-time lateral acceleration of the vehicle is penalized to reduce unnecessary and frequent lane changes.

[0050] By directing the reward function It can motivate the intelligent agent to perform smooth lane change operations at the right time, avoid unnecessary and sudden lane changes, and thus improve driving safety and road traffic efficiency.

[0051] Based on the multi-agent deep deterministic policy gradient MADDPG algorithm, the movement state of the agent in the mixed traffic ramp merging area is adjusted.

[0052] Specifically, the multi-agent deep deterministic policy gradient MADDPG algorithm adopts centralized training and distributed execution, such as Figure 4 As shown in the figure, each agent has two networks: Actor network and Critic network. The Actor network includes Actor training network and Actor target network, and the Critic network includes Critic training network and Critic target network. The Critic network of each agent can obtain the strategy information of other agents, and all agents share a centralized Critic network, which provides guidance for the Actor network of each agent during the training process.

[0053] In the execution phase, each agent's Actor network makes decisions independently to achieve decentralized execution. Each agent node maintains and updates the local Actor network separately and outputs corresponding actions based on its own observation information.

[0054] When the Critic training network of agent i is updated, the objective function is first calculated :

[0055] in, The reward value corresponding to agent i, is the discount factor; Represents the Critic target network of agent i, giving the next global observation space and all agent actions Q value.

[0056] According to the objective function Calculating the loss function :

[0057] in, is the batch size, which indicates the number of samples drawn from the experience replay pool; Represents the Critic training network of agent i, given in the current global observation space and all agent actions Q value.

[0058] Afterwards, gradient descent is used to update the parameters of the Critic training network of agent i:

[0059] in, The parameters of the Critic training network, is the learning rate of the Critic network; Represents the loss function About parameters gradient.

[0060] When the Actor training network of agent i is updated, the loss function is first calculated :

[0061] in, is the batch size, which indicates the number of samples drawn from the experience replay pool; Represents the Critic training network of agent i; Represents the Actor training network of agent i, outputting agent i in the current global observation space The action below; Parameters of the Actor training network for agent i; is the action of other agents.

[0062] Afterwards, the Actor training network parameters of agent i are updated using gradient ascent:

[0063] in, The parameters of the Actor training network for agent i, is the learning rate of the Actor network, Represents the loss function About parameters gradient.

[0064] Afterwards, the target network of agent i and the Actor target network are soft updated:

[0065] in, is the critic target network parameter of agent i, are the training network parameters of agent i, is the target network parameter of the Actor of agent i, are the training network parameters of agent i, is the soft update coefficient.

[0066] Based on the multi-agent deep deterministic policy gradient MADDPG algorithm, the Actor network based on quadratic neurons obtains additional nonlinear mapping from the quadratic aggregation function, provides more powerful feature representation capabilities, constructs a multi-agent deep policy gradient BQ-MADDPG algorithm based on quadratic neurons, and trains the agents.

[0067] Specifically, a secondary neuron layer is added to the original Actor network, such as Figure 5 As mentioned above, stronger feature extraction capabilities can help the agent learn better strategies faster. The calculation of the secondary neuron layer hidden layer 1 and hidden layer 2 is as follows:

[0068] in, is the input dimension, i.e. the observation space dimension of the environment, and are the trainable weights in the neurons, and is the input vector.

[0069] To allow for more trainable weights, another hidden layer 3 is introduced to allow for a higher number of trainable weights as follows:

[0070] in, is the weight matrix, is the bias vector, is the number of neurons.

[0071]

[0072] in, is the dimension of the input state, that is, the dimension of the agent’s observation space.

[0073] No activation function that would introduce additional nonlinearity is used inside the secondary neuron layer, which makes it possible to directly derive closed-form partial derivatives of any parameter in the network, while reducing the amount of computation and converging faster.

[0074] The multi-agent deep strategy gradient BQ-MADDPG algorithm with quadratic neurons is used for training. During the training process, different numbers of intelligent networked vehicles (CAVs) and human-driven vehicles (HDVs) are randomly generated at different locations on the main road and ramp. The initial speeds of intelligent networked vehicles (CAVs) and human-driven vehicles (HDVs) are randomly selected between 22m / s and 25m / s, the maximum speed is 30m / s, and the longitudinal acceleration is between [-4.5 ,3.5 ] interval, the lateral acceleration is in [-1.2 ,1.2 ] interval, the update time is 100Hz, that is, the simulation step length is 0.1s. The lateral and longitudinal vehicle motion models of the human-driven car HDV are SL2015 and IDM respectively. The motion of the intelligent connected car CAV is controlled by the multi-agent deep policy gradient BQ-MADDPG algorithm of secondary neurons. The training hyper parameters of the multi-agent deep policy gradient BQ-MADDPG algorithm of secondary neurons are shown in Table 1.

[0075]

[0076] Example 2 The present invention discloses a multi-lane ramp merging area vehicle control system based on multi-agents, comprising: a simulation module, an acquisition module, a first execution module, a construction module, and a second execution module.

[0077] Simulation module: used to construct a simulation system for a multi-lane mixed traffic flow ramp merging area; acquisition module: used to obtain the motion state of the agent based on the simulation system of the multi-lane mixed traffic flow ramp merging area; first execution module: used to adjust the motion state of the agent in the mixed traffic ramp merging area based on the multi-agent deep deterministic policy gradient MADDPG algorithm; construction module: used to construct a multi-agent deep strategic gradient BQ-MADDPG algorithm based on quadratic neurons based on the Actor network; second execution module: used in the multi-agent deep strategic gradient BQ-MADDPG algorithm based on quadratic neurons to adjust the motion state of the agent in the multi-lane ramp merging area until the agent completely leaves the multi-lane ramp merging area.

[0078] The multi-agent-based multi-lane ramp merging area vehicle control system provided by the present invention can implement method steps consistent with the above-mentioned method, and will not be described in detail.

[0079] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A multi-agent based vehicle control method for a multi-lane ramp merging area, characterized in that: The following steps are involved: Construct a simulation system for ramp merging areas with multi-lane mixed traffic flows; The simulation system based on the ramp merging area of ​​multi-lane mixed traffic flow obtains the motion state of the intelligent body; Based on the multi-agent deep deterministic policy gradient MADDPG algorithm, the motion state of the agent in the mixed traffic ramp merging area is adjusted; Based on the Actor network of quadratic neurons, a multi-agent deep strategy gradient BQ-MADDPG algorithm based on quadratic neurons is constructed; In the multi-agent deep strategic gradient BQ-MADDPG algorithm based on quadratic neurons, the movement state of the agent in the multi-lane ramp merging area is adjusted until the agent completely leaves the multi-lane ramp merging area.

2. The multi-agent based multi-lane ramp merging area vehicle control method according to claim 1, characterized in that: The motion state of the intelligent agent includes: the state space corresponding to each intelligent agent , Action Space And the reward function .

3. The multi-agent based multi-lane ramp merging area vehicle control method according to claim 2, characterized in that: The state space corresponding to each intelligent agent Includes: the current vehicle and the four vehicles closest to the current vehicle within the perception range; in, is the current vehicle status information, and is the status information of the two vehicles closest to the current vehicle. and It is the status information of the two vehicles behind the current vehicle.

4. The multi-agent based multi-lane ramp merging area vehicle control method according to claim 2, characterized in that: The action space corresponding to each intelligent agent include: in, is the current vehicle's longitudinal acceleration, -4.5 3.5 ; is the current vehicle's lateral acceleration, -1.2 1.2 .

5. The multi-agent based multi-lane ramp merging area vehicle control method according to claim 2, characterized in that: The reward function corresponding to each agent include: in, , , , , is a constant term, is the collision reward function of the vehicle, is the vehicle's speed reward function, is the vehicle’s comfort reward function, is the vehicle’s merging distance reward function, is the vehicle's choreographed reward function.

6. The multi-agent based multi-lane ramp merging area vehicle control method according to claim 1, characterized in that: The method of adjusting the motion state of the agent in the mixed traffic ramp merging area based on the multi-agent deep deterministic policy gradient MADDPG algorithm includes: each agent adjusting the motion state in the mixed traffic ramp merging area according to the Actor network and the Critic network; The Actor network includes an Actor training network and an Actor target network, and the Critic network includes a Critic training network and a Critic target network.

7. The multi-agent based multi-lane ramp merging area vehicle control method according to claim 6, characterized in that: When updating the Critic training network, the objective function is first calculated. , using the objective function Calculating the loss function , and then use gradient descent to update the parameters of the agent Critic training network.

8. The multi-agent based vehicle control method for a multi-lane ramp merging area according to claim 6, characterized in that: When the Actor training network is updated, the loss function is first calculated , and then use gradient ascent to update the parameters of the agent Actor training network.

9. The multi-agent based vehicle control method for a multi-lane ramp merging area according to claim 1, characterized in that: The multi-lane mixed traffic flow ramp merging area includes a two-lane main road and a single-lane ramp.

10. Multi-lane ramp merging area vehicle control system based on multi-agent, including: Simulation module: used to build a simulation system for the merging area of ​​multi-lane mixed traffic flow ramps; Acquisition module: used to acquire the motion state of the intelligent body based on the simulation system of the ramp merging area of ​​multi-lane mixed traffic flow; The first execution module is used to adjust the motion state of the agent in the mixed traffic ramp merging area based on the multi-agent deep deterministic policy gradient MADDPG algorithm; Building module: used for Actor network based on quadratic neurons, building multi-agent deep strategy gradient BQ-MADDPG algorithm based on quadratic neurons; The second execution module is used in the multi-agent deep strategic gradient BQ-MADDPG algorithm based on quadratic neurons to adjust the movement state of the agent in the multi-lane ramp merging area until the agent completely leaves the multi-lane ramp merging area.

Citation Information

Cited By

  • Intelligent vehicle control method based on deep reinforcement learning and driving style

    CN120245996A

  • Intelligent cooperative control method, system and equipment for highway ramps and medium

    CN121171044A

  • Main line multi-lane high-speed ramp confluence area vehicle control method based on deep reinforcement learning

    CN121260032A