Lane Merging Control Method for Vehicle Ramp Entrance Based on Deep Reinforcement Learning
Through the deep reinforcement learning method, the state space and reward function are designed, and the Actor-Critic network is built, which solves the accuracy and multiple factors of ramp merging control in the existing technology, and realizes safe, efficient and comfortable vehicle ramp merging control.
Patent Information
- Application Number
- CN202310178262.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-28
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2043-02-28
AI Technical Summary
The existing machine learning-based vehicle ramp merging control method is difficult to establish an accurate traffic prediction model in complex scenarios, resulting in a decrease in the accuracy of ramp merging control and it is difficult to take into account safety, efficiency and comfort.
Using a method based on deep reinforcement learning, the Actor network and Critic network are built by designing state space, action space and reward functions, the near-end strategy optimization algorithm is used to train the network, and combined with SUMO traffic simulation software to realize vehicle ramp entrance confluence control, considering comfort, efficiency and safety.
It improves the safety and efficiency of vehicle ramp merging, enhances passenger comfort, simplifies control methods, and avoids the problem of state space explosion.
Smart Images

Figure CN116215532B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for merging control at a ramp entrance of a vehicle, and more particularly to a method for merging control at a ramp entrance of a vehicle based on deep reinforcement learning. It belongs to the technical field of deep learning and also belongs to the technical field of artificial intelligence control of vehicles. Background Art
[0002] Ramp merging is one of the main reasons for traffic congestion on urban elevated roads and highways. For vehicles, various factors in the surrounding environment must be comprehensively considered at the ramp merging entrance, waiting for a suitable merging opportunity, and at the same time determining the degree and time of acceleration based on the judgment of the safety distance and the speed of the main road vehicles. Careless operation by the driver is extremely likely to cause traffic accidents, resulting in severe traffic congestion, reducing traffic efficiency, increasing collision risks, increasing travel time, and bringing discomfort to passengers. Even if the vehicle successfully completes the ramp entrance merging, it may not necessarily be globally optimal in the actual process, and it is difficult to balance and ensure safety, efficiency, and comfort. Therefore, the vehicle ramp entrance merging control method has strong practical significance and research value.
[0003] With the continuous development of artificial intelligence, intelligent transportation has received extensive attention from scholars at home and abroad. Developing intelligent transportation is an important means to build a modern comprehensive transportation system of "safe and reliable, convenient and efficient, green and intelligent, open and shared", and it is an active response to the new trends of emerging information technology and Internet development. The existing merging control methods based on machine learning mainly focus on model predictive control. Generally, feature variables need to be extracted from actual traffic data first, and then a traffic flow model is established. Due to the high randomness of the actual traffic conditions, it is difficult to establish an accurate traffic prediction model, so it is difficult to achieve good modeling results in complex scenarios, resulting in a decrease in the accuracy of ramp merging control. Deep reinforcement learning is to continuously interact between the intelligent agent (the vehicle to be merged) and other vehicles in a dynamic and complex environment, and learn the optimal control strategy according to the feedback of environmental information. It can be flexibly applied to the control of driving a vehicle into a ramp merging, so as to solve problems such as road congestion and traffic accidents. Therefore, comprehensively considering the longitudinal and lateral motion control of vehicles, taking into account factors such as safety, comfort, and efficiency, a method and system for merging control at a ramp entrance of a vehicle based on deep reinforcement learning of PPO are provided to achieve high-reliability, high-efficiency, and high-comfort ramp merging, and provide a reference solution for solving the bottleneck technology problems of intelligent transportation autonomous vehicles. Summary of the Invention
[0004] A vehicle ramp - entrance merging control method based on deep reinforcement learning. The purpose of this invention is to overcome the deficiencies of existing ramp - merging schemes. It provides a vehicle ramp - entrance merging control method based on deep reinforcement learning, realizes the merging control of autonomous driving vehicles on ramps, enhances the safety of vehicle driving in the merging area, improves the efficiency of vehicle merging actions, and ensures the comfort of passengers, thereby improving the traffic flow efficiency.
[0005] The technical solution of this invention is realized as follows:
[0006] A vehicle ramp - entrance merging control method based on deep reinforcement learning, characterized by using SUMO to construct a ramp - merging simulation environment and obtain relevant state information, designing a state space, an action space, and a reward function, constructing an Actor network and a Critic network based on the Proximal Policy Optimization algorithm, and iteratively training the network until convergence. Finally, it interacts with SUMO through the TraCI interface to complete the ramp - entrance merging behavior, including the following steps:
[0007] Step 1: Use the SUMO traffic simulation software to build a highway ramp - merging section and obtain the state information of the road environment, the ego - vehicle, and surrounding vehicles. The state information includes vehicle ID, lateral position, longitudinal position, lateral speed, longitudinal speed, and acceleration.
[0008] Step 2: Use the state information of the vehicle to design a state space, an action space, and a reward function:
[0009] (1) The state space S is composed of the continuous states of the ego - vehicle and the vehicles in front of and behind it on the main lane. S = {x e ,x f ,x r}, where x e represents the ego - vehicle state, and x e = [p x ,p y ,sp x ,sp y ,a], where p x ,p y represent the lateral and longitudinal positions of the ego - vehicle respectively, sp x ,sp y represent the lateral and longitudinal speeds of the ego - vehicle respectively, a represents the acceleration of the ego - vehicle, and x f ,x r represent the states of the vehicles in front of and behind on the main lane respectively. x i = [d rd,i ,p x,i ,sp y,i ,a i , i ∈ {f, r}, where d rd,i represents the relative distance between the vehicle in that lane and the ego - vehicle, px,i , sp y,i , a i are the lateral position, longitudinal speed, and acceleration of the vehicle, respectively;
[0010] (2) The action space U consists of the vehicle action set u, where u = {1, 2}. 1 represents immediate ramp merging, and 2 represents pausing merging;
[0011] (3) The reward function R consists of a comfort sub - reward, an efficiency sub - reward, and a safety sub - reward, and the expression is: R = μ c R c (t) + μ e R eff (t) + μ u R unsafety , where R c (t), R eff (t), R unsafety represent the comfort sub - reward, efficiency sub - reward, and safety sub - reward, respectively, and μ c , μ e , μ u are the weights corresponding to the sub - rewards, respectively, and μ c + μ e + μ u = 1, where:
[0012] 1) The comfort sub - reward, with the expression: R c (t) = -α·a x (t) 2 - β·a y (t) 2 , a x , a y are the lateral and longitudinal accelerations, respectively, and α, β are the weights corresponding to the lateral and longitudinal comfort;
[0013] 2) The efficiency sub - reward, with the expression: R eff (t) = w t ·R time (t) + w m ·R merge (t) + w s ·R speed (t), R time (t) is the sub - reward based on time, R merge (t) represents the difference between the lateral position of the vehicle itself and the target lateral position, R speed (t) represents the difference between the longitudinal speed of the vehicle itself and the target longitudinal speed, and is adjusted by w t , w m and w s ;
[0014] 3) Safety sub-reward, the expression is: When a collision occurs, the reward value of the current action in the current state is -100; when the distance is lower than the safe distance d s When the merging action u=1 or the merging action u=2 is selected, the R near_collision Parameter to represent the reward value at this time, R near_collision The expression is:
[0015]
[0016] p y,e represents the longitudinal position of the vehicle, p y,f p y,r Respectively represent the longitudinal positions of the front and rear vehicles in the main lane;
[0017] Step 3: Based on the proximal policy optimization algorithm, construct the Actor network and the Critic network. The Actor network is a policy network responsible for interacting with the environment to obtain action strategies; the Critic network is an evaluation network responsible for evaluating the strategy. The policy network adjusts the parameter θ to output the action, and the evaluation network guides the Actor network to converge in the direction of greater cumulative rewards; initialize the Actor network, the Critic network and the network parameters, train the network until convergence, and establish a vehicle ramp entrance merging control model;
[0018] Step 4: Input the real-time vehicle status information generated by the SUMO simulation software into the trained vehicle ramp entrance merging control algorithm model through the TraCI interface, obtain the corresponding decision-making behavior, and return it to SUMO to execute vehicle merging.
[0019] Compared with the prior art, the advantages of the present invention are obvious, mainly manifested in:
[0020] 1. The system and method of the present invention are no longer limited to the traditional merging action that only considers the influencing factors of the collision rate. Other reward standards such as efficiency and comfort are added, which can effectively improve the traffic efficiency of bottleneck sections and ensure passenger comfort.
[0021] 2. The existing vehicle ramp entrance merging control technology is complex, mainly because the model required to describe the highway traffic flow is complex. The deep reinforcement learning proposed in the present invention improves the control behavior by mining the features of historical data, thus eliminating the need to build a complex traffic model and simplifying the control method.
[0022] 3. When dealing with complex state spaces, existing deep reinforcement learning methods are prone to fall into the dilemma of state space explosion, causing dimensionality disaster; the PPO deep reinforcement learning method using the actor-critic architecture in the present invention can effectively solve this problem. Brief Description of the Drawings
[0023] The present invention has attached Figure 2 drawings.
[0024] Figure 1 is a schematic diagram of ramp merging of the present invention;
[0025] Figure 2 is a flowchart of the vehicle ramp merging control method of the present invention. Detailed Embodiments
[0026] Such as Figure 1 , 2 shown, a vehicle ramp entry merging control method based on deep reinforcement learning, characterized in that SUMO is used to construct a ramp merging simulation environment and obtain relevant state information, a state space, an action space, and a reward function are designed, an Actor network and a Critic network are constructed based on the proximal policy optimization algorithm, and the network is iteratively trained until convergence, and finally, the TraCI interface is used to interact with SUMO to complete the ramp entry merging behavior, including the following steps:
[0027] Step 1: Use the SUMO traffic simulation software to build a highway ramp merging section and obtain the state information of the road environment, the vehicle itself, and the surrounding environmental vehicles. The state information includes vehicle ID, lateral position, longitudinal position, lateral speed, longitudinal speed, and acceleration;
[0028] Step 2: Use the state information of the vehicle to design a state space, an action space, and a reward function:
[0029] (1) The state space S is composed of the continuous states of the vehicle itself and the front and rear vehicles on the main lane. S = {x e , x f , x r}, where x e represents the state of the vehicle itself, and x e = [p x , p y , sp x , sp y , a], where p x , p y respectively represent the lateral and longitudinal positions of the vehicle itself, sp x , sp y respectively represent the lateral and longitudinal speeds of the vehicle itself, a represents the acceleration of the vehicle itself, and x f , x r are the states of the front and rear vehicles on the main lane respectively, and x i = [d rd,i , p x,i , sp y,i , a i , i ∈ {f, r}, where drd,i Indicates the relative distance between the vehicle in this lane and the host vehicle, p x,i , sp y,i , a i are respectively the lateral position, longitudinal speed and acceleration of the vehicle;
[0030] (2) The action space A is composed of the vehicle action set a, a = {1, 2}, where 1 indicates immediate ramp merging and 2 indicates pausing merging;
[0031] (3) The reward function R is composed of a comfort sub - reward, an efficiency sub - reward, and a safety sub - reward. The expression is: R = μ c R c (t) + μ e R eff (t) + μ u R unsafety , where R c (t), R eff (t), R unsafety represent the comfort sub - reward, efficiency sub - reward, and safety sub - reward respectively, and μ c , μ e , μ u are respectively the weights corresponding to the sub - rewards, and μ c + μ e + μ u = 1, where:
[0032] 1) The comfort sub - reward, the expression is: R c (t) = -α·a x (t) 2 -β·a y (t) 2 , a x , a y are respectively the lateral and longitudinal accelerations, and α and β are respectively the weights corresponding to lateral and longitudinal comfort;
[0033] 2) The efficiency sub - reward, the expression is: R eff (t) = w t ·R time (t) + w m ·R merge (t) + w s ·R speed (t), R time (t) is the sub - reward based on time, R merge (t) represents the difference between the lateral position of the host vehicle and the target lateral position, R speed (t) represents the difference between the longitudinal speed of the host vehicle and the target longitudinal speed, and is adjusted by w t , w m and w s ;
[0034] 3) Safety sub - reward, with the expression: When a collision occurs, the reward value of the current action in the current state is - 100; when the distance is lower than the safety distance d s At this time, according to the situation of choosing the merging action u = 1 or the suspended merging action u = 2, the reward value at this time is represented by the R near_collision parameter, and the R near_collision expression is:
[0035]
[0036] p y,e represents the longitudinal position of the ego - vehicle, and p y,f p y,r respectively represent the longitudinal positions of the vehicles in front and behind on the main lane;
[0037] Step 3: Based on the Proximal Policy Optimization algorithm, construct the Actor network and the Critic network. The Actor network is the policy network, responsible for interacting with the environment to obtain the action strategy; the Critic network is the evaluation network, responsible for evaluating the policy. The policy network adjusts the parameter θ to output actions, and the evaluation network guides the Actor network to converge in the direction of a larger cumulative reward; initialize the Actor network, the Critic network, and the network parameters, train the network until convergence, and establish a vehicle ramp - entrance merging control model;
[0038] Step 4: Input the real - time vehicle state information generated by the SUMO simulation software into the trained vehicle ramp - entrance merging control algorithm model through the TraCI interface, obtain the corresponding decision - making behavior, and return to SUMO to execute vehicle merging.
Claims
1. A vehicle ramp entrance merging control method based on deep reinforcement learning, characterized in that: Build a ramp merging simulation environment using SUMO and obtain relevant status information. Design the state space, action space, and reward function. Based on the Proximal Policy Optimization algorithm, construct the Actor network and the Critic network, and iteratively train the networks until convergence. Finally, interact with SUMO through the TraCI interface to complete the ramp entry merging behavior, including the following steps: Step 1: Use the SUMO traffic simulation software to build a highway ramp merging section and obtain the status information of the road environment, the ego vehicle, and the surrounding environment vehicles. The status information includes vehicle ID, lateral position, longitudinal position, lateral speed, longitudinal speed, and acceleration; Step 2: Use the vehicle status information to design the state space, action space, and reward function: (1) The state space S is composed of the continuous states of the ego vehicle and the vehicles in front of and behind it on the main lane, S = {x e , x f , x r}, where x e represents the state of the ego vehicle, and x e = [p x , p y , sp x , sp y , a], where p x , p y represent the lateral and longitudinal positions of the ego vehicle respectively, sp x , sp y represent the lateral and longitudinal speeds of the ego vehicle respectively, a represents the acceleration of the ego vehicle, and x f , x r are the states of the vehicles in front of and behind the main lane respectively. x i = [d rd,i , p x,i , sp y,i , a i , i ∈ {f, r}, where d rd,i represents the relative distance between the vehicle in that lane and the ego vehicle, and p x,i , sp y,i , a i are the lateral position, longitudinal speed, and acceleration of that vehicle respectively; (2) The action space U is composed of the vehicle action set u, where u = {1, 2}. 1 represents immediate ramp merging, and 2 represents pausing the merging; (3) The reward function R is composed of a comfort sub-reward, an efficiency sub-reward, and a safety sub-reward, and its expression is: R = μ c R c (t) + μ e R eff (t) + μ u R unsafety , where R c (t), R eff (t), R unsafety represent the comfort sub-reward, the efficiency sub-reward, and the safety sub-reward respectively, and μ c , μ e , μ u are the weights of the corresponding sub-rewards respectively, and μ c + μ e + μ u = 1, where: 1) Comfort sub - reward, with the expression: R c (t) = -α·a x (t) 2 -β·a y (t) 2 where a x and a y are the transverse and longitudinal accelerations respectively, and α and β are the weights corresponding to the transverse and longitudinal comfort levels respectively; 2) Efficiency sub - reward, with the expression: R eff (t) = w t ·R time (t) + w m ·R merge (t) + w s ·R speed (t), where R time (t) is the sub - reward based on time, and R merge (t) represents the difference between the lateral position of the vehicle itself and the target lateral position, and R speed (t) represents the difference between the longitudinal speed of the vehicle itself and the target longitudinal speed, and is adjusted by w t , w m and w s to adjust the weights; 3) Safety sub - reward, the expression is: When a collision occurs, the reward value of the current action in the current state is - 100; when the distance is lower than the safety distance d s At this time, according to the case of choosing the merging action u = 1 or the pause - merging action u = 2, the reward value at this time is represented by the R near_collision parameter, and the R near_collision expression is: p y,e represents the longitudinal position of the host vehicle, p y,f p y,r respectively represent the longitudinal positions of the vehicles in front of and behind in the main lane; Step 3: Based on the Proximal Policy Optimization algorithm, construct the Actor network and the Critic network. The Actor network is the policy network, responsible for interacting with the environment to obtain the action policy; the Critic network is the evaluation network, responsible for evaluating the policy. The policy network adjusts the parameter θ to output the action, and the evaluation network guides the Actor network to converge in the direction of greater cumulative reward; initialize the Actor network, Critic network, and network parameters, train the network until convergence, and establish a vehicle ramp entry merging control model; Step 4: Input the real-time vehicle status information generated by the SUMO simulation software into the trained vehicle ramp entry merging control algorithm model through the TraCI interface, obtain the corresponding decision behavior, and return to SUMO to execute vehicle merging.
Citation Information
Patent Citations
Autonomous driving rule learning method based on deep reinforcement learning
CN111222630A
Method for improving traffic passing efficiency by utilizing intelligent network connection vehicle
CN112700642A