Fully automatic driving intersection far-leading turn-round path optimization method based on deep reinforcement learning
Through the method based on deep reinforcement learning, the far-lead turn point setting at the intersection is dynamically adjusted, which solves the problem of autonomous driving vehicles selecting the optimal turn path at the intersection, and achieves the effect of reducing queueing and improving traffic efficiency.
Patent Information
- Application Number
- CN202510175746.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-23
AI Technical Summary
The existing intersection design and traffic flow control methods cannot fully utilize the advantages of autonomous driving technology, especially in the setting of remote head turnover at the intersection, it is difficult to optimize the position of the turnover to reduce congestion and improve traffic efficiency.
The optimization method of the far-lead head turnover path of the fully autonomous driving intersection based on deep reinforcement learning is adopted. By training the autonomous driving vehicle to select a suitable turnover path, dynamically adjust the setting of the far-lead head turnover point at the intersection to reduce queueing and improve traffic efficiency.
Effectively reduce the queuing phenomenon at intersections, improve traffic efficiency, optimize traffic flow in a fully autonomous driving environment, and improve the overall operating efficiency of intersections.
Smart Images

Figure CN120024352A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent transportation and path optimization, and in particular to a method for optimizing the path of a fully automatic driving intersection remote U-turn based on deep reinforcement learning. Background Art
[0002] With the rapid development of intelligent network technology and autonomous driving technology, the application prospects of connected autonomous vehicles (CAV) in transportation systems are becoming increasingly broad. With its precise control and real-time decision-making capabilities, autonomous driving technology can significantly improve the efficiency of traffic flow, reduce energy consumption, and enhance traffic safety.
[0003] Autonomous vehicles have the ability to perceive the surrounding environment in real time and make accurate decisions, so they perform much better than manually driven vehicles at intersections. However, most existing intersection designs are still based on the behavioral assumptions of manually driven vehicles and do not fully consider the characteristics of autonomous vehicles. As a result, in a fully autonomous driving environment, traditional intersection design and traffic flow control methods cannot fully utilize the advantages of autonomous driving technology, restricting the realization of its potential.
[0004] Traditional intersection design methods mostly rely on static traffic flow models, which are difficult to cope with dynamic changes in traffic flow and give full play to the characteristics of autonomous vehicles. Especially in the setting of intersection remote U-turns, how to reduce intersection congestion and improve traffic efficiency by optimizing the location of the U-turn has become an important issue that needs to be studied urgently. Summary of the invention
[0005] Purpose of the invention: The purpose of the present invention is to provide a method for optimizing the remote U-turn path at a fully autonomous driving intersection based on deep reinforcement learning, so as to dynamically adjust the remote U-turn point setting at the intersection, effectively reduce the queuing phenomenon at the intersection, and improve the traffic efficiency, thereby optimizing the traffic flow in the fully autonomous driving environment and improving the overall operation efficiency of the intersection.
[0006] Technical solution: A method for optimizing the remote U-turn path at a fully autonomous driving intersection based on deep reinforcement learning. Through the remote U-turn setting, combined with the real-time traffic flow data of the intersection, the deep reinforcement learning algorithm is used to train the autonomous driving vehicle to select a suitable U-turn path; the remote U-turn setting refers to opening the central dividing strip of the main road to prohibit vehicles on the secondary road from going straight and turning left: the straight traffic flow of the secondary road enters the U-turn opening by turning right and then merges into the main road; the left-turning traffic flow of the secondary road also enters the U-turn opening by turning right, then merges into the main road and goes straight through the intersection with the main road traffic flow; the steps are as follows:
[0007] S1, taking each vehicle as an intelligent agent and aiming to select the optimal long-distance U-turn path, an intelligent agent model based on deep reinforcement learning is constructed; in each time step, the number of queued vehicles at the U-turn entrance of each vehicle and the previous action are taken as the state S(t) of time step t, the vehicle's choice of maintaining the current U-turn path or selecting the next U-turn path is taken as the action A(t) of time step t, and the negative number of queued vehicles on the selected U-turn path is taken as the reward R(t) of time step t;
[0008] S2, initialize the learning rate, sample extraction number, experience replay pool capacity, exploration rate, discount factor and training rounds;
[0009] S3, using the observed state as the network input, the agent selects an action based on the current state. After executing the action, the agent obtains an immediate reward R(t) from the environment; the (S(t), A(t), R(t), S(t+1)) generated by each interaction is stored in the experience replay pool for subsequent training; when the number of training times reaches the preset value, the training is completed and the agent model is saved;
[0010] S4, in the new scenario, each vehicle makes real-time decisions by loading the trained agent model.
[0011] Furthermore, the agent selects action A(t) at time step t based on the current state, and calculates the reward R(t) at time step t according to the number of queued vehicles on the path selected by the target vehicle. The agent maximizes the reward by reducing the number of queued vehicles. After the vehicle selects the U-turn path, the traffic flow state changes to a new state S(t+1). According to the reward function value, the U-turn path A(t+1) is reselected, and this cycle is repeated until the optimal U-turn path is finally selected. Each U-turn path corresponds to a different U-turn point distance.
[0012] Furthermore, the state space S(t) is a state vector:
[0013] S(t)={current_QueueLength(t),A(t-1)}
[0014] Among them, current_QueueLength(t) is the number of queued vehicles on the current path at time step t; A(t-1) is the previous action.
[0015] Furthermore, the agent's action space has two options:
[0016] A(t) = 0 means maintaining the current U-turn path;
[0017] A(t) = 1 means selecting the next U-turn path;
[0018] Each path selection corresponds to a specific U-turn point distance, and the specific mapping formula is as follows:
[0019]
[0020] Where D(·) is the distance of the U-turn selected by the target vehicle at state t, A(t) is the action selected by the target vehicle at time step t, and current_edge is the current road section of the target vehicle; E0 It is the entrance to the secondary road. Ei is the road section where the i-th U-turn is located, 1≤i≤7.
[0021] Furthermore, the expression of reward R(t) is as follows:
[0022] R(t)=-next_QueueLength(next_edge(A(t),current_edge))
[0023] Among them, next_QueueLength(·) represents the number of queued vehicles on the selected path, and next_edge(A(t), current_edge) is the next U-turn path selected by the action.
[0024] Compared with the prior art, the present invention has the following significant effects:
[0025] The present invention sets up multiple remote U-turn paths in the central dividing strip of the main road at the intersection for the autonomous driving vehicles to choose. The remote U-turn path optimization model of the intersection is constructed based on the DQN algorithm in deep reinforcement learning, and a unique state, action and reward setting scheme is designed. Among them, the state of each vehicle includes the number of queued vehicles in the lane and the last action selected; the action of each vehicle is designed to choose to keep the current U-turn path and choose the next U-turn path; the reward is designed to achieve the purpose of reducing the average vehicle delay, and the negative number of the number of vehicles queuing on the U-turn path selected by the vehicle is used as the reward, so that the vehicle can choose the U-turn path with fewer queued vehicles to reduce the average vehicle delay, and at the same time solve the problem of the autonomous driving vehicle selecting the optimal remote U-turn path at the intersection. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 is a flow chart of the present invention;
[0027] Figure 2 It is a simulation scene diagram in the embodiment;
[0028] Figure 3 Schematic diagram of the total reward changing with the number of training times in the embodiment. DETAILED DESCRIPTION
[0029] The present invention is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0030] The present invention proposes a method for optimizing the remote U-turn path at a fully automated driving intersection based on deep reinforcement learning. The flowchart is as follows: Figure 1 As shown in the figure, this method combines the real-time traffic flow data of the intersection and uses the deep reinforcement learning algorithm to train the autonomous driving vehicle to select the appropriate U-turn path to adapt to the changing traffic demand. Through this dynamic optimization method, it can effectively reduce the queuing phenomenon at the intersection and improve the traffic efficiency, thereby optimizing the traffic flow in the fully autonomous driving environment and improving the overall operation efficiency of the intersection.
[0031] like Figure 2 As shown in the figure, the remote U-turn setting refers to opening the central dividing strip of the main road to prohibit vehicles on the secondary road from going straight or turning left: the straight traffic on the secondary road enters the U-turn by turning right and then merges into the main road; the left-turning traffic on the secondary road also enters the U-turn by turning right, then merges into the main road and goes straight through the intersection with the main road traffic, thereby effectively reducing traffic conflicts and queues in the intersection. For a fully autonomous driving environment, autonomous vehicles can make efficient and accurate path selection under complex traffic conditions. Therefore, autonomous vehicles can choose the optimal U-turn at the intersection to improve traffic efficiency. The intelligent optimization method of deep reinforcement learning (DRL) is applied, through real-time interaction with the traffic environment, it can self-learn and adjust strategies in the ever-changing traffic flow in order to find the optimal intersection remote U-turn path selection solution.
[0032] In this embodiment, a cross intersection is constructed using the urban traffic simulation software SUMO. Figure 2 As shown in the figure, the main road is 500 meters long in the east-west direction, with a left-turn lane, two straight lanes, and a right-turn lane; the secondary road is 500 meters long in the north-south direction, with a left-turn lane, a straight lane, and a right-turn lane. Seven U-turns are set every 50 meters at the east entrance. An hourly flow rate of 800pcu / h is set at the east entrance of the intersection, and an hourly flow rate of 500pcu / h is set at the south entrance. The left-turn rate is 30%, the vehicle generation time is 300s, and the simulation time is 500s. The CACC following model is used as the autonomous driving following model, and the initial speed of the vehicle is set to 50km / h. The implementation steps of the remote U-turn path optimization method for the autonomous driving intersection are as follows:
[0033] Step 1: Design the overall framework of deep reinforcement learning (DRL);
[0034] The optimization problem of the long-distance U-turn path at the intersection of autonomous driving is described as a deep reinforcement learning problem. Each vehicle is regarded as an intelligent agent. The goal of the intelligent agent is to select the optimal long-distance U-turn path and reduce the average delay of autonomous driving vehicles passing through the intersection. The present invention designs the intelligent agent model based on the Deep Q-Network (DQN) algorithm in deep reinforcement learning. The DQN algorithm combines deep learning with Q-learning. By using a neural network to approximate the Q-value function, the intelligent agent can learn the optimal U-turn path selection strategy through interaction with the environment. The intelligent agent uses this framework to make the best decision in a specific scenario by exploring different actions and observing the environmental feedback obtained. The DQN algorithm uses a deep neural network (DNN) to approximate the value function to improve the efficiency of the DQN algorithm, so that the intelligent agent can learn and optimize the decision-making process more effectively. The intelligent agent learns to select the optimal U-turn path through interaction with the traffic environment. The intelligent agent selects actions according to the current traffic status and learns through rewards from environmental feedback. Specifically, in each time step, the number of queued vehicles at the current U-turn intersection of each vehicle and the previous action are used as the state space S(t) of time step t. The task of the agent is to select the action A(t) of time step t based on the current state, that is, to keep the current U-turn path (action 0) or to select the next U-turn path (action 1). The path selection corresponds to different U-turn point distances (such as 50m, 100m, 150m, etc.). The agent calculates the reward R(t) of time step t based on the number of queued vehicles generated by the U-turn path selected by the target vehicle. The reward is the negative number of queued vehicles on the selected U-turn path. The agent maximizes the reward by reducing the number of queued vehicles, thereby improving the traffic efficiency of the intersection. After the vehicle selects the U-turn path, the traffic flow state changes to the new state S(t+1). According to the reward function value, the action A(t+1) is re-executed (i.e., the U-turn path is selected). This cycle is repeated, and the optimal U-turn path is finally selected.
[0035] The number of queued vehicles at the U-turn entrance of each vehicle and the previous action are taken as the state S(t) at time step t, the vehicle's choice of maintaining the current path or selecting the next path is taken as the action A(t) at time step t, and the negative number of queued vehicles on the selected U-turn path is taken as the reward R(t) at time step t.
[0036] Step 2, state setting;
[0037] The design of the state specifically includes the number of queued vehicles in the lane and the last action A(t-1).
[0038] Then the state space can be represented as a state vector:
[0039] S(t)={current_QueueLength(t),A(t-1)} (1)
[0040] Among them, current_QueueLength(t) is the number of queued vehicles on the current path at time step t; A(t-1) is the previous action.
[0041] Step 3, action setting;
[0042] The action space of the agent consists of two choices: keep the current U-turn path or choose the next U-turn path. The action A(t) of the agent can take values {0,1}. Specifically:
[0043] A(t)=0: keep the current U-turn path;
[0044] A(t)=1: select the next U-turn path;
[0045] Each path selection corresponds to a specific U-turn distance, which is mapped according to the characteristics of the road section. For example, assuming that the current road section is the secondary road entrance of the intersection, the target vehicle can choose a U-turn distance of 0 meters or 50 meters. Similarly, other road sections will be mapped to the corresponding distance according to the set rules. The specific mapping formula of the action is as follows:
[0046]
[0047] Where D(·) is the distance between the U-turn path selected by the target vehicle at state t and the intersection, A(t) is the action selected by the target vehicle at time step t, current_edge is the current section of the target vehicle, E0 is the secondary entrance, and Ei is the section where the i-th (1≤i≤7) U-turn is located.
[0048] Thus, depending on the current road section, a selected action will correspond to a U-turn point distance, and these distances determine the U-turn path selected by the vehicle.
[0049] Step 4, reward setting;
[0050] The agent's reward R(t) is optimized based on the negative of the number of queued vehicles on the U-turn path chosen by the target vehicle. The agent tries to pass the intersection faster by choosing a U-turn path with fewer queued vehicles.
[0051] The reward R(t) is calculated using the following reward function:
[0052] R(t)=-next_QueueLength(next_edge(A(t),current_edge)) (3)
[0053] Among them, next_QueueLength(·) represents the number of queued vehicles on the selected path; next_edge(A(t), current_edge) is the next U-turn path (i.e., the next U-turn intersection) selected by the action, and the distance of this path from the intersection is obtained according to the mapping rule between the current road section and the action A(t) (Formula (2)).
[0054] This reward function encourages the agent to choose a path that can reduce the number of vehicles in the queue, thereby optimizing the traffic efficiency of the intersection.
[0055] Step 5, initialize hyper parameters;
[0056] Hyperparameters include learning rate, sample extraction number, experience replay pool capacity, exploration rate, discount factor, and training rounds. The hyperparameter settings in this implementation are shown in the following table:
[0057] Table 1 Hyperparameter settings
[0058]
[0059] After each interaction with the environment, the agent stores the current state, the action taken, the reward obtained, and the next state in the experience replay pool as a sample.
[0060] Step 6, deep Q network training;
[0061] During training, the network input is the observed state. The agent selects actions based on the current state, using the ε greedy strategy. When making an action decision, a random action is randomly selected with probability ε, and the action that maximizes the Q value is selected with probability 1-ε. The Q value refers to the expected return of selecting an action in a certain state, that is, the cumulative number of negative queued vehicles in all future time steps after selecting action A from state S.
[0062] After executing the action, the agent obtains an immediate reward R(t) from the environment. The (S(t), A(t), R(t), S(t+1)) generated by each interaction is stored as a sample in the experience replay pool for subsequent training.
[0063] When the number of samples in the sample pool reaches the set sample capacity, a batch of training samples are randomly selected from the playback pool for training to enhance the independence and diversity of the samples.
[0064] Calculate the target Q value according to the Bellman equation:
[0065]
[0066] Where y(t) is the target Q value, S(t+1) is the state at time step t+1, γ is the discount factor, which indicates the decay of future rewards, and R(t) is the immediate reward obtained at time step t. is the maximum Q value corresponding to the optimal action in the next state S(t+1); A' refers to the optimal action; θ - Represents the parameters of the target network.
[0067] Loss function calculation, using mean square error (MSE) to calculate the difference between the predicted Q value and the target Q value:
[0068]
[0069] Among them, m is the number of samples drawn; is the target Q value; Q θ (S(t), A(t)) is the predicted Q value, which is the Q value of the optimal action selected by the trained deep Q network based on the current state.
[0070] Update network parameters. Use Adam optimizer to update neural network parameters. The optimizer uses learning rate to adjust parameter updates:
[0071]
[0072] Among them, θ is the parameter of the neural network; α is the learning rate; is the gradient of the loss function with respect to the parameters.
[0073] Step 7, save and apply the deep Q network;
[0074] When the number of training times reaches the preset value, the training is completed and the result is Figure 3 The converged total reward image is shown. The deep Q network will be saved as a file for subsequent new scene applications. After the training is completed, the parameter weights of the deep Q network are obtained. In the new scene, each vehicle can make real-time decisions by loading the corresponding Q network file, so as to achieve the purpose of letting the autonomous driving vehicle choose the optimal long-distance U-turn path.
[0075] In the simulation scenario, the corresponding Q network parameter weight file is loaded for each vehicle to make real-time decisions. In the end, 10 vehicles chose to go directly through the intersection, and 13, 8, 7, 4, 2, 1, and 0 vehicles chose to turn around at 50, 100, 150, 200, 250, 300, and 350 meters away from the intersection, respectively.
[0076] In this embodiment, the average vehicle delay is used as the output indicator to analyze the optimization effect of the U-turn path optimization function. In order to more clearly observe the optimization effect of the present invention, the simulation experiment uses the first-come-first-served scheme at an unsignalized intersection as a control group to compare with the U-turn path optimization model proposed by the present invention. Finally, the average vehicle delay of the first-come-first-served scheme is 47.43 seconds, and the average vehicle delay of the intersection U-turn path optimization scheme is 34.96 seconds. After optimization, the average vehicle delay is reduced by 26%.
Claims
1. A method for optimizing the remote U-turn path at an intersection for fully autonomous driving based on deep reinforcement learning, characterized in that: Through the remote U-turn setting, combined with the real-time traffic flow data of the intersection, the deep reinforcement learning algorithm is used to train the autonomous driving vehicle to select the appropriate U-turn path; the remote U-turn setting refers to opening the central dividing strip of the main road to prohibit the vehicles on the secondary road from going straight and turning left: the straight traffic flow on the secondary road enters the U-turn opening by turning right and then merges into the main road; the left-turning traffic flow on the secondary road also enters the U-turn opening by turning right, then merges into the main road and goes straight through the intersection with the main road traffic flow; The steps include: S1, taking each vehicle as an intelligent agent and aiming to select the optimal long-distance U-turn path, an intelligent agent model based on deep reinforcement learning is constructed; in each time step, the number of queued vehicles at the U-turn entrance of each vehicle and the previous action are taken as the state S(t) of time step t, the vehicle's choice of maintaining the current U-turn path or selecting the next U-turn path is taken as the action A(t) of time step t, and the negative number of queued vehicles on the selected U-turn path is taken as the reward R(t) of time step t; S2, initialize the learning rate, sample extraction number, experience replay pool capacity, exploration rate, discount factor and training rounds; S3, using the observed state as the network input, the agent selects an action based on the current state. After executing the action, the agent obtains an immediate reward R(t) from the environment; the (S(t), A(t), R(t), S(t+1)) generated by each interaction is stored in the experience replay pool for subsequent training; when the number of training times reaches the preset value, the training is completed and the agent model is saved; S4, in the new scenario, each vehicle makes real-time decisions by loading the trained agent model.
2. The method for optimizing the remote U-turn path at an intersection for fully automatic driving based on deep reinforcement learning according to claim 1, characterized in that: The agent selects the action A(t) at time step t based on the current state, and calculates the reward R(t) at time step t according to the number of queued vehicles on the path selected by the target vehicle. The agent maximizes the reward by reducing the number of queued vehicles. After the vehicle selects the U-turn path, the traffic flow state changes to the new state S(t+1). According to the reward function value, the U-turn path A(t) is reselected. t +1), and the cycle is repeated, and finally the optimal U-turn path is selected; wherein each U-turn path corresponds to a different U-turn point distance.
3. The method for optimizing the remote U-turn path at an intersection for fully automatic driving based on deep reinforcement learning according to claim 2 is characterized in that: The state space S(t) is a state vector: S(t)={current_QueueLength(t),A(t-1)} Among them, current_QueueLength(t) is the number of queued vehicles on the current path at time step t; A(t-1) is the previous action.
4. The method for optimizing the remote U-turn path at an intersection for fully automatic driving based on deep reinforcement learning according to claim 2 is characterized in that: There are two options for the agent's action space: A(t) = 0 means maintaining the current U-turn path; A(t) = 1 means selecting the next U-turn path; Each path selection corresponds to a specific U-turn point distance, and the specific mapping formula is as follows: Where D(·) is the distance of the U-turn selected by the target vehicle at state t, A(t) is the action selected by the target vehicle at time step t, and current_edge is the current road section of the target vehicle; E0 It is the entrance to the secondary road. Ei is the road section where the i-th U-turn is located, 1≤i≤7.
5. The method for optimizing the remote U-turn path at an intersection for fully automatic driving based on deep reinforcement learning according to claim 2 is characterized in that: The expression of reward R(t) is as follows: R(t)=-next_QueueLength(next_edge(A(t),current_edge)) Among them, next_QueueLength(·) represents the number of queued vehicles on the selected path, and next_edge(A(t), current_edge) is the next U-turn path selected by the action.