Intersection entrance lane manual driving and automatic driving hybrid fleet queuing configuration optimization method
By applying SAC algorithm to optimize the queuing configuration of hybrid fleets in the intersection environment, the problem of coordinated operation of manual driving and autonomous driving vehicles in the hybrid environment is solved, and the fleet traffic efficiency and safety is improved.
Patent Information
- Application Number
- CN202510175745.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-23
AI Technical Summary
In a mixed travel environment, the uncertainty of artificially driven vehicles limits the optimization of autonomous vehicles at intersections and the improvement of overall traffic efficiency. It is difficult for traditional traffic management methods to take into account the coordinated operation of manual driving and autonomous vehicles, especially in dynamic traffic environments.
The SAC algorithm is used to optimize the queuing configuration of the hybrid fleet. Through the design of the agent and deep reinforcement learning, the different arrangements and combinations of CAV and HV are comprehensively considered, and the queuing configuration of the fleet is optimized to improve the stability and operation efficiency of the fleet through the intersection.
By optimizing the queuing configuration of the hybrid fleet, the synergy between CAV and HV is improved, the fleet's traffic efficiency and safety are improved, traffic congestion and accident rates are reduced, and the efficiency and safety of the traffic system are enhanced.
Smart Images

Figure CN120032502A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent transportation technology, and in particular to a method for optimizing the queue configuration of a mixed fleet of manual driving and automatic driving at an intersection entrance. Background Art
[0002] With the continuous advancement of intelligent network technology and autonomous driving technology, many studies predict that the rapid development of connected autonomous vehicles (CAV) may bring unprecedented changes to the transportation field, including the improvement of traffic efficiency and traffic safety, and the reduction of energy consumption. In urban traffic scenarios, intersections are key nodes of traffic flow, and their efficiency optimization is of great significance to achieving the above goals. Although autonomous driving technology is developing rapidly, since the development of CAV still requires a certain amount of time, human vehicles (HV) will not be completely replaced by CAV in the future, and HV and CAV will coexist in a long transition period. However, since HV does not have the function of network connection, many studies have shown that in mixed traffic environments, the uncertainty of HV driving behavior will significantly limit the potential advantages of CAV in intersection optimization and overall traffic efficiency improvement.
[0003] In order to enable CAV vehicles to fully exert their advantages, many studies have explored the collaborative cooperation between manual driving and automatic driving in mixed traffic environments. For example, some scholars have proposed the concept of a "1+n" mixed convoy, which consists of a leading CAV and n following HVs. The trajectory of the convoy is guided by optimizing the driving trajectory of the leading vehicle, thereby establishing an optimal signal control framework for intersections. Similarly, many scholars have indirectly influenced and controlled HV vehicles by dispatching CAV vehicles at intersections in mixed traffic environments, effectively coordinating the traffic of vehicles at intersections. Existing studies have shown that the collaboration between CAV and HV has a positive effect on improving the traffic efficiency of intersections in mixed traffic environments. However, existing studies mostly focus on scenarios where CAV is the leading vehicle, and are less applicable to more common scenarios, such as random arrangement of CAV and HV, and CAV cannot always be in the leading position. In mixed convoys, there are significant differences in the driving behaviors of manual and automatic driving vehicles, and traditional traffic management methods are difficult to take into account the collaborative operation of the two. In addition, the change of traffic flow on the road is a dynamic process. Traditional intersection management methods are usually based on fixed models and are difficult to cope with the complexity of dynamic environments. For example, linear regression models assume that there is a linear relationship between traffic flow and time or other factors, so it is difficult to capture complex nonlinear dynamic characteristics. Summary of the invention
[0004] Purpose of the invention: The purpose of the present invention is to provide a method for optimizing the queuing configuration of a mixed fleet of manually driven and automatically driven vehicles at an intersection entrance. By utilizing the random exploration ability and dynamic adaptation characteristics of the SAC algorithm and comprehensively considering different arrangements and combinations of CAVs and HVs, the optimal queuing configuration of the mixed fleet is found to improve the stability and operation efficiency of the fleet passing through the intersection.
[0005] Technical solution: A method for optimizing the queuing configuration of a mixed fleet of manual and automatic driving vehicles at an intersection entrance, comprising the following steps:
[0006] S1, construct the overall framework of the SAC model;
[0007] S2, design the agent, set the state S(t), action A(t) and reward R(t);
[0008] S3, using the urban traffic simulation software SUMO to build a simulation environment, the CACC model is used to simulate the car-following behavior of the autonomous driving vehicles in the simulation, and the Krauss car-following model is used to simulate the behavior characteristics of the manually driven vehicles;
[0009] S4, training the SAC model, and after reaching the training stop condition, outputting the optimized Actor network for actually guiding the mixed fleet queuing configuration adjustment in the embodiment.
[0010] Furthermore, in the overall framework of the SAC model, both the Actor network and the Critic network are composed of two layers of fully connected networks, each with 256 neurons in the hidden layer and ReLU as the activation function, and the corresponding results are generated through the output layer, where the Actor outputs the mean and standard deviation of the action, and the Critic outputs the Q value of the state-action pair.
[0011] Furthermore, in step S2, a state representation method based on multidimensional feature extraction is designed to capture the dynamic relationship between vehicles and their impact on traffic efficiency; the maximum fleet size is set to N max , the state of each vehicle contains two key features: vehicle type u and speed v; where vehicle type u = 1 represents a connected autonomous vehicle, and u = 0 represents a manually driven vehicle; vehicle speed v is the instantaneous speed value of the vehicle; by arranging the feature vectors of all team members, the global state vector of the team is constructed:
[0012]
[0013] Among them, S(t) represents the overall state of the mixed fleet at time t; u i represents the vehicle type of the i-th vehicle in the fleet; v i represents the speed of the i-th vehicle in the convoy;
[0014] Design an action setting method, including vehicle speed adjustment and vehicle lane change; the agent outputs the action set A(t) of the entire mixed fleet based on the input represented by the state:
[0015]
[0016] Among them, A(t) represents the set of vehicle actions of the mixed fleet at time t; a i ∈[0,1], represents the action decision value of the i-th vehicle in the fleet;
[0017] For the speed adjustment of the i-th vehicle, set the target speed of vehicle i to v i ', the calculation formula is:
[0018] v i '=a i ·v max
[0019] Among them, v max Indicates the maximum speed limit allowed on the road;
[0020] When the target speed of the i-th vehicle exceeds the speed v of the preceding vehicle i-1 When , vehicle i needs to adjust its position in the convoy by changing lanes and overtaking;
[0021] A reward function is designed based on the average speed and speed variance of the fleet to guide the learning process of the agent. The actions taken by the agent at each time step can get corresponding reward feedback at the next time step, which is used to evaluate and improve the action strategy. At each time step t, the speed set of all vehicles in the fleet is obtained:
[0022]
[0023] Among them, V represents the instantaneous speed value set of all vehicles in the fleet;
[0024] According to the speed set V, the average speed of the fleet is calculated as:
[0025]
[0026] in, is the average speed of all vehicles in the convoy. The larger the value, the higher the traffic efficiency of the intersection.
[0027] Speed variance is used to represent the fluctuation of vehicle speed in a fleet, and its calculation formula is:
[0028]
[0029] Among them, the velocity variance v varIt indicates the discrete degree of the speed of all vehicles in the convoy. The smaller the value, the closer the speed of each vehicle in the convoy is.
[0030] Then the reward function expression is:
[0031]
[0032] Among them, R(t) represents the reward value of the mixed fleet at time t, which is defined as the difference between the average speed of the fleet and the speed variance, where the average speed is used as a positive reward and the speed variance is used as a negative reward; α, β∈[0,1] represent the weight coefficients of the average speed and speed variance in the reward function, respectively.
[0033] Furthermore, by real-time monitoring of vehicle information in the simulation environment, the characteristic value of the state S(t) is dynamically filled to ensure that each state representation contains the latest information of the fleet at the current moment; max The mixed fleet state of the vehicles is set, padded with zeros to maintain a fixed dimension of the state vector.
[0034] Furthermore, the lane-changing constraints are as follows:
[0035] Speed constraint: v i '>v i-1
[0036] Among them, v i-1 is the speed of the vehicle ahead of vehicle i; when the target vehicle speed v i 'When the speed exceeds that of the vehicle in front, lane changing becomes a necessary condition; safety distance constraint: d ≥ d min
[0037] Where d represents the distance between the lane-changing vehicle and the nearest vehicle in the target lane, d min Indicates the minimum safe distance that the two can accept;
[0038] When the vehicle satisfies both of the above constraints, the lane-changing operation is performed.
[0039] Furthermore, in step S4, the SUMO simulation environment is initialized to obtain the initial state S(t) of the mixed vehicle fleet at the first entrance of the intersection; the agent samples the current state through the Actor network and generates continuous actions A(t); after executing action A(t), the vehicle fleet state is updated, and the simulation environment returns the reward R(t) and the next state S(t+1); the agent stores the (S(t), A(t), R(t), S(t+1)) obtained after executing the action at each time step as a group of samples in the experience replay pool. When the sample size is sufficient, the agent randomly extracts a batch of samples from the experience replay pool for training;
[0040] In each training cycle, the agent obtains the status information of the current mixed fleet through the SUMO simulation environment. This status information is passed as input to the Actor network and Critic network in the SAC model to generate prediction results.
[0041] Compared with the prior art, the present invention has the following significant effects:
[0042] 1. The present invention optimizes the platoon configuration in a mixed traffic environment. The application object is a mixed fleet of manually driven and automatically driven vehicles, rather than a purely automatically driven fleet. It can take into account the behavioral differences of different types of vehicles and give full play to the synergy of automatically driven and manually driven vehicles, thereby improving the traffic efficiency and safety of the entire fleet;
[0043] 2. The present invention optimizes the queuing configuration of mixed fleets based on the SAC algorithm in deep reinforcement learning, so that CAVs and HVs can achieve efficient collaboration in a more diversified arrangement, thereby achieving speed consistency and smoothness within the fleet, improving intersection traffic efficiency, reducing traffic congestion, reducing accident rates, and further enhancing the efficiency and safety of the transportation system;
[0044] 3. The present invention adopts the SAC algorithm to solve the problem of optimizing the queuing configuration of a fleet in a mixed traffic environment. First, in terms of state representation, the state of each vehicle is composed of the vehicle type (CAV or HV) and the vehicle speed, which can fully describe the real-time characteristics of the mixed fleet. Secondly, in terms of action design, the present invention not only includes the adjustment of vehicle speed, but also introduces the vehicle lane changing operation to achieve the position optimization of the fleet members on the entrance lane, ensuring the smoothness and efficiency of the entire fleet. In addition, the lane changing operation also adds safety distance and speed constraints to ensure safety during the lane changing process. Finally, the design of the reward function fully considers the overall operation effect of the fleet. By comprehensively evaluating the average speed and speed variance of the fleet, the stability and smoothness of the fleet are balanced, and the realization effect of the optimization goal is improved. The state representation, action space and reward function design scheme of the present invention are more flexible and applicable, and can effectively cope with complex changes in a dynamic traffic environment, thereby achieving efficient optimization of the queuing configuration of a mixed fleet. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 is a flow chart of the present invention;
[0046] Figure 2 It is a reward curve diagram in the embodiment;
[0047] Figure 3 This is a mixed fleet queuing configuration diagram before optimization in the embodiment;
[0048] Figure 4 This is a queuing configuration diagram of a mixed fleet after optimization in the embodiment. DETAILED DESCRIPTION
[0049] The present invention is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0050] In order to solve the problems existing in the prior art, the present invention selects the SAC algorithm to optimize the queuing configuration of mixed fleets of manual and automatic driving. By utilizing the random exploration characteristics and efficient learning ability of the SAC algorithm, it can not only handle complex situations in mixed traffic environments, but also improve the coordination effect of the fleet, and ultimately provide an innovative and effective solution for dynamic traffic management. As an emerging algorithm in the deep reinforcement learning (DRL) method, the Soft Actor-Critic (SAC) algorithm has the main advantage of being able to efficiently process continuous states and action spaces. Compared with the traditional DRL algorithm, the SAC algorithm introduces the maximum entropy objective. While ensuring the efficiency of the strategy, it enhances the random exploration ability of the algorithm, thereby avoiding the problem of easily falling into suboptimal strategies. Based on the above advantages,
[0051] In this embodiment, a cross intersection was constructed using the urban traffic simulation software SUMO, and the maximum speed limit allowed on the road was 13.89m / s. In order to eliminate interference from signal control, vehicle conflicts, and other traffic participants, only a mixed fleet consisting of manually driven vehicles and self-driving vehicles was established at the west entrance, including 8 HVs and 7 CAVs, and no signal control was set up. The CAV vehicles in the simulation all adopted the CACC following model, and the HV vehicles all adopted the Krauss following model. In this embodiment, the average speed and speed variance of the mixed fleet passing through the intersection are used as evaluation indicators, the fleet queue configuration is optimized based on the SAC algorithm, and the optimization effect is analyzed. In order to balance the contribution of the average speed and speed variance to the reward value, the weight coefficients of the two in the reward function are set to 0.5. In order to intuitively demonstrate the effect of the present invention, the numerical changes of relevant indicators before and after optimization are compared. As shown Figure 1 As shown, it is a flow chart of the method for optimizing the queue configuration of mixed vehicles entering the intersection from the west in this embodiment, including the following steps:
[0052] Step 1, design the overall framework of the SAC model;
[0053] Construct a policy network (Actor) and a value network (Critic). The present invention uses the SAC algorithm to solve the queuing configuration optimization problem of a mixed fleet. The goal of the intelligent agent is to improve the stability and efficiency of the mixed fleet passing through the intersection. The SAC algorithm framework mainly includes a policy network (Actor) and a value network (Critic). The main function of the Actor is to generate the optimal strategy under the current state. The output includes the mean and logarithmic standard deviation of the action, and the action is generated by normal distribution sampling. In training, the entropy regularization term is used to balance exploration and utilization. Critic is used to evaluate the quality of the Actor strategy, using a dual network structure to estimate the long-term return (Q value) of the state-action pair, and combining with the target network to achieve stable training. The Actor network and the Critic network are both composed of two layers of fully connected networks, each with 256 neurons in the hidden layer, and the activation function is ReLU, and the corresponding results are generated through the output layer, where the Actor outputs the mean and standard deviation of the action, and the Critic outputs the Q value of the state-action pair. In order to effectively store and utilize the experience data generated by the interaction between the intelligent agent and the environment, the present invention introduces an experience replay pool (ReplayBuffer). The agent stores the state, action, reward and next state obtained after performing an action at each time step as a set of samples in the experience replay pool. When the sample size is sufficient, the agent will randomly select a batch of samples from the experience replay pool for training to break the time correlation between the data and improve the stability and efficiency of training.
[0054] Based on the characteristics of the mixed fleet and the scene setting of the west entrance of the intersection in the embodiment, the Actor network and the Critic network are constructed. The embodiment sets the maximum fleet size to 15 vehicles, and the state setting of each vehicle includes two pieces of information: vehicle type and speed. Therefore, the state input dimension is 30 dimensions and the action output dimension is 15 dimensions. Both the Actor network and the Critic network are set up with two layers of fully connected networks, and the hidden layer is set up with 256 neurons, and the activation function is ReLU.
[0055] Step 2: Design the agent and set the state S(t), action A(t) and reward R(t);
[0056] The state S(t) is set. Aiming at the mixed fleet queuing configuration optimization problem, the present invention designs a state representation method based on multi-dimensional feature extraction to capture the dynamic relationship between vehicles and their impact on traffic efficiency. Specifically, the maximum fleet size is set to N max, the state of each vehicle contains two key features: vehicle type u and speed v. Among them, vehicle type u is defined as a binary variable, u = 1 represents a connected autonomous vehicle (CAV), and u = 0 represents a manually driven vehicle (HV); vehicle speed v is the instantaneous speed value of the vehicle. By arranging the feature vectors of all team members, the global state vector of the team is constructed:
[0057]
[0058] Among them, S(t) represents the overall state of the mixed fleet at time t; u i represents the vehicle type of the i-th vehicle in the fleet; v i represents the speed of the i-th vehicle in the convoy.
[0059] In practical applications, by real-time monitoring of vehicle information in the simulation environment, the characteristic value of the state S(t) is dynamically filled to ensure that each state representation contains the latest information of the fleet at the current moment; max The mixed fleet state of the vehicles is set and padded with zeros to maintain a fixed dimension of the state vector, thus ensuring the consistency of the algorithm input.
[0060] Action A(t) is set. In order to dynamically adjust the position of vehicles in a mixed convoy on the entrance lane of an intersection, thereby optimizing the convoy configuration, the present invention designs an action setting method, which includes two parts: vehicle speed adjustment and vehicle lane change. Specifically, the agent outputs the action set A(t) of the entire mixed convoy based on the input represented by the state:
[0061]
[0062] Among them, A(t) represents the set of vehicle actions of the mixed fleet at time t; a i ∈[0,1], represents the action decision value of the i-th vehicle in the fleet.
[0063] For the speed adjustment of the i-th vehicle, set the target speed of vehicle i to v i ', the calculation formula is:
[0064] v i '=a i ·v max (3)
[0065] Among them, v max Indicates the maximum speed limit allowed on the road.
[0066] When the target speed of the i-th vehicle exceeds the speed v of the preceding vehicle i-1 , vehicle i needs to adjust its position in the convoy by changing lanes to overtake, thereby optimizing the overall convoy configuration. To ensure the safety and smoothness of lane changing, the present invention designs the following lane changing constraints:
[0067] 1) Speed constraints:
[0068] v i '>v i-1 (4)
[0069] Among them, v i-1 is the speed of the vehicle ahead of vehicle i. When the target speed exceeds the speed of the vehicle ahead, lane changing becomes a necessary condition.
[0070] 2) Safety distance constraints:
[0071] d≥d min (5)
[0072] Where d represents the distance between the lane-changing vehicle and the nearest vehicle in the target lane, d min It represents the minimum safe distance between the two. This constraint ensures that the lane change operation will not cause danger to other vehicles in the target lane.
[0073] When the vehicle satisfies both of the above constraints, a lane change operation is performed, thereby optimizing the dynamic interaction between vehicles. The action setting strategy of the present invention ensures that the deep reinforcement learning agent can effectively adjust the motion state of each vehicle in the fleet, thereby achieving overall optimization of the mixed fleet configuration under dynamic traffic environments.
[0074] Reward R(t) is set. In order to optimize the queuing configuration of mixed convoys and improve the stability and traffic efficiency of convoys passing through intersections, the present invention designs a reward function based on the average speed and speed variance of the convoy to guide the learning process of the agent. The actions taken by the agent at each time step can get corresponding reward feedback at the next time step, and are used to evaluate and improve the action strategy. At each time step t, the speed set of all vehicles in the convoy is obtained:
[0075]
[0076] Where V represents the instantaneous speed value set of all vehicles in the convoy; v i represents the speed of the i-th vehicle in the convoy.
[0077] According to the speed set V, the average speed of the fleet is calculated as:
[0078]
[0079] in, It is the average speed of all vehicles in the fleet. The larger the value, the higher the traffic efficiency of the intersection. Therefore, this item is a positive reward.
[0080] The present invention uses speed variance to represent the volatility of vehicle speed in a fleet, and the calculation formula is:
[0081]
[0082] Among them, the velocity variance v var It indicates the degree of discreteness of the speeds of all vehicles in the convoy. The smaller the value, the closer the speeds of the vehicles in the convoy are, thereby improving the speed consistency within the convoy and enhancing the safety and stability of the convoy operation.
[0083] Since speed fluctuations will have a negative impact on the fleet configuration, this item is a negative reward. In order to balance the need for improving the overall speed of the fleet and the stability of the speed, the reward function designed by the present invention is expressed as:
[0084]
[0085] Wherein, R(t) represents the reward value of the mixed fleet at time t, which is defined as the difference between the average speed of the fleet and the speed variance, where the average speed is used as a positive reward and the speed variance is used as a negative reward; α, β∈[0,1], respectively represent the weight coefficients of the average speed and speed variance in the reward function, and their values reflect the degree of influence of the two factors on the reward function. In practical applications, the present invention can set the weight coefficient according to specific needs to balance the contribution of the two factors to the reward value. The reward function of the present invention achieves the goal of optimizing the configuration of mixed fleets in a dynamic traffic environment by encouraging an overall increase in fleet speed and suppressing speed fluctuations.
[0086] Step 3, constructing a simulation model;
[0087] The present invention uses the urban traffic simulation software SUMO to simulate the cross intersection, and implements the proposed mixed fleet optimization method through Python programming, and realizes the real-time control of the vehicle through the Traci interface. In order to eliminate the interference of signal control, vehicle conflicts and other traffic participants, the present invention only establishes a group of mixed fleets of manual driving and automatic driving on one entrance road for analysis, and does not set signal control. The CAV vehicles in the simulation all use the CACC (Cooperative Adaptive Cruise Control) model to simulate the vehicle following behavior. Unlike the traditional following model, the CACC model combines vehicle-to-vehicle communication and allows vehicles to collaborate, so that the calculation of safe speed and vehicle spacing is more accurate, and can more accurately reflect the interactive behavior of CAV. HV vehicles all use the Krauss following model to simulate the behavioral characteristics of manually driven vehicles. The Krauss following model is a classic microscopic traffic flow model that can effectively describe the acceleration, deceleration and dynamic interaction characteristics of manually driven vehicles with the preceding vehicle.
[0088] Step 4: Create an experience replay pool.
[0089] An experience replay pool is constructed to store data such as the state, action, reward, and next state generated by the interaction between the agent and the environment. The capacity of the experience replay pool is set to 200 to ensure the diversity of samples.
[0090] Step 5: Randomly initialize network parameters and experience replay pool.
[0091] Randomly initialize the parameters of the Actor network and the Critic network, and synchronously initialize the parameters of the target network so that their initial parameters are consistent with the corresponding main network. At the same time, initialize the experience replay pool to store the environment interaction data later.
[0092] Step 6: Environmental interaction and sample collection.
[0093] Initialize the SUMO simulation environment and obtain the initial state S(t) of the mixed fleet at the west entrance of the intersection. The agent samples the current state through the Actor network and generates continuous actions A(t). After executing action A(t), the fleet state is updated, and the simulation environment returns the reward R(t) and the next state S(t+1). The agent stores the (S(t), A(t), R(t), S(t+1)) obtained after executing the action at each time step as a set of samples in the experience replay pool. When the sample size is sufficient, the agent will randomly select a batch of samples from the experience replay pool for training.
[0094] Step 7: SAC model training and network parameter update.
[0095] The SAC model training process of the present invention includes two key links: forward propagation and back propagation, which are specifically trained through the policy network (Actor) and the value network (Critic). In each training cycle, the agent obtains the state information of the current mixed fleet through the SUMO simulation environment. The state information is passed as input to the Actor network and the Critic network in the SAC model to generate prediction results.
[0096] A1) Forward propagation process: In the forward propagation stage, the state information is first passed to the Actor network, and an action value in a continuous action space is output (i.e., the vehicle's speed adjustment strategy). This action is selected based on the current fleet state and aims to optimize the overall efficiency of the fleet. At the same time, the Critic network outputs the Q value in this state based on the same state input, indicating the long-term reward corresponding to this state. The outputs of the Actor network and the Critic network are both processed by the ReLU activation function to increase the nonlinear representation ability of the network and ensure that the agent can learn effective behavioral strategies in complex traffic environments. As a result, the SAC model can generate the corresponding action value and the expected reward of the action.
[0097] A2) Reward calculation and loss function: Based on the actions generated by the Actor network, the agent obtains rewards through interaction with the SUMO simulation environment. The reward function takes into account the average speed and speed variance of the fleet. A state with a higher reward indicates that the overall operation efficiency and stability of the fleet are better. For the calculation of the loss function, it consists of two parts: one is the Q-value loss calculated based on the Critic network, and the other is the action selection loss generated by the Actor network. Specifically, the loss of the Critic network is calculated by minimizing the mean square error (MSE) between the estimated Q-value and the target Q-value, while the loss of the Actor network is optimized by maximizing the Q-value.
[0098] A3) Back propagation and parameter update: The loss function obtained by forward propagation is used for back propagation. During the back propagation process, the SAC model starts from the output layer, transfers the loss backward, calculates the gradient of each layer by the chain rule, and uses the gradient descent method to update the model parameters. The parameter update of the Critic network is based on the loss of the Q value, while the parameter update of the Actor network is based on the joint optimization of the policy gradient and the Q value. The SAC algorithm uses a soft update mechanism to update the parameters of the target network. The soft update mechanism smoothly updates the current network parameters to the target network at a certain ratio, thereby reducing the training instability that may be caused during the update process. In addition, in order to improve the training efficiency, the present invention uses the Adam optimizer for parameter update. The Adam optimizer adaptively adjusts the learning rate according to the historical gradient information, thereby effectively avoiding the problem of gradient disappearance or explosion, and accelerating the convergence of the model.
[0099] By back-propagating and optimizing the parameters of the Critic network and the Actor network, the SAC agent can gradually optimize its strategy and enhance the efficiency and stability of mixed convoys passing through intersections. During multiple training sessions, the Actor network continuously learns how to choose the best action, while the Critic network continuously optimizes the prediction of future rewards, ultimately achieving the optimization of the queuing configuration of mixed convoys of manual and autonomous driving at the entrance of the intersection.
[0100] The agent randomly samples a batch of interaction data from the experience replay pool, inputs it into the Critic network to calculate the Q-value error of the current state-action pair, and constructs the loss function of the value network. The Actor network optimizes the policy distribution through forward propagation, and combines the maximum entropy goal to improve the stability and exploration ability of the policy. The back-propagation algorithm is used to calculate the gradient, and the Adam optimizer is used to update the Actor and Critic network parameters. To improve stability, the soft update mechanism is used to gradually adjust the target network parameters to ensure the smoothness of the training process.
[0101] Step 8: Set the training termination conditions.
[0102] Repeat the above interaction and training steps in the SUMO simulation environment. The maximum number of training rounds set in the embodiment is 100. When the number of training rounds reaches 100, as shown in the attached Figure 2 As shown, the reward curve converges, the training can be stopped and the optimized Actor network can be output to actually guide the adjustment of the mixed fleet platoon configuration in the embodiment.
[0103] Through the above steps, the trained Actor network can be used to optimize the queuing configuration of the mixed fleet consisting of 15 vehicles in the embodiment. The optimized fleet queuing configuration is shown in the attached figure. Figure 4 As shown in the figure, the platoon configuration without SAC algorithm optimization is shown in the figure. Figure 3 As shown. Before optimization, the average speed of the fleet passing through the intersection was 11.89m / s, and the speed variance was 4.02; after optimization, the average speed of the fleet passing through the intersection increased to 12.85m / s, and the speed variance decreased to 2.59. Compared with before optimization, the average speed of the optimized fleet increased by about 8%, and the speed variance decreased by about 55%. By comparing the numerical changes of relevant indicators before and after optimization, it can be seen that the SAC model is used to optimize the queuing configuration of the mixed fleet of manual driving and automatic driving at the entrance of the intersection, which effectively improves the driving speed of the fleet and significantly reduces the speed difference, which can improve the traffic efficiency of the intersection and alleviate traffic congestion.
Claims
1. A method for optimizing the queuing configuration of a mixed fleet of manual and automatic driving vehicles at an intersection entrance, characterized in that: The steps include: S1, construct the overall framework of the SAC model; S2, design the agent, set the state S(t), action A(t) and reward R(t); S3, using the urban traffic simulation software SUMO to build a simulation environment, the CACC model is used to simulate the car-following behavior of the autonomous driving vehicles in the simulation, and the Krauss car-following model is used to simulate the behavior characteristics of the manually driven vehicles; S4, training the SAC model, and after reaching the training stop condition, outputting the optimized Actor network for actually guiding the mixed fleet queuing configuration adjustment in the embodiment.
2. The method for optimizing the queuing configuration of a mixed fleet of manual and automatic driving vehicles at an intersection entrance according to claim 1, characterized in that: In the overall framework of the SAC model, both the Actor network and the Critic network are composed of two layers of fully connected networks. The hidden layer has 256 neurons and the activation function is ReLU. The corresponding results are generated through the output layer, where the Actor outputs the mean and standard deviation of the action, and the Critic outputs the Q value of the state-action pair.
3. The method for optimizing the queuing configuration of a mixed fleet of manual and automatic driving vehicles at an intersection entrance according to claim 1, characterized in that: In step S2, a state representation method based on multi-dimensional feature extraction is designed to capture the dynamic relationship between vehicles and their impact on traffic efficiency; Set the maximum fleet size to N max , the state of each vehicle contains two key features: vehicle type u and speed v; where vehicle type u = 1 represents a connected autonomous vehicle, and u = 0 represents a manually driven vehicle; vehicle speed v is the instantaneous speed value of the vehicle; by arranging the feature vectors of all team members, the global state vector of the team is constructed: Among them, S(t) represents the overall state of the mixed fleet at time t; u i represents the vehicle type of the i-th vehicle in the fleet; v i represents the speed of the i-th vehicle in the convoy; Design an action setting method, including vehicle speed adjustment and vehicle lane change; the agent outputs the action set A(t) of the entire mixed fleet based on the input represented by the state: Among them, A(t) represents the set of vehicle actions of the mixed fleet at time t; a i ∈[0,1], represents the action decision value of the i-th vehicle in the fleet; For the speed adjustment of the i-th vehicle, set the target speed of vehicle i to v i ', the calculation formula is: in i '=a i ·in max Among them, v max Indicates the maximum speed limit allowed on the road; When the target speed of the i-th vehicle exceeds the speed v of the preceding vehicle i-1 When , vehicle i needs to adjust its position in the convoy by changing lanes and overtaking; A reward function is designed based on the average speed and speed variance of the fleet to guide the learning process of the agent. The actions taken by the agent at each time step can get corresponding reward feedback at the next time step, which is used to evaluate and improve the action strategy. At each time step t, the speed set of all vehicles in the fleet is obtained: Among them, V represents the instantaneous speed value set of all vehicles in the fleet; According to the speed set V, the average speed of the fleet is calculated as: in, is the average speed of all vehicles in the convoy. The larger the value, the higher the traffic efficiency of the intersection. Speed variance is used to represent the fluctuation of vehicle speed in a fleet, and its calculation formula is: Among them, the velocity variance v var It indicates the discrete degree of the speed of all vehicles in the convoy. The smaller the value, the closer the speed of each vehicle in the convoy is. Then the reward function expression is: Among them, R(t) represents the reward value of the mixed fleet at time t, which is defined as the difference between the average speed of the fleet and the speed variance, where the average speed is used as a positive reward and the speed variance is used as a negative reward; α, β∈[0,1] represent the weight coefficients of the average speed and speed variance in the reward function, respectively.
4. The method for optimizing the queuing configuration of a mixed fleet of manual and automatic driving vehicles at an intersection entrance according to claim 3, characterized in that: By real-time monitoring of vehicle information in the simulation environment, the characteristic value of the state S(t) is dynamically filled to ensure that each state representation contains the latest information of the fleet at the current moment; max The mixed fleet state of the vehicles is set, padded with zeros to maintain a fixed dimension of the state vector.
5. The method for optimizing the queuing configuration of a mixed fleet of manual and automatic driving vehicles at an intersection entrance according to claim 3, characterized in that: The lane-changing constraints are as follows: Speed constraint: v i '>v i-1 Among them, v i-1 is the speed of the vehicle ahead of vehicle i; when the target speed v i 'When you exceed the speed of the vehicle ahead, changing lanes becomes a necessity; Safety distance constraint: d≥d min Where d represents the distance between the lane-changing vehicle and the nearest vehicle in the target lane, d min Indicates the minimum safe distance that the two can accept; When the vehicle satisfies both of the above constraints, the lane-changing operation is performed.
6. The method for optimizing the queuing configuration of a mixed fleet of manual and automatic driving vehicles at an intersection entrance according to claim 1, characterized in that: In step S4, the SUMO simulation environment is initialized to obtain the initial state S(t) of the mixed vehicle fleet at the first entrance of the intersection; the agent samples the current state through the Actor network and generates continuous actions A(t); after executing action A(t), the vehicle fleet state is updated, and the simulation environment returns the reward R(t) and the next state S(t+1); the agent stores the (S(t), A(t), R(t), S(t+1)) obtained after executing the action at each time step as a group of samples in the experience replay pool. When the sample size is sufficient, the agent randomly extracts a batch of samples from the experience replay pool for training; In each training cycle, the agent obtains the status information of the current mixed fleet through the SUMO simulation environment. This status information is passed as input to the Actor network and Critic network in the SAC model to generate prediction results.
Citation Information
Cited By
Decision planning method for autonomous vehicle based on deep reinforcement learning
CN120396986A