Non-signalized intersection multi-agent scheduling optimization method based on deep reinforcement learning

By applying deep reinforcement learning technology with deep reinforcement learning in signal-free intersections, the agent collaborates on decision-making and planning driving strategies, solving the problem that traditional scheduling methods are difficult to meet real-time and efficient needs, achieving rapid vehicle response and global optimization, and improving the efficiency and safety of traffic flow.

CN120108204APending Publication Date: 2025-06-06NANTONG UNIV

Patent Information

Application Number
CN202510177776.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

With the increasing number of vehicles and complex traffic at no signal intersection, traditional traffic scheduling methods are difficult to meet the needs of real-time and efficientness, resulting in vehicle interweaving, traffic conflicts and deadlocks.

Method used

Multi-agent reinforcement learning technology based on deep reinforcement learning is adopted. By equip each vehicle with an independent agent, it can coordinate decision-making and plan driving strategies based on obtaining information from itself and the surrounding environment to achieve efficient scheduling and traffic optimization within the intersection. Specific methods include initializing the agent, building an online network and target network, updating Critic network parameters, designing reward functions, using experience playback pools, and deadlock detection and processing mechanisms.

Benefits of technology

It realizes rapid response and global optimization of vehicles in signal-free intersection scenarios, reduces vehicle delay time, alleviates congestion, improves overall traffic efficiency, and ensures the safety and smoothness of traffic flow.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108204A_ABST
    Figure CN120108204A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of intelligent transportation, and particularly relates to a non-signalized intersection multi-agent scheduling optimization method based on deep reinforcement learning. The method comprises the following steps: establishing an online network and a target network based on an Actor-Critic structure, wherein the online network and the target network are used for dispatching and controlling intelligent networked vehicles; a joint reward function is designed by combining the characteristics of the non-signalized intersection, and factors such as safe distance, traffic efficiency, speed deviation and collision punishment are comprehensively considered; historical interaction data of the vehicle is stored through an experience playback pool, time sequence correlation is broken, and the sample utilization rate is increased; through a deadlock detection and processing mechanism, it is ensured that the vehicle is prevented from being deadlocked in a complex traffic environment; and providing an optimal scheduling strategy for each intelligent network connection vehicle by using the trained model, wherein the optimal scheduling strategy comprises driving speed and passing time suggestions. According to the method disclosed by the invention, the traffic safety of vehicles is effectively ensured, and efficient scheduling and traffic control of the non-signalized intersection are realized while the waiting time is shortened and the traffic jam is relieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent transportation, and specifically relates to a multi-agent scheduling optimization method for unsignalized intersections based on deep reinforcement learning. Background Art

[0002] As the intelligent networked transportation system becomes increasingly perfect, the problems faced by urban road networks, such as the increase in vehicles and complex traffic flow, are becoming increasingly prominent. Especially in the scenario of unsignalized intersections, traditional timing control, induction control and adaptive control can no longer meet the needs of real-time and efficient traffic scheduling. With the popularization of vehicle autonomous driving and vehicle-road collaboration technology, intersections are bottleneck sections of traffic flow, and their efficient operation directly affects the smoothness and safety of the entire urban transportation system; however, unsignalized intersections lack the command of traffic lights, which makes it easy for vehicles to interweave, traffic conflicts and even deadlocks. On the other hand, they put forward higher requirements for vehicle collaboration and scheduling, requiring a solution that can respond quickly and take into account overall optimization.

[0003] In this context, multi-agent reinforcement learning technology is becoming a key breakthrough in traffic scheduling at unsignalized intersections. By equipping each vehicle or a group of vehicles with an independent agent, the vehicles can make collaborative decisions and plan driving strategies based on their own and surrounding environment information (such as position, speed, acceleration, and the status of neighboring vehicles) to achieve efficient scheduling and traffic optimization within the entire intersection. Among them, the multi-agent deep deterministic policy gradient (MADDPG) algorithm combines the advantages of deep learning and reinforcement learning. It can learn strategies for each agent separately, while maintaining a shared critic network for value evaluation, thereby achieving stable and effective training in a complex multi-agent interaction environment.

[0004] Compared with traditional signal control methods, the multi-agent collaborative scheduling optimization method based on MADDPG has both real-time and flexibility: on the one hand, vehicles can respond to changes in surrounding traffic information in seconds or even sub-seconds, accurately adjust vehicle speed or decide on the order of passage; on the other hand, the online learning mechanism of the agent can continuously optimize the decision-making strategy according to dynamic changes such as traffic load and emergencies, and reduce the probability of vehicle delays and conflicts. In addition, by modeling and jointly optimizing the interaction between vehicles, it can also effectively avoid stagnation or congestion caused by vehicles competing for the same conflict area, bringing a more efficient and safer way of organizing traffic flow for unsignalized intersections. Combined with the Internet of Vehicles technology and adaptive scheduling algorithms, information sharing and collaboration between agents are also smoother, which is of great significance to the development and upgrading of future urban transportation networks. Summary of the invention

[0005] The purpose of the present invention is to overcome the shortcomings of the prior art and propose a multi-agent scheduling optimization method for unsignalized intersections based on deep reinforcement learning. By introducing a multi-agent reinforcement learning framework, distributed scheduling and speed control are performed on intelligent connected vehicles at unsignalized intersections. Under the premise of ensuring safety and avoiding deadlock, each agent learns the driving strategy through a deep deterministic policy gradient method to achieve coordinated optimization of vehicle traffic order and speed, thereby effectively reducing vehicle delay time at intersections, alleviating congestion and improving overall traffic efficiency. This method makes full use of perception, communication, positioning and other technologies to obtain real-time information about vehicles and the environment, and combines the collaborative decision-making mechanism of multiple agents to dynamically allocate road rights in complex traffic environments, reduce unnecessary waiting, and provide reliable speed regulation and scheduling strategies for intelligent connected transportation systems.

[0006] In order to achieve the above-mentioned purpose, the present invention adopts the following technical solution: a multi-agent scheduling optimization method for unsignalized intersections based on deep reinforcement learning, comprising the following steps:

[0007] S1. Initialize the agent: Initialize the motion state S (p, v, a, r) ​​of each vehicle in the system, where p is the distance between the vehicle and the conflict point of the intersection; v is the current speed of the vehicle; a is the current acceleration of the vehicle; r is the driving intention of the vehicle, including left turn, straight and right turn; through initialization, the system accurately obtains the position information of the vehicle relative to the core area of ​​the intersection and its driving mode in the subsequent training and decision-making process, providing basic data for subsequent actions including acceleration and deceleration and decision-making;

[0008] S2. Online network and target network construction: In the multi-agent deep deterministic policy gradient MADDPG algorithm, each agent contains two main networks, the online network and the target network; the online network is responsible for real-time action prediction and value function estimation, and the target network is used to stabilize the training process and reduce learning oscillations;

[0009] Online network: This network consists of an Actor network and a Critic network. The current state S of the vehicle is used as the input of the Actor network. After calculation by the Actor network, a deterministic action A is output, which is the predicted acceleration of the vehicle. The input of the Critic network consists of two parts, including state information S and action information A. After calculation by the network, the value estimate Q(S,A) of taking action A under state S is output.

[0010] Target network: This network consists of an Actor target network and a Critic target network. The input of the Actor target network is the state S' at the next moment, and the output is the predicted action A' at the next moment. (S', A') is used as the input of the Critic target network, and the Critic target network calculates the corresponding output Q'(S', A').

[0011] The target network has the same structure as the online network, but its parameters are not synchronized with the online network in real time; instead, a soft update strategy is adopted. In each training, the parameters of the target network only make an exponential smooth transition to the parameters of the online network with a small step size τ.

[0012] S3. Update the parameters of the Critic network: Update the parameters of the online Critic network according to the TD-error temporal difference error. Specifically, first sample a batch of (s, a, r, s′, done) interaction samples in the experience replay pool, then calculate the TD-error, and then update the parameters of the Critic network by minimizing the TD-error. The TD-error formula is as follows:

[0013] TD-error=(R+γ·Q'(S',A')-Q(S,A))

[0014] Among them, Q′(S′,A′) is the value estimate of the next moment output by the target Critic network; Q(S,A) is the value prediction of the current action by the online Critic network; γ is the discount factor; R is the immediate reward;

[0015] S4, reward function: The reward function includes the safety distance reward t distance , penalty jerk caused by acceleration change, traffic efficiency incentive d distance , speed deviation reward v reward and collision penalty collision, and by weighting or combining these factors, provide targeted feedback to each vehicle at an unsignalized intersection;

[0016] S5, Experience Replay Pool: By storing, randomly sampling and reusing the interactive information of the agent in the environment, the stability of training and data utilization are improved. The specific steps are as follows:

[0017] S5-1, initialize the buffer, create a fixed-capacity storage structure, and use a double-ended queue structure to store the experience data generated by each step of interaction;

[0018] S5-2, store interaction information. After each step of environmental interaction, pack the current vehicle state s, action a, reward r, vehicle state s′ at the next moment, and whether the interaction is done into an experience tuple (s, a, r, s′, done), and append it to the buffer. When the buffer is full, remove the earliest experience to make room.

[0019] S5-3, breaking data correlation, using random sampling in the training phase to extract batch experience from the replay pool (s i ,a i,r i ,s i ′,done i ), since random sampling has nothing to do with the collection sequence, it can eliminate the strong correlation in time to a certain extent, making the network training more stable;

[0020] S5-4, multiple reuse and update: In each round or every several steps of training, the agent samples different batches of experience from the replay pool multiple times to update the policy network and value network;

[0021] S6. Deadlock detection and processing: A deadlock detection and processing mechanism is designed to avoid vehicles waiting in a circle near the conflict point. After detecting that vehicles are waiting in a circle, the corresponding sequence or signs are analyzed and adjusted to release the deadlock state in time, so that the blocked vehicles can return to the normal passable sequence, ensuring the order and smoothness of the overall traffic flow.

[0022] As the preferred technical solution of the present invention, the online network in S2 includes an Actor network and a Critic network. The specific structures of the two networks are as follows:

[0023] S2.1, Actor network input layer: The input layer has 20 nodes, covering the status information of the current vehicle and surrounding vehicles. First, the layer is normalized to make each feature consistent in numerical scale, and then the data is mapped to a 64-dimensional output;

[0024] S2.2, the first hidden layer of the Actor network: receives the 64-dimensional features from the input layer, and after the first layer of full connection mapping, the output is 64-dimensional;

[0025]

[0026] in n 1 is the number of units in this layer, b 1 is bias;

[0027] S2.3, Actor network second hidden layer: the 64-dimensional output of the first hidden layer is remapped to a 1-dimensional feature, h 1 The mapping process is as follows:

[0028] h 2 =ReLU(W 2 h 1 +b 2 )

[0029] in Further extract high-order interaction features;

[0030] S2.4, Actor network output layer: Map the above 1D representation to an action scalar to predict the acceleration of the vehicle. The specific expression of the action scalar is as follows:

[0031] a raw =W 3 h 2 +b 3

[0032] To ensure that the output action is within the predetermined range, the hyperbolic tangent activation function is applied and then multiplied by a constant w, w = 3; the final action is obtained;

[0033] a=w·tanh(a raw )

[0034] Since the output range of tanh(·) is (-1,1), the action a∈(-3,3); this control ensures that the network output complies with the conventional traffic physics constraints;

[0035] S2.5, Critic network input layer: receives the state vector S of dimension 20 as input and normalizes it through the layer to ensure that different features are consistent in numerical scale; then, the normalized state vector Mapped to a hidden dimension of 64;

[0036] S2.6, the first hidden layer of the Critic network: the above hidden state feature is recorded as h state , and then concatenated with the current action A, dimension 4, to form a fusion vector h fusion ;h state With h fusion The specific formula is as follows:

[0037]

[0038] h fusion =concat(h state ,A)

[0039] Where W state is the weight matrix of the first hidden layer, b state is the bias term of the first hidden layer;

[0040] S2.7, the second hidden layer of the Critic network: fusion Further applying the fully connected mapping and activation function, we can get a deeper fusion representation h fusion2 ;

[0041] h fusion2 =ReLU(W fusion h fusion +b fusion )

[0042] Where W fusion is the weight matrix of the second hidden layer, b fusion is the bias term of the second hidden layer;

[0043] S2.8, Critic network output layer: h is transformed into fusion2 Mapped to a scalar Q value, which is used to evaluate the value under a given state and action, thus providing a basis for action selection and strategy update;

[0044] Q(S,A)=W out h fusion2 +b out

[0045] Where W out is the output layer weight matrix, b out is the output layer bias term.

[0046] As a preferred technical solution of the present invention, the setting content of the reward function in S4 is as follows, which is used to provide effective feedback for multiple agents to learn the optimal strategy at an unsignaled intersection:

[0047] S4.1. Safety distance reward t distance :t distance is the relative distance between the vehicle and the vehicle in front; if this distance is less than 4, the reward will increase, and the increase in the reward depends on the inverse of the distance, calculated by the tanh function, indicating that the vehicle should try to keep a longer distance to avoid collision;

[0048]

[0049] The reward is negative, and a smaller distance will result in a larger negative reward, encouraging the vehicle to keep a certain distance from the vehicle in front;

[0050] S4.2, acceleration change penalty jerk: jerk is the rate of change of vehicle acceleration. Too high acceleration change will cause unstable behavior, so there is a penalty term:

[0051]

[0052] The severity of the penalty is proportional to the size of the jerk, and its purpose is to encourage smooth acceleration. w =3;

[0053] S4.3, Traffic efficiency rewardd distance : When the vehicle is at a spatial distance d from other potential conflicting vehicles or target areas distance If the distance is large or the vehicle successfully passes the conflict point, there will be a certain positive incentive; if the distance is too small, a negative penalty will be imposed to guide the vehicle to ensure the necessary safety margin;

[0054]

[0055] The reward is a logarithmic function based on the distance, which encourages vehicles to maintain a certain distance. Vehicles that are too close or too far will affect the reward.

[0056] S4.4, speed deviation reward v reward : v is the current speed of the vehicle, v m is the minimum speed of the vehicle, a M and a m are the maximum and minimum acceleration respectively; this reward is adjusted by the speed of the vehicle. The closer the speed is to the target speed, the greater the reward. w =2;

[0057]

[0058] S4.5, collision penalty collision: the penalty value when a collision occurs is -10. If no collision occurs, the reward is calculated in the normal way;

[0059]

[0060] S4.6. Joint Reward R 0 : The various rewards and penalties such as safety distance reward, acceleration change penalty, traffic efficiency reward, speed deviation reward and collision penalty are summed up to evaluate and update the comprehensive performance of the agent in the unsignalized intersection; in order to avoid extreme fluctuations in the reward value and ensure training stability, the total reward value of each time is limited to the range of [-20,20];

[0061]

[0062] R=min(20,max(-20,R 0 )).

[0063] As the preferred technical solution of the present invention, the specific method of deadlock detection and processing in S6 is as follows:

[0064] S6.1, deadlock detection: During vehicle driving, when it is detected that multiple vehicles are waiting for each other and any one of them cannot move forward, forming a closed waiting loop, the vehicle is determined to be deadlocked. The specific judgment conditions are as follows:

[0065] S6.1.1, Circular determination of virtual preceding vehicles: During the driving process, each vehicle will record a "virtual preceding vehicle" identification, that is, the vehicle in front of the vehicle in the virtual lane sequence; when the system detects that the "virtual preceding vehicle" of a certain vehicle is the same as the vehicle itself, it means that the vehicles in the queue have shown signs of circular waiting;

[0066] S6.1.2, Check whether a loop is formed: Start from a vehicle and search forward along its "virtual front vehicle" chain step by step. If it eventually returns to the vehicle itself, it means that the current group of vehicles forms a loop, that is, a deadlock occurs;

[0067] S6.1.3. Process all vehicles in the loop: Once a loop is detected, use a while flag: loop to set the lock flag of all vehicles in the loop to True to record that they are in a deadlock state; then, apply a forced unlocking strategy to the key vehicles to quickly break the mutual waiting situation and restore normal traffic;

[0068] S6.2, deadlock processing: When it is detected that the vehicles form a loop and enter a deadlock state, it is necessary to unlock the vehicles in the loop through a specific mechanism to prevent the vehicles from being stuck in permanent waiting for each other. The specific process is as follows:

[0069] S6.2.1. Deadlock detected and lock a : After the loop is confirmed, the lock flags of all vehicles in the loop are set to True; at the same time, at least one pair of vehicles are assigned lock a =+1 and lock a = -1, used to distinguish vehicles that need to be actively adjusted during the unlocking phase;

[0070] S6.2.2 Deadlock unlocking mechanism: In the deadlock state, the actual acceleration target of the vehicle a According to lock a Fine-tune; if lock a = +1, the vehicle will be slightly adjusted upward based on the current acceleration; if lock a = -1, then a deceleration strategy is adopted to break the closed loop of waiting for each other;

[0071] target a =current a +lock a

[0072] S6.2.3, reset flag: After each simulation step, the system will reset the vehicle's lock flag to False and lock a Restore to 0 to ensure that the status of each vehicle can be re-evaluated and the traffic strategy can be updated in time in the next step.

[0073] The multi-agent scheduling optimization method for unsignalized intersections based on deep reinforcement learning described in the present invention adopts the above technical solution and has the following technical effects compared with the prior art:

[0074] (1) The present invention builds an Actor network and a Critic network. By constructing an Actor network (responsible for generating actions) and a Critic network (responsible for evaluating the value of actions) based on a deep neural network for each intelligent agent, the present invention realizes refined decision-making of vehicles under different traffic conditions, thereby improving the accuracy of action prediction and training stability in unsignaled intersection scenarios.

[0075] (2) The present invention designs a joint reward system that comprehensively sets a variety of reward factors and penalty factors for the vehicle's driving process at the intersection, including safety distance, acceleration change, speed deviation, and collision penalty, so that the vehicle can maximize its traffic efficiency while taking safety into consideration, and avoid overly conservative or aggressive driving strategies.

[0076] (3) The present invention designs a deadlock detection and deadlock handling mechanism. By detecting the loop of the virtual lane queue, it accurately identifies the deadlock state of vehicles due to mutual waiting, and proposes a corresponding unlocking strategy: adjusting the acceleration of the vehicles in the loop and resetting the key mark, so as to quickly break the waiting loop, avoid vehicles from being permanently blocked, and significantly improve the traffic efficiency of the intersection.

[0077] (4) The present invention applies the MADDPG algorithm to realize multi-agent collaborative control, regards each vehicle as an independent intelligent agent and conducts collaborative training. Through global information interaction and strategy sharing, vehicles in unsignalized intersections can be adaptively dispatched and give way, taking into account the safety and traffic efficiency of the overall traffic flow. BRIEF DESCRIPTION OF THE DRAWINGS

[0078] Figure 1 A schematic diagram of a flow chart of a multi-agent scheduling optimization method for an unsignalized intersection based on deep reinforcement learning according to the present invention;

[0079] Figure 2 A virtual scene of a two-way four-lane unsignalized intersection according to the present invention;

[0080] Figure 3 It is a graph showing the distance between the vehicle and the conflict point changing over time;

[0081] Figure 4 This is a graph showing the change of vehicle speed over time on each lane involved in the present invention. DETAILED DESCRIPTION

[0082] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present application, some symbols in the embodiments of the present application are first explained to facilitate understanding by those skilled in the art.

[0083] like Figure 1 As shown, the multi-agent scheduling optimization method for unsignalized intersections based on deep reinforcement learning includes the following steps:

[0084] Step 1) (Initialize the agent): Create a virtual scene of a two-way four-lane unsignaled intersection such as Figure 2 As shown, the agent is initialized. The initial speed of each vehicle is set to 10m / s, and its driving intentions include turning left, going straight, and turning right, which are mapped to 0, 1, and 2 respectively; the lane width is cw =2.5m; in the setting of lane conflict relations, the lanes that conflict with lane 0 are [14,4,13,12,9,10,5], the lanes that conflict with lane 1 are [14,13,8,4,5,6,12], the lanes that conflict with lane 2 are [14,13,8,4,5,6,7], the lanes that conflict with lane 3 are

[14] , the lanes that conflict with lane 4 are [2,8,1,0,13,14,9], the lanes that conflict with lane 5 are [2,1,12,8,9,10,0], the lanes that conflict with lane 6 are [2,1,12,8,9,10,11], and the lanes that conflict with lane 7 are [2,1,12,8,9,10,12]. The lane that conflicts with lane 8 is [6,12,5,4,1,2,13], the lane that conflicts with lane 9 is [6,5,0,12,13,14,4], the lane that conflicts with lane 10 is [6,5,0,12,13,14,15], the lane that conflicts with lane 11 is [6], the lane that conflicts with lane 12 is [10,0,9,8,5,6,1], the lane that conflicts with lane 13 is [10,9,4,0,1,2,8], the lane that conflicts with lane 14 is [10,9,4,0,1,2,3], and the lane that conflicts with lane 15 is

[10]

[0085] Step 2) (Online network and target network): In the MADDPG algorithm, each agent contains two main networks, the online network and the target network. The online network is responsible for real-time action prediction and value function estimation, while the target network is used to stabilize the training process and reduce learning oscillations.

[0086] Step 2-1) (Training process): Write the initial hyperparameters of the model into the args.txt file, which contains (actor lr =0.0001, batch size =128, collision thr =2, critic lr =0.001, gamma = 0.8, lane num =8, learn start =20000, mat path='arvTimeNewVeh_for_train.mat', model = 'MADDPG', num episodes =1, num units =64, o agentnum =4, save rate =1, seq maxstep =12, trans r =0.998) and other information. It includes key parameters such as the learning rate of the Actor network and the Critic network, the source of training data, the number of iterations, the collision threshold, etc. The relevant data recorded during the training process are shown in Table 1:

[0087] Table 1 Training data

[0088]

[0089] Step 2-2) (Testing process): After training is completed, the model parameters will be saved in a file named "index.cptk". During testing, the model parameters are loaded from this file and then the environment is evaluated and inferred. The relevant test results and key indicators are shown in Table 2:

[0090] Table 2 Test data

[0091]

[0092]

[0093] Step 3) (Update Critic network parameters): Update the parameters of the online Critic network through the calculation of TD-error to make its estimation of the state-action value (Q value) more accurate. During the update process, the Critic network receives the current action and target value as input, and adjusts the parameters according to the gradient descent optimization strategy. The specific process data is as follows, CurrentAction(act_batch):

[0094] [[0.27276698],[-0.46654604],[-0.13992999],[-0.03665939].[3.],[0.14632147],[-0.11674444],[-0.01875685],[-0.01813661],[0.35182745],[-0.22684622],[3.],[-0.0293106],[3.

[0095] ],[0.33290177],[-0.27257816],[3.],[-0.02271797],[-0.06558337],[-0.37584846]]; TargetValue:

[0096] [[3.82066470e-01],[2.55410127e+00],[2.27651744e+00],[2.28416547e+00],[3.99474210e+00],[2 .37166658e+00],[2.28461303e+00],[2.28998521e+00],[2.45062300e+00],[-2.14768431e+01],[9.36 360753e-01],[3.99474273e+00],[2.12916730e+00],[-3.29848242e+01],[-6.52671537e+00],[5.6142 3004e-01],[1.82575739e+00],[1.27184735e-01],[-1.71530677e+01],[2.24428944e+00]]; TD-error:

[0097] [[0.52063529],[2.73378211],[2.45692952],[2.24299012],[3.98884936],[2.55387273],[2.26137023],[2.26192263],[2.59914314],[21.29501317], [0.95180851],[4.17391506],[2.31121454],[32.80309286],[6.34478727],[0.74383929],[1.78324885],[0.30675002],[16.97211773],[2.42467845]].

[0098] Step 4) (Reward function): Joint reward R 0 It is calculated by combining multiple rewards and penalties such as safety distance reward, acceleration change penalty, traffic efficiency reward, speed deviation reward and collision penalty. The design of these rewards and penalties fully considers the safety, efficiency and comfort of the vehicle during the passage, and sums up various rewards and penalties to form the final joint reward. The specific update process is as follows:

[0099] Step 4-1)(Safety Distance Reward(t distance )):t distance=[-2.163953413738653,-2.163953413738653,-2.163953413738653,-2.163953413738653,-2.163953413738653,-2.163953413738653,-2.163953413738653,-1.5355643918365445,-1.5354213840340718,-2.163953413738653,- 1.824152674000544,-1.8239931932137838,-2.163953413738653,-2.139516573207878,-2.1393429694247668,-2.163953413738653,-2.5451272764553807,-2.5449387688448177,-2.163953413738653,-2.9734626637786783];

[0100] Step 4-2) (penalty for acceleration change (jerk)):

[0101] jerk=[0.0005422948462178518,2.3761531073331455e-05,0.0010904396637960623,0.75,1.390420804396921 8e-05,0.75,-0.0,0.7689636595276929,-0.0,-0.0,-0.0,-0.0,-0.0,-0.0,-0.0,-0.0,-0.0,-0.0,-0.0,-0.0].

[0102] Step 4-3) (Traffic efficiency reward (d distance )):d distance=[-1.11098113241536,-1.1258906855011732,-1.1408448250282024,-1.155843818380946,-1.170887935354866,-1.1859774481854297,-1.201112631577586,-1.2162937627356905,-1.2315211213938835,-1.2467949898469284,-1.2621156529815254,-1. 2621156529815254,-5.3921766054673546,-1.2621156529815254,-4.948450642501732,-1.2621156529815254,-4.5424093284421385,-1.2621156529815254,-4.168542895641996,-1.2621156529815254], the reward reflects the trend of the distance between the vehicle and the conflict point over time. The larger the value, the more efficiently the vehicle passes through the intersection. The specific continuous curve of the distance between the vehicle and the conflict point over time is as follows: Figure 3 shown.

[0103] Step 4-4) (Speed ​​Deviation Reward (v reward )):v reward =[1.335395838283609,1.3400064516013799,1.3423126170078667,1.43333333333333336,1.3344579682992173,1.4333333333333336,1.533333333333339,1.326230702510608,1.5333333333333339,1.63333333333333334,1.3384693101249834,1.63 The above results reflect the speed deviation reward of the vehicle at different time steps. At the same time, the continuous curve of the vehicle’s speed in each lane changes with time is shown in Figure 3. Figure 4As shown, the dynamic change trend of vehicle speed is demonstrated, which further verifies the effect of speed deviation reward in optimizing vehicle traffic.

[0104] Step 4-5) (Collision Penalty): The penalty value is -10 when a collision occurs. If no collision occurs, the reward continues to be calculated in the normal way.

[0105] Steps 4-6) (Joint Reward R 0 ):R 0 =[-0.8296868037290686,-0.8285147851471206,-0.8291328877806594,-4.505515581336914,-5.3580333681808145,0.01937991959468066,-5.214658417118433,-4.496415408530568,-0.630620080405319,-5.174484572392290 5,-5.278033623826636,-0.5306200804053189,-5.135008453151371,-5.238557504585717,-0.43062008040531863,-5.096244799078844,-5.199793850513191,-0.3306200804053183,-5.058208680178575,-5.161757731612921].

[0106] Step 5) (Experience replay pool): Improve the stability of training and data utilization by storing, randomly sampling and reusing the interaction information of the agent in the environment. The specific steps are as follows:

[0107] Step 5-1) (Initialize buffer): Maximum capacity of buffer size =500000; the number of experience batches extracted each time during sampling size =32; the number of steps to be collected before starting learning start =2000; the total number of steps to be sampled during the entire training process is step=100000.

[0108] Step 5-2) (Store interaction information): After each step of environmental interaction, the current vehicle state s, execution action a, reward r, vehicle state s′ at the next moment, and whether the process is done are packaged into an experience tuple (s, a, r, s′, done) and appended to the buffer.

[0109] Step 5-3) (Break data correlation): In the training phase, random sampling is used to extract batch experience (s) from the playback pool i ,a i ,r i ,s i ′,done i ). Since random sampling has nothing to do with the collection sequence, it can eliminate the strong temporal correlation to a certain extent, making the network training more stable.

[0110] Step 5-4) (Multiple reuse and update): In each round or every few steps of training, the agent can sample different batches of experience from the replay pool multiple times to update the policy network and value network.

[0111] Step 6) (Deadlock detection and processing): When a deadlock is detected, the actual acceleration target of the vehicle is a According to lock a Fine-tune and adjust the data target a =[1.3725348111746987,2,2.372534811174699,2,3.372534811174699,2,4,2,4,2,4,2,-1.1066041013823538,1.2124865732709602,-2.106604101382354,2.21248657327096,-3.106604101382354,3.2124865 7327096,-4,4,-4,4,-4,4,-4,4,-4,4,-1.154050830921937,0.9427275818584404,-2.1540508309219373,1.9427275818584404,-3.1540508309219373,2.94272758185844,-4,3.94272758185844,-4,4,-4,4].

[0112] After verification by examples, this method can provide accurate driving speed and scheduling suggestions for each intelligent connected vehicle through the MADDPG algorithm. Through the intelligent multi-agent collaborative scheduling strategy, vehicles can achieve efficient and safe passage in the non-signaled intersection environment, avoid the occurrence of deadlock, and effectively reduce waiting time and improve the overall traffic efficiency. This method significantly alleviates the problem of traffic congestion and provides an efficient and reliable solution for future intelligent transportation systems.

[0113] The specific implementation scheme described above further describes in detail the purpose, technical scheme and beneficial effects of the present invention. It should be understood that the above is only a specific implementation scheme of the present invention and is not intended to limit the scope of the present invention. Any equivalent changes and modifications made by any technician in the field without departing from the concept and principle of the present invention should fall within the scope of protection of the present invention.

Claims

1. A multi-agent scheduling optimization method for unsignalized intersections based on deep reinforcement learning, characterized in that: The following steps are involved: S1. Initialize the agent: Initialize the motion state S (p, v, a, r) ​​of each vehicle in the system, where p is the distance between the vehicle and the conflict point of the intersection; v is the current speed of the vehicle; a is the current acceleration of the vehicle; r is the driving intention of the vehicle, including left turn, straight and right turn; through initialization, the system accurately obtains the position information of the vehicle relative to the core area of ​​the intersection and its driving mode in the subsequent training and decision-making process, providing basic data for subsequent actions including acceleration and deceleration and decision-making; S2. Online network and target network construction: In the multi-agent deep deterministic policy gradient MADDPG algorithm, each agent contains two main networks, the online network and the target network; the online network is responsible for real-time action prediction and value function estimation, and the target network is used to stabilize the training process and reduce learning oscillations; Online network: This network consists of an Actor network and a Critic network. The current state S of the vehicle is used as the input of the Actor network. After calculation by the Actor network, a deterministic action A is output, which is the predicted acceleration of the vehicle. The input of the Critic network consists of two parts, including state information S and action information A. After calculation by the network, the value estimate Q(S,A) of taking action A under state S is output. Target network: This network consists of an Actor target network and a Critic target network. The input of the Actor target network is the state S' at the next moment, and the output is the predicted action A' at the next moment. (S', A') is used as the input of the Critic target network, and the Critic target network calculates the corresponding output Q'(S', A'). The target network has the same structure as the online network, but its parameters are not synchronized with the online network in real time; instead, a soft update strategy is adopted. In each training, the parameters of the target network only make an exponential smooth transition to the parameters of the online network with a small step size τ. S3. Update the parameters of the Critic network: Update the parameters of the online Critic network according to the TD-error temporal difference error. Specifically, first sample a batch of (s, a, r, s′, done) interaction samples in the experience replay pool, then calculate the TD-error, and then update the parameters of the Critic network by minimizing the TD-error. The TD-error formula is as follows: TD-error=(R+γ·Q'(S',A')-Q(S,A)) Among them, Q′(S′,A′) is the value estimate of the next moment output by the target Critic network; Q(S,A) is the value prediction of the current action by the online Critic network; γ is the discount factor; R is the immediate reward; S4, reward function: The reward function includes the safety distance reward t distance , penalty jerk caused by acceleration change, traffic efficiency incentive d distance , speed deviation reward v reward and collision penalty collision, and by weighting or combining these factors, provide targeted feedback to each vehicle at an unsignalized intersection; S5, Experience Replay Pool: By storing, randomly sampling and reusing the interactive information of the agent in the environment, the stability of training and data utilization are improved. The specific steps are as follows: S5-1, initialize the buffer, create a fixed-capacity storage structure, and use a double-ended queue structure to store the experience data generated by each step of interaction; S5-2, store interaction information. After each step of environmental interaction, pack the current vehicle state s, action a, reward r, vehicle state s′ at the next moment, and whether the interaction is done into an experience tuple (s, a, r, s′, done), and append it to the buffer. When the buffer is full, remove the earliest experience to make room. S5-3, breaking data correlation, using random sampling in the training phase to extract batch experience from the replay pool (s i ,a i ,r i ,s i ′,done i ), since random sampling has nothing to do with the collection sequence, it can eliminate the strong correlation in time to a certain extent, making the network training more stable; S5-4, multiple reuse and update: In each round or every several steps of training, the agent samples different batches of experience from the replay pool multiple times to update the policy network and value network; S6. Deadlock detection and processing: A deadlock detection and processing mechanism is designed to avoid vehicles waiting in a circle near the conflict point. After detecting that vehicles are waiting in a circle, the corresponding sequence or signs are analyzed and adjusted to release the deadlock state in time, so that the blocked vehicles can return to the normal passable sequence, ensuring the order and smoothness of the overall traffic flow.

2. The multi-agent scheduling optimization method for unsignalized intersections based on deep reinforcement learning according to claim 1 is characterized in that: The online network in S2 includes the Actor network and the Critic network. The specific structures of the two networks are as follows: S2.1, Actor network input layer: The input layer has 20 nodes, covering the status information of the current vehicle and surrounding vehicles. First, the layer is normalized to make each feature consistent in numerical scale, and then the data is mapped to a 64-dimensional output; S2.2, the first hidden layer of the Actor network: receives the 64-dimensional features from the input layer, and after the first layer of full connection mapping, the output is 64-dimensional; in n1 is the number of units in this layer, b1 is the bias; S2.3, Actor network second hidden layer: The 64-dimensional output of the first hidden layer is remapped to a 1-dimensional feature. The mapping process of h1 is as follows: h2=ReLU(W2h1+b2) in Further extract high-order interaction features; S2.4, Actor network output layer: Map the above 1D representation to an action scalar to predict the acceleration of the vehicle. The specific expression of the action scalar is as follows: <h2 style=";text-align:left;direction:ltr">a<h2 style=";text-align:left;direction:ltr"> raw <h2 style=";text-align:left;direction:ltr"> =W3h2+b3 To ensure that the output action is within the predetermined range, the hyperbolic tangent activation function is applied and then multiplied by a constant w, w = 3; the final action is obtained; a=w·tanh(a raw ) Since the output range of tanh(·) is (-1,1), the action a∈(-3,3); this control ensures that the network output complies with the conventional traffic physics constraints; S2.5, Critic network input layer: receives the state vector S of dimension 20 as input and normalizes it through the layer to ensure that different features are consistent in numerical scale; then, the normalized state vector Mapped to a hidden dimension of 64; S2.6, Critic network first hidden layer: the above hidden state feature is recorded as h state , and then concatenated with the current action A, dimension 4, to form a fusion vector h fusion ;h state With h fusion The specific formula is as follows: h fusion =concat(h state ,A) Where W state is the weight matrix of the first hidden layer, b state is the bias term of the first hidden layer; S2.7, the second hidden layer of the Critic network: fusion Further applying the fully connected mapping and activation function, we can get a deeper fusion representation h fusion2 ; h fusion2 =ReLU(W fusion h fusion +b fusion ) Where W fusion is the weight matrix of the second hidden layer, b fusion is the bias term of the second hidden layer; S2.8, Critic network output layer: h is transformed into fusion2 Mapped to a scalar Q value, which is used to evaluate the value under a given state and action, thus providing a basis for action selection and strategy update; Q(S,A)=W out h fusion2 +b out Where W out is the output layer weight matrix, b out is the output layer bias term.

3. The multi-agent scheduling optimization method for unsignalized intersections based on deep reinforcement learning according to claim 2 is characterized in that: The reward function in S4 is set as follows, which is used to provide effective feedback for multi-agents to learn the optimal strategy at unsignaled intersections: S4.

1. Safety distance reward t distance :t distance is the relative distance between the vehicle and the vehicle in front; if this distance is less than 4, the reward will increase, and the increase in the reward depends on the inverse of the distance, calculated by the tanh function, indicating that the vehicle should try to keep a longer distance to avoid collision; The reward is negative, and a smaller distance will result in a larger negative reward, encouraging the vehicle to keep a certain distance from the vehicle in front; S4.2, acceleration change penalty jerk: jerk is the rate of change of vehicle acceleration. Too high acceleration change will cause unstable behavior, so there is a penalty term: The severity of the penalty is proportional to the size of the jerk, and its purpose is to encourage smooth acceleration. w =3; S4.3, Traffic efficiency rewardd distance : When the vehicle is at a spatial distance d from other potential conflicting vehicles or target areas distance If the distance is large or the vehicle successfully passes the conflict point, there will be a certain positive incentive; if the distance is too small, a negative penalty will be imposed to guide the vehicle to ensure the necessary safety margin; The reward is a logarithmic function based on the distance, which encourages vehicles to maintain a certain distance. Vehicles that are too close or too far will affect the reward. S4.4, speed deviation reward v reward : v is the current speed of the vehicle, v m is the minimum speed of the vehicle, a M and a m are the maximum and minimum acceleration respectively; this reward is adjusted by the speed of the vehicle. The closer the speed is to the target speed, the greater the reward. w =2; S4.5, collision penalty collision: the penalty value when a collision occurs is -10. If no collision occurs, the reward is calculated in the normal way; S4.6, joint reward R0: sum up the various rewards and penalties such as safety distance reward, acceleration change penalty, traffic efficiency reward, speed deviation reward and collision penalty to evaluate and update the comprehensive performance of the agent in the unsignalized intersection; To avoid extreme fluctuations in reward values ​​and ensure training stability, the total reward value for each run is limited to the range of [-20, 20]. R=min(20,max(-20,R0)).

4. The multi-agent scheduling optimization method for unsignalized intersections based on deep reinforcement learning according to claim 3 is characterized in that: The specific methods of deadlock detection and processing in S6 are as follows: S6.1, deadlock detection: During vehicle driving, when it is detected that multiple vehicles are waiting for each other and any one of them cannot move forward, forming a closed waiting loop, the vehicle is determined to be deadlocked. The specific judgment conditions are as follows: S6.1.1, Circular determination of virtual preceding vehicles: During the driving process, each vehicle will record a "virtual preceding vehicle" mark, that is, the vehicle in front of the vehicle in the virtual lane sequence; when the system detects that the "virtual preceding vehicle" of a certain vehicle is the same as the vehicle itself, it means that the vehicles in the queue have shown signs of circular waiting; S6.1.2, Check whether a loop is formed: Start from a vehicle and search forward along its "virtual front vehicle" chain step by step. If it eventually returns to the vehicle itself, it means that the current group of vehicles forms a loop, that is, a deadlock occurs; S6.1.

3. Process all vehicles in the loop: Once a loop is detected, use a while flag: loop to set the lock flag of all vehicles in the loop to True to record that they are in a deadlock state; then, apply a forced unlocking strategy to the key vehicles to quickly break the mutual waiting situation and restore normal traffic; S6.2, deadlock processing: When it is detected that the vehicles form a loop and enter a deadlock state, it is necessary to unlock the vehicles in the loop through a specific mechanism to prevent the vehicles from being stuck in permanent waiting for each other. The specific process is as follows: S6.2.

1. Deadlock detected and lock a : After the loop is confirmed, the lock flags of all vehicles in the loop are set to True; at the same time, at least one pair of vehicles are assigned lock a =+1 and lock a = -1, used to distinguish vehicles that need to be actively adjusted during the unlocking phase; S6.2.2 Deadlock unlocking mechanism: In the deadlock state, the actual acceleration target of the vehicle a According to lock a Fine-tune; if lock a = +1, the vehicle will be slightly adjusted upward based on the current acceleration; if lock a = -1, then a deceleration strategy is adopted to break the closed loop of waiting for each other; target a =current a +lock a S6.2.3, reset flag: After each simulation step, the system will reset the vehicle's lock flag to False and lock a Restore to 0 to ensure that the status of each vehicle can be re-evaluated and the traffic strategy can be updated in time in the next step.

Citation Information

Patent Citations

  • Mixed traffic flow control method based on deep reinforcement learning, medium and equipment

    CN115100850A

  • Non-signalized intersection cooperative control method based on multi-agent constraint strategy optimization

    CN115440042A

  • Intelligent network connection automobile cooperative control method and device in non-signalized intersection scene

    CN118397854A

  • Method, system and equipment for controlling traffic of vehicles at non-signalized intersection

    CN119068660A

  • Traffic control at an intersection

    US20240371264A1

Cited By

  • Non-signalized intersection end-to-end vehicle motion control method based on DRL

    CN120353172A

  • Bus and station interactive scheduling method and system based on V2X and deep reinforcement learning decision, medium and equipment

    CN121096160A

  • Tunnel construction intersection non-stop avoidance decision-making method based on MADDPG-DC

    CN121811695A

  • Small sample electricity consumption prediction method based on mutual information feature screening

    CN121858968A