Vehicle track and traffic signal partition joint optimization method based on multi-head attention reinforcement learning
Through the multi-head attention reinforcement learning method, state space and action space are constructed, and multi-dimensional reward function is designed, which solves the problem of separation between traffic signals and vehicle trajectory optimization, realizes dynamic coordinated optimization of traffic signals and vehicle trajectory, and improves traffic efficiency.
Patent Information
- Application Number
- CN202510525715.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-07-22
AI Technical Summary
In the existing signal intersection optimization control, the optimization processing of traffic signals and vehicle trajectory is a separation module, lacking dynamic interaction and deep coordination, and failing to feedback the signal timing changes in real time, resulting in low traffic efficiency.
The multi-headed attention reinforcement learning method is adopted to construct a state space that integrates the multi-headed attention mechanism, design a matching action space, and build a multi-dimensional reward function, combining the auto-headed attention mechanism and the attention mechanism to process the original traffic state characteristics, and realize dynamic collaborative optimization of CAVs and traffic signal control system.
The calculation efficiency is improved by more than 12% and the intersection traffic efficiency is more than 10%, and the dynamic coordinated optimization of traffic signals and vehicle trajectories is realized, improving overall traffic efficiency.
Smart Images

Figure CN120356332A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of traffic engineering. Specifically, it relates to a joint optimization method for vehicle trajectories and traffic signal zoning based on multi-head attention reinforcement learning. Background Technique
[0002] In the optimization control of signal intersections, the research on the optimization of traffic signals and vehicle trajectories has gradually developed from a separate mode to a collaborative mode. However, in the proposed collaborative optimization framework, traffic signals and trajectory optimization are actually still treated as "separate" modules. The optimization of signal timing depends on the prediction of traffic flow and historical data, while vehicle trajectory optimization is carried out independently based on these preset signal conditions, failing to feedback the dynamic changes of signal timing to the trajectory optimization process in real time, not forming a dynamic closed loop, and lacking true dynamic interaction and deep collaboration. In contrast, the introduction of deep reinforcement learning (DRL) provides a new solution to this problem, breaking through the limitations of traditional separate optimization methods. Using vehicle-to-everything (V2X) technology, connected automated vehicles (CAVs) can share traffic information in real time. Through dynamic learning and feedback mechanisms, DRL can adjust vehicle trajectories and signal timing in real time, forming a closed-loop control system. Such a system can adjust traffic signals and vehicle trajectories in real time in a changing traffic environment, achieve true dynamic collaboration, and improve the overall traffic efficiency.
[0003] Based on this, the present invention aims to: (1) improve the computational efficiency by constructing a state space that integrates a multi-head attention mechanism autoencoder; (2) construct a matching action space to coordinate and optimize the longitudinal control and lateral decision-making of CAVs; (3) construct a multi-dimensional reward function to guide CAVs and the traffic signal control system to collaboratively learn the optimal control strategy. The present invention aims to adjust traffic signals and vehicle trajectories in real time in a changing traffic environment, achieve dynamic collaboration, and improve the overall passing efficiency of intersections. Summary of the Invention
[0004] To solve the above problems, the present invention proposes a joint optimization method for vehicle trajectories and traffic signal zoning based on multi-head attention reinforcement learning.
[0005] The present invention relates to a joint optimization method for vehicle trajectories and traffic signal zoning based on multi-head attention reinforcement learning, including the following steps:
[0006] Step 1: Construct a state space that integrates a multi-head attention mechanism autoencoder;
[0007] Step 2: Construct a matching action space;
[0008] Step 3: Construct a multi-dimensional reward function;
[0009] Step 4: Design an evaluation index of "comprehensive score".
[0010] Further, in Step 1, the original traffic state feature X includes the speed v of the vehicle ahead n-1 , the speed difference Δv from the vehicle ahead, the expected safety distance L s , the jerk j t , the shortest queue length L among all candidate lanes q,min , the queue length L of the current lane qj , the distance d to the stop line l , the distance d from the vehicle ahead p , the acceleration ρ t , the speed v t There are a total of 10 traffic state features, that is, X = (v n-1 , Δv, L s , j t , L q,min , L q,j , d l , d p , ρ t , v t ).
[0011] Considering the dimensionality problem caused by the exponential growth of the interaction space of multiple agents, an autoencoder with multi-head attention (AMHA) is used to reduce the dimension of the original traffic state features.
[0012] (1) Perform a linear transformation on the original traffic state feature X:
[0013] Q h = XW Q
[0014] K h = XW K
[0015] V h = XW V
[0016] In the formula, W Q , W K , W V are the transformation weight matrices of the query, key, and value respectively; Q h is the query matrix of the h-th attention head, describing the feature information that the current agent needs to focus on when trying to interact with other agents; K hV is the key matrix of the hth attention head, providing information labels for the entire feature space to assist other agents in determining whether they need to pay attention to this agent; h is the value matrix of the h-th attention head, representing the specific feature information of each agent.
[0017] (2) Compute the outputs of multiple attention heads in parallel and fuse them:
[0018]
[0019] MutilHead(Q,K,V)=Concat(head1,...,head h )W O
[0020] In the formula, Q h K h T is the correlation of the h-th attention head on different feature dimensions; d X is the dimension of the original feature X, H is the number of attention heads, and It is used to suppress the excessive size of the dot product result to enhance numerical stability; then, after the Softmax normalization operation, the weighted normalized attention score can be obtained and V h Perform weighted summation to generate the final attention output; W O It is the projection transformation matrix, which is responsible for fusing the outputs of multiple attention heads to improve the ability to express information.
[0021] (3) In order for AMHA to effectively learn traffic state characteristics and keep the key information of the original traffic state X as much as possible, it needs to be trained. The objective function is defined as:
[0022] L(ψ)=minE(XX′) 2
[0023] Where Ψ is the AMHA network parameter; X' is the reconstructed traffic state feature.
[0024] Furthermore, in step 2, CAVs perform following and lane changing operations by adjusting the throttle and steering wheel, while traffic lights achieve joint control by adjusting the green light duration. The action space is defined as:
[0025] A={A CAV ,A light}
[0026] In the formula, A CAV is the action set of CAV, A light A set of actions for a signal light.
[0027] The actions of CAVs are jointly determined by lane selection and acceleration, which is defined as a matching action space, that is, select a lane-changing action from the finite set lane = {l1, l2, l3} and match an acceleration for each action. A complete action can be represented by a tuple (l i , ρ), that is:
[0028]
[0029] where l1 = nc represents no lane change, l2 = lc represents a left lane-changing behavior, l3 = rc represents a right lane-changing behavior, and ρ is the acceleration of the CAV, m / s 2 .
[0030] The action space of the traffic signal is defined as:
[0031] A light = {T g ′}
[0032] where T g ’ is the green light adjustment time.
[0033] Furthermore, in step three, for the coupling relationship between the lateral decision-making, longitudinal control of CAVs and the traffic signal phase duration, a reward mechanism including six key dimensions of lane selection r j,t , vehicle formation r p,t , traffic efficiency r v,t , smooth operation r ρ,t , driving safety r l,t , and release efficiency r k,t is constructed, that is, the total reward R t = r j,t + r p,t + r v,t + r ρ,t + r l,t + r k,t , to guide CAVs to cooperate with the traffic signal control system to learn the optimal control strategy.
[0034] (1) Lane selection reward
[0035] In the variable lane-changing area, when CAVs face multiple candidate lanes, they should preferentially select the lane with a shorter queue length. The following lane selection reward function is designed:
[0036]
[0037] where L qj is the current queue length of lane j, veh; L q,min is the shortest queue length among the candidate lanes, veh.
[0038] (2) Vehicle formation reward
[0039] In the vehicle platoon area, to reduce the interference of HDVs and guide CAVs to form a platoon, the following CAV platoon reward function is designed:
[0040]
[0041] where p n-1 is the type of the leading vehicle.
[0042] (3) Throughput efficiency reward
[0043] To improve the ability of CAVs to pass through intersections without stagnation and encourage them to maintain a high driving speed when they are far from intersections, the following throughput efficiency reward function is designed:
[0044]
[0045] where a and b are constant coefficients, taking values of -2 / (v1 - v0) 3 and 3 / (v1 - v0) 2 respectively; v0 and v1 are the minimum and maximum speed limits of the road, in m / s; v t is the speed of the CAV at the t-th time step, in m / s.
[0046] (4) Smooth operation reward
[0047] Excessive acceleration or severe speed fluctuations will cause the CAVs to operate unstably and affect passenger comfort. Based on acceleration and jerk, a smooth operation reward is designed:
[0048]
[0049] where ρ + is the maximum acceleration limit of the road, in m / s 2 ; ρ t is the acceleration of the CAV at the t-th time step, in m / s 2 ; j t is the jerk of the CAV at the t-th time step, in m / s 3 .
[0050] (5) Driving safety reward
[0051] During the following process of CAVs, an excessive headway reduces the throughput efficiency, while a too short headway affects driving safety. To balance safety and throughput efficiency, this paper uses the Intelligent Driver Model (IDM) to calculate the expected safety distance and designs a driving safety reward function to encourage CAVs to maintain a reasonable following distance:
[0052]
[0053] where c and d are coefficients, taking values of -2 / (L s -s0) 3 and 3 / (L s -s0) 2 respectively, dp is the distance to the vehicle ahead, m; s0 is the minimum safety distance, m; L s is the expected safety distance calculated according to the IDM model, m; its calculation formula is:
[0054]
[0055] where ω is the expected time headway to the vehicle ahead, m; Δv is the speed difference from the vehicle ahead, m / s; ρ- is the maximum deceleration, m / s 2 .
[0056] (6) Release efficiency reward
[0057] To quantify the release efficiency of traffic signals, in this paper, a reward function is designed by comparing the differences in vehicle density at intersections before and after release, so as to encourage the signal lights to adaptively adjust the green light time according to the actual number of released vehicles and achieve the collaborative optimization of CAVs and signal control:
[0058]
[0059] where Δn i is the number of vehicles changing before and after release in the lane controlled by the i-th signal phase, veh; x is the observation length, m.
[0060] (7) Considering the above several factors, the total reward function is:
[0061] R t = r j,t + r p,t + r v,t + r ρ,t + r l,t + r k,t
[0062] Furthermore, in step four, a simulation scenario reference for the intersection of urban road mixed traffic flow is constructed Figure 2 As shown, the trained model is tested under different CAV penetration rates. To evaluate the model from three aspects of efficiency, comfort, and safety, an evaluation index of "comprehensive score" is designed to comprehensively evaluate the performance of the model. The calculation formula is as follows:
[0063]
[0064] where is the average speed, m / s; For average delay, s; j is the average acceleration change rate, m / s 3 。
[0065] Beneficial effects
[0066] The present invention provides a joint optimization method for vehicle trajectories and traffic signal zoning based on multi-head attention reinforcement learning. By combining an autoencoder and an attention mechanism to process the original traffic state features and setting the state space as the model input, a matching action space is proposed, and a multi-dimensional reward function is designed. Finally, a "comprehensive score" evaluation index is constructed to evaluate the model. Compared with the existing joint optimization methods, this method can improve the calculation efficiency by more than 12% and the intersection passing efficiency by more than 10%. Description of the drawings
[0067] Figure 1 is the overall flowchart of the present invention.
[0068] Figure 2 is the simulation scenario of the urban road mixed traffic flow intersection built by the present invention. Specific implementation manners
[0069] The following is combined with Figures 1 to 2 to specifically describe the implementation manners of the present invention.
[0070] The present invention provides a joint optimization method for vehicle trajectories and traffic signal zoning based on multi-head attention reinforcement learning. The process overview is as follows:
[0071] First, construct a state space based on the multi-agent soft actor-critic (MASAC) segmented joint optimization strategy, and use an autoencoder with multi-head attention (AMHA) to reduce the dimension of the original traffic state features.
[0072] Then, set action spaces for CAVs and traffic signals respectively, and construct a matching action space for CAVs to coordinate horizontal decision-making and vertical control.
[0073] Secondly, design a reward guidance mechanism including six key dimensions: lane selection, vehicle formation, passing efficiency, smooth operation, driving safety, and release efficiency.
[0074] Finally, construct a "comprehensive score" evaluation index to comprehensively evaluate the control effectiveness of the present invention from three aspects: safety, efficiency, and comfort.
[0075] The joint optimization method for vehicle trajectories and traffic signal zoning based on multi-head attention reinforcement learning has an optimization process as shown in Figure 1 , and the specific steps are as follows:
[0076] Step 1: Construct the state space of the fusion multi-head attention mechanism autoencoder
[0077] The original traffic state feature X includes the speed v of the vehicle in front n-1 , the speed difference Δv from the vehicle in front, the desired safety distance L s , the jerk j t , the shortest queue length L among all candidate lanes q,min , the queue length L of the current lane qj , the distance d to the stop line l , the distance d from the vehicle in front p , the acceleration ρ t , the speed v t There are a total of 10 traffic state features, that is, X = (v n-1 , Δv, L s , j t , L q,min , L q,j , d l , d p , ρ t , v t ).
[0078] Considering the dimensionality problem caused by the exponential growth of the interaction space of multiple agents, an autoencoder with multi-head attention (AMHA) is used to reduce the dimension of the original traffic state features.
[0079] (1) Perform a linear transformation on the original traffic state feature X:
[0080] Q h = XW Q
[0081] K h = XW K
[0082] V h = XW V
[0083] In the formula, W Q , W K , W V are the transformation weight matrices of the query, key, and value respectively; Q h is the query matrix of the h-th attention head, describing the feature information that the current agent needs to focus on when trying to interact with other agents; K h is the key matrix of the h-th attention head, providing information labels for the entire feature space to assist other agents in determining whether to pay attention to this agent; V his the value matrix of the h-th attention head, representing the specific feature information of each agent.
[0084] If trained Then Q1=K1=V1=(55.436 60.586)
[0085] (2) Compute the outputs of multiple attention heads in parallel and fuse them:
[0086]
[0087] MutilHead(Q,K,V)=Concat(head1,...,head h )W O
[0088] In the formula, Q h K h T is the correlation of the h-th attention head on different feature dimensions; d X is the dimension of the original feature X, H is the number of attention heads, and It is used to suppress the excessive size of the dot product result to enhance numerical stability; then, after the Softmax normalization operation, the weighted normalized attention score can be obtained and V h Perform weighted summation to generate the final attention output; W O It is the projection transformation matrix, which is responsible for fusing the outputs of multiple attention heads to improve the ability to express information.
[0089] If calculated Attention1(Q1,K1,V1)=(55.436 60.586), assuming that the second attention head Attention2(Q2,K2,V2)=(50.764 58.249) and Then MutilHead(Q,K,V)=(393.05 415.55 438.06 460.56 483.06)
[0090] (3) In order for AMHA to effectively learn traffic state characteristics and keep the key information of the original traffic state X as much as possible, it needs to be trained. The objective function is defined as:
[0091] L(ψ)=minE(XX′) 2
[0092] Where Ψ is the AMHA network parameter; X' is the reconstructed traffic state feature.
[0093] After training, the dimensionality-reduced traffic state feature S is obtained. Among them, Feature 1 is the vehicle dynamic interaction characteristic, which extracts the interaction information between CAVs and surrounding vehicles and reflects its ability to maintain a safe distance and relative speed through speed adjustment. Feature 2 is the vehicle dynamics characteristic, which focuses on the speed change trend, emphasizes the overall movement rhythm rather than the instantaneous acceleration, and reflects the stability and dynamics characteristics of the vehicle. Feature 3 is the spatial position adjustment ability, which depicts the position adjustment ability of CAVs in the road environment and captures the non-linear relationship related to spatial position. Feature 4 is the global traffic state characteristic, which measures the traffic congestion degree and lane selection strategy. Feature 5 is the multi-feature collaborative optimization ability, which reflects the synergistic effect between vehicle acceleration decisions and the global traffic environment and synthesizes the interaction relationship between multiple features.
[0094] Step 2: Construct a matching action space
[0095] CAVs perform car-following and lane-changing operations by adjusting the throttle and steering wheel, while traffic lights achieve joint control through the adjustment of green light duration. The action space is defined as:
[0096] A = {A CAV , A light}
[0097] In the formula, A CAV is the action set of CAVs, and A light is the action set of traffic lights.
[0098] The actions of CAVs are jointly determined by lane selection and acceleration, and are defined as a matching action space, that is, select a lane-changing action from the finite set lane = {l1, l2, l3} and match an acceleration for each action. A complete action can be represented by a tuple (l i , ρ), that is:
[0099]
[0100] In the formula, l1 = nc means no lane change is performed, l2 = lc means a left lane change behavior, l3 = rc means a right lane change behavior, and ρ is the acceleration of the CAV, m / s 2 .
[0101] The action space of traffic lights is defined as:
[0102] A light = {T g '}
[0103] In the formula, T g ' is the green light adjustment time.
[0104] Step 3: Construct a multi-dimensional reward function
[0105] A reward mechanism with six key dimensions, namely lane selection \(r\) j,t , vehicle platooning \(r\) p,t , traffic efficiency \(r\) v,t , smooth operation \(r\) ρ,t , driving safety \(r\) l,t , release efficiency \(r\) k,t , is constructed for the coupling relationship between the lateral decision-making, longitudinal control of CAVs and the traffic signal phase duration, i.e., the total reward \(R\) t \(= r\) j,t \(+ r\) p,t \(+ r\) v,t \(+ r\) ρ,t \(+ r\) l,t \(+ r\) k,t , to guide CAVs and the traffic signal control system to co-learn the optimal control strategy.
[0106] (1) Lane selection reward
[0107] In the lane-changing area, when CAVs face multiple candidate lanes, they should preferentially choose the lane with a shorter queue length. The following lane selection reward function is designed:
[0108]
[0109] where \(L\) qj is the current queue length of lane \(j\), in veh; \(L\) q,min is the shortest queue length among the candidate lanes, in veh.
[0110] Then \(r\) j,t \(= -0.46\), that is, the current lane queue length is 3, while the shortest queue length of the candidate lane is 2. Therefore, this CAV is punished to encourage it to change lanes to the lane with a shorter queue length.
[0111] (2) Vehicle platooning reward
[0112] In the vehicle platooning area, to reduce the interference of HDVs and guide CAVs to form platoons, the following CAVs platooning reward function is designed:
[0113]
[0114] where \(p\) n-1 is the type of the leading vehicle.
[0115] Then \(r\) p,t \(= 1\), that is, the type of the leading vehicle is CAV. Therefore, this CAV is rewarded to encourage the formation of CAVs queues.
[0116] (3) Traffic efficiency reward
[0117] To improve the ability of CAVs to pass through intersections without stopping and encourage them to maintain a high driving speed when they are far from intersections, the following passing efficiency reward function is designed:
[0118]
[0119] In the formula, a and b are constant coefficients, with values of -2 / (v1 - v0) 3 and 3 / (v1 - v0) 2 ; v0 and v1 are the minimum and maximum speed limits of the road, in m / s; v t is the speed of the CAV at the t-th time step, in m / s.
[0120] Then r v,t = 0.73, which means the current vehicle speed is appropriate, but it can still be accelerated appropriately.
[0121] (4) Smooth operation reward
[0122] Excessive acceleration or severe speed fluctuations will cause the CAVs to operate unstably and affect passenger comfort. Based on acceleration and jerk, a smooth operation reward is designed:
[0123]
[0124] In the formula, ρ + is the maximum acceleration limit of the road, in m / s 2 ; ρ t is the acceleration of the CAV at the t-th time step, in m / s 2 ; j t is the jerk of the CAV at the t-th time step, in m / s 3 .
[0125] Then r ρ,t = -1.52, which means the current vehicle driving comfort is poor, and it is encouraged for the CAV to drive with a smoother acceleration.
[0126] (5) Driving safety reward
[0127] During the following process of CAVs, an excessive headway reduces the passing efficiency, while a too short headway affects driving safety. To balance safety and passing efficiency, this paper uses the Intelligent Driver Model (IDM) to calculate the expected safety distance and designs a driving safety reward function to encourage CAVs to maintain a reasonable following distance:
[0128]
[0129] In the formula, c and d are coefficients, with values of -2 / (L s-s0) 3 and 3 / (L s -s0) 2 , where dp is the distance to the vehicle ahead, m; s0 is the minimum safety distance (taking 2 m), m; L s is the expected safety distance calculated according to the IDM model, m; its calculation formula is:
[0130]
[0131] In the formula, ω is the expected time headway to the vehicle ahead, m; Δv is the speed difference from the vehicle ahead, m / s; ρ- is the maximum deceleration, m / s 2 .
[0132] Then there is r l,t = 0.99, that is, the distance between the current vehicle and the vehicle ahead is relatively close to the expected safety distance, thus obtaining a reward.
[0133] (6) Reward for release efficiency
[0134] To quantify the release efficiency of traffic signals, this paper designs a reward function by comparing the difference in vehicle density at the intersection before and after release, so as to encourage the signal lights to adaptively adjust the green light time according to the actual number of released vehicles and achieve the collaborative optimization of CAVs and signal control:
[0135]
[0136] In the formula, Δn i is the number of changed vehicles before and after release in the lane controlled by the i-th signal phase, veh; x is the observation length (taking 200 m), m.
[0137] Then there is r k,t = 0.35, that is, the current joint optimization effect of the signal is average, and it is encouraged to release as many vehicles as possible in each phase.
[0138] (7) Considering the above several factors, the total reward function is:
[0139] R t = r j,t + r p,t + r v,t + r ρ,t + r l,t + r k,t
[0140] Then there is R = 1.09, that is, the total reward obtained by this CAV is 1.09.
[0141] Step Four: Design the evaluation index of "comprehensive score"
[0142] Constructed a reference for the simulation scenario of the intersection of mixed traffic flow on urban roads Figure 2As shown, the trained model is tested under different CAV penetration rates. To evaluate the model from three aspects: efficiency, comfort, and safety, a "comprehensive score" evaluation index is designed to comprehensively evaluate the model performance. The calculation formula is as follows:
[0143]
[0144] In the formula, is the average speed, m / s; is the average delay, s; j is the average acceleration change rate, m / s 3 .
[0145] Table 1 Model operation results
[0146]
[0147] The above content of the present invention is only the preferred embodiment of the present invention and is not used to limit the implementation of the present invention. Those of ordinary skill in the art can easily make corresponding changes or modifications according to the main idea and spirit of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope required by the claims.
Claims
1. The described joint optimization method for vehicle trajectories and traffic signal zoning based on multi-head attention reinforcement learning is characterized in that, The steps include: Step 1: Construct the state space of the autoencoder fused with multi-head attention mechanism; Step 2: Construct a matching action space; Step 3: Construct a multi-dimensional reward function; Step 4: Design the "comprehensive score" evaluation index.
2. The vehicle trajectory and traffic signal zoning joint optimization method based on multi-head attention reinforcement learning according to claim 1, characterized in that In Step 1, the original traffic state features X include the speed v of the vehicle ahead n-1 , the speed difference Δv from the vehicle ahead, and the desired safety distance L s , the jerk j t , the shortest queue length L among all candidate lanes q,min , the queue length L of the current lane qj , the distance d to the stop line l , the distance d from the vehicle ahead p , the acceleration ρ t , the speed v t There are a total of 10 traffic state features, i.e., X = (v n-1 , Δv, L s , j t , L q,min , L q,j , d l , d p , ρ t , v t ). Taking into account the dimensionality problem caused by the exponential growth of the interaction space of multiple agents, an autoencoder with multi-head attention (AMHA) is used to reduce the dimensionality of the original traffic state features. (1) Perform linear transformation on the original traffic state feature X: Q h = XW Q K h = XW K V h = XW V where, W Q , W K , W V are the transformation weight matrices for query, key, and value respectively; Q h is the query matrix of the h-th attention head, describing the feature information that the current agent needs to focus on when attempting to interact with other agents; K h is the key matrix of the h-th attention head, providing information labels for the entire feature space to assist other agents in determining whether to pay attention to this agent; V h is the value matrix of the h-th attention head, representing the specific feature information of each agent. (2) Compute the outputs of multiple attention heads in parallel and fuse them: MutilHead(Q,K,V)=Concat(head1,...,head h )W O where Q h K h T is the correlation of the h-th attention head in different feature dimensions; d X is the dimension of the original feature X, H is the number of attention heads, and the in the denominator is used to suppress the excessive dot product result to enhance numerical stability; Subsequently, through the Softmax normalization operation, the attention scores with normalized weights can be obtained, and V h is weighted and summed to generate the final attention output; W O is a projection transformation matrix that is responsible for fusing the outputs of multiple attention heads to enhance the information expression ability. (3) In order for AMHA to effectively learn traffic state characteristics and keep the key information of the original traffic state X as much as possible, it needs to be trained. The objective function is defined as: L(ψ) = min E(X - X′) 2 Where Ψ is the AMHA network parameter; X' is the reconstructed traffic state feature.
3. The vehicle trajectory and traffic signal zoning joint optimization method based on multi-head attention reinforcement learning according to claim 1, characterized in that In step 2, CAVs perform following and lane changing operations by adjusting the throttle and steering wheel, while traffic lights achieve joint control by adjusting the green light duration. The action space is defined as: A = {A CAV , A light} where, A CAV is the action set of the CAV, and A light is the action set of the signal lamp. The actions of CAVs are jointly determined by lane selection and acceleration, defined as a matching action space, that is, selecting a lane-changing action from a finite set lane = {l1, l2, l3} and matching an acceleration for each action. A complete action can be represented by a tuple (l i , ρ), that is: where \(l_1 = nc\) represents not executing a lane change, \(l_2 = lc\) represents a left lane change behavior, \(l_3 = rc\) represents a right lane change behavior, and \(\rho\) is the acceleration of the CAV, in m / s 2 . The action space of the signal light is defined as: A light = {T g '} Where T g ’ is the green light adjustment time.
4. The vehicle trajectory and traffic signal zoning joint optimization method based on multi-head attention reinforcement learning according to claim 1, characterized in that In Step 3, aiming at the coupling relationship among the lateral decision-making, longitudinal control of CAVs and traffic signal phase duration, a reward mechanism including six key dimensions of lane selection r j,t , vehicle platoon r p,t , traffic efficiency r v,t , smooth operation r ρ,t , driving safety r l,t , release efficiency r k,t is constructed, that is, the total reward R t = r j,t + r p,t + r v,t + r ρ,t + r l,t + r k,t , to guide CAVs to cooperate with the traffic signal control system to learn the optimal control strategy. (1) Lane selection reward In the variable lane area, when CAVs face multiple candidate lanes, they should give priority to the lane with shorter queue length. The following lane selection reward function is designed: where, L qj is the current queue length of lane j, in veh; L q,min is the shortest queue length among the candidate lanes, in veh. (2) Vehicle Formation Rewards In the vehicle formation area, in order to reduce the interference of HDVs and guide CAVs to form a formation, the following CAVs formation reward function is designed: where p n-1 is the type of the vehicle ahead. (3) Traffic efficiency reward In order to improve the ability of CAVs to pass through intersections without stagnation and encourage them to maintain a higher driving speed when they are far away from the intersection, the following traffic efficiency reward function is designed: where a and b are constant coefficients, taking values of -2 / (v1 - v0) and 3 / (v1 - v0), respectively 3 ; v0 and v1 are the minimum and maximum speed limits of the road, in m / s; v 2 is the speed of the CAV at the t-th time step, in m / s. t (4) Rewards for Smooth Operation Excessive acceleration or drastic speed fluctuations can cause CAVs to operate unstably, affecting passenger comfort. Smooth operation rewards are designed based on acceleration and jerk: where ρ + is the maximum allowable acceleration of the road, m / s 2 ; ρ t is the acceleration of the CAV at the \(t\)-th time step, in m / s 2 ; \(j\) t is the jerk of the CAV at the \(t\)-th time step, in m / s 3 . (5) Driving safety rewards When CAVs are following a car, a large distance between cars reduces traffic efficiency, while a short distance between cars affects driving safety. To balance safety and traffic efficiency, this paper uses an intelligent driver model (IDM) to calculate the expected safety distance and designs a driving safety reward function to encourage CAVs to maintain a reasonable following distance: where c and d are coefficients, taking values of -2 / (L s -s0) 3 and 3 / (L s -s0) 2 , dp is the distance to the vehicle ahead, m; s0 is the minimum safety distance, m; L s is the expected safety distance calculated according to the IDM model, m; its calculation formula is: Where ω is the expected time headway of the preceding vehicle, m; Δv is the speed difference from the vehicle ahead, in m / s; ρ - is the maximum deceleration, in m / s 2 . (6) Release efficiency reward In order to quantify the efficiency of traffic signal release, this paper designs a reward function by comparing the difference in vehicle density at the intersection before and after release, so as to encourage the traffic light to adaptively adjust the green light time according to the actual number of vehicles released, and realize the coordinated optimization of CAVs and signal control: where Δn i is the number of vehicles released before and after in the lane controlled by the i-th signal phase, veh; x is the observation length, m. (7) Considering the above factors, the total reward function is: R t =r j,t +r p,t +r v,t +r ρ,t +r l,t +r k,t 5. The vehicle trajectory and traffic signal zoning joint optimization method based on multi-head attention reinforcement learning according to claim 1, characterized in that In step 4, a simulation scenario of mixed traffic flow intersection on urban roads was constructed as shown in Figure 2. The trained model was tested under different CAV penetration rates. In order to evaluate the model from three aspects, namely efficiency, comfort and safety, an "overall score" evaluation index was designed to evaluate the model performance as a whole. The calculation formula is as follows: Wherein, is the average speed, m / s; is the average delay, s; j is the average acceleration change rate, m / s 3 .