A Complex Traffic Cooperative Control Method Based on Heterogeneous Multi-Agent Reinforcement Learning

Through heterogeneous graph structure and multi-level exploration strategy, the problem of collaborative control of multiple types of agents in hybrid traffic flow is solved, and safety and efficiency are improved.

CN120071623BActive Publication Date: 2025-07-29KUNMING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510446830.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-29
Estimated Expiration
2045-04-10

AI Technical Summary

Technical Problem

The prior art is difficult to effectively coordinate the control of human-driving vehicles, autonomous vehicles and signal lights in hybrid traffic flows, and lacks sufficient modeling of interactions between multiple types of agents, resulting in insufficient strategic optimization accuracy and generalization capabilities.

Method used

The heterogeneous graph structure is used to integrate observation information of different types of agents, design heterogeneous graph encoder, action generation network and value evaluation network, and combine multi-level exploration and utilization strategies to achieve coordinated control of heterogeneous multi-agents.

Benefits of technology

Effectively reduce collision risks, alleviate traffic congestion, improve the safety and traffic efficiency of traffic flow, and adapt to different road scales and traffic structures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071623B_ABST
    Figure CN120071623B_ABST
Patent Text Reader

Abstract

The present invention discloses a complex traffic collaborative control method based on heterogeneous multi-agent reinforcement learning, belonging to the technical field of intelligent transportation, including: designing a heterogeneous hybrid traffic scenario abstraction modeling method based on a heterogeneous graph structure for three types of heterogeneous agents, namely human-driven vehicles, autonomous vehicles, and traffic lights, to achieve an efficient communication mechanism among various objects; designing an encoder applicable to the heterogeneous graph structure, which can efficiently embed and integrate the observation features of heterogeneous objects; through sharing the heterogeneous graph encoder, establishing a multi-agent action generation network and a value evaluation network based on a heterogeneous graph convolutional deep network for the action generation of each agent and the value evaluation of the generated actions; constructing a deep reinforcement learning method applicable to heterogeneous multi-agents, training the model, inputting the environmental observation data of each heterogeneous multi-agent, and outputting the control quantities of all controlled objects to achieve the collaborative control of heterogeneous multi-agents in a mixed traffic flow scenario.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent transportation, and particularly to a complex traffic collaborative control method based on heterogeneous multi-agent reinforcement learning. Background Art

[0002] With the continuous maturity of autonomous driving technology, the intelligence level of vehicles is increasing day by day. However, within a foreseeable long period, human-driven vehicles and autonomous driving vehicles will still coexist in the road environment, forming a complex mixed traffic flow. Traditional traffic control methods are mostly based on a single type of vehicle or only consider signal control separately, lacking sufficient modeling of the interaction between multiple types of agents, and it is difficult to meet the real-time and efficient collaborative control requirements in complex environments. In addition, existing research based on multi-agent reinforcement learning usually assumes that agents are homogeneous, and it cannot well accommodate the multi-dimensional characteristics of heterogeneous entities such as autonomous driving vehicles, human-driven vehicles, and traffic lights, thus restricting the accuracy and generalization ability of policy optimization. Summary of the Invention

[0003] (1) Technical Problems to be Solved

[0004] Aiming at the deficiencies of the prior art, the present invention effectively integrates the observation information of different types of agents by using a heterogeneous graph structure, and realizes the learning of collaborative control strategies for heterogeneous multi-agents through a shared heterogeneous graph encoder and corresponding action generation network and value evaluation network. At the same time, independent experience storage and exploration strategies are designed for different types of agents, which can significantly improve the learning efficiency and control performance. Through the above method, it is possible to effectively reduce the collision risk, alleviate traffic congestion, and improve the safety and traffic efficiency of the overall traffic flow in a mixed traffic scenario.

[0005] (2) Technical Solutions

[0006] To achieve the above object, the present invention provides the following technical solution: A complex traffic collaborative control method based on heterogeneous multi-agent reinforcement learning, including the following steps:

[0007] S1. Design a heterogeneous multi-agent graph structure including autonomous driving vehicles, human-driven vehicles, and traffic lights;

[0008] S2. Based on the heterogeneous multi-agent graph structure in step S1, construct a heterogeneous observation encoder for autonomous driving vehicles and traffic lights, input the environmental observation data of heterogeneous multi-agents, and output the encoded environmental observation data;

[0009] S3. Design an action generation network and a corresponding value evaluation network for heterogeneous multi - agents, and share the heterogeneous observation encoder in step S2 for the action generation network and the value evaluation network. That is, the action generation network outputs the actions of the heterogeneous multi - agents according to the encoded environmental observation data, and the value evaluation network predicts the corresponding value based on the actions of the heterogeneous multi - agents.

[0010] S4. Construct a multi - level exploration and exploitation strategy suitable for the actions of heterogeneous multi - agents, and the output exploration value acts on the actions generated by the action generation network in step S3 for exploration regularization.

[0011] S5. Design two experience pools for autonomous vehicles and traffic lights, and store the environmental observation data, actions, and rewards of the agents. Specifically, each heterogeneous agent (autonomous vehicle, traffic light) outputs an action according to its action generation network. After the environment executes this action, new observations and corresponding rewards are obtained, and the quadruple information at the t - th moment is stored in the experience pool. For autonomous vehicles, record the quadruple information at the t - th moment and store it in the experience pool of the autonomous vehicle, where and respectively represent the environmental observation data, the action generated by the action generation network, the obtained reward, and the environmental observation data of the i - th autonomous vehicle at the (t + 1) - th moment at the t - th moment; for traffic lights, record the quadruple information at the t - th moment and store it in the experience pool of the traffic light, where and respectively represent the environmental observation data, the action generated by the action generation network, the obtained reward, and the environmental observation data of the k - th traffic light at the (t + 1) - th moment at the t - th moment; it should be noted that, in order to avoid the untimely feedback of the reward signal, in the present invention, the experience of autonomous vehicles is stored once per step, while the experience of traffic lights is stored once after the end of one phase cycle.

[0012] S6. Establish a heterogeneous multi - agent deep reinforcement learning framework. During the learning process, collect the experience of the heterogeneous multi - agents, input it into the action generation network designed in step S3, and obtain the action value through the corresponding value evaluation network. Update the value evaluation network according to the accuracy of the value evaluation, and update the action generation network according to the magnitude of the value of the action generated by the action generation network, so as to continuously iterate to realize the update of the network, that is, the action learning of the heterogeneous multi - agents.

[0013] S7. Save the model parameters after learning is completed, load and deploy the heterogeneous multi - agent collaborative control model, and input the environmental observation data of each heterogeneous multi - agent to output the control quantities of all controlled objects.

[0014] As a preferred technical solution of the present invention, step S1 includes the following steps:

[0015] S1.1. The autonomous vehicle and the signal light identify the surrounding object human-driven vehicles, and judge whether they should be linked to each other according to the distance threshold d threshold = 10m. If it is less than the distance threshold, it is considered that the autonomous vehicle or the signal light should be connected to the human-driven vehicle, otherwise it is not connected. Since the node types in this structure are not completely consistent, the generated graph structure belongs to a heterogeneous graph;

[0016] S1.2. Convert the heterogeneous graph in step S1.1 into a mathematical expression: use to represent the set of autonomous vehicle nodes, where represents the autonomous vehicle numbered 1 AD 、2 AD ……n AD , and n AD represents the number of autonomous vehicles; use to represent the set of human-driven vehicle nodes, where represents the human-driven vehicle numbered 1 HD 、2 HD ……n HD , and n HD represents the number of human-driven vehicles; use to represent the set of signal light nodes, where represents the signal light numbered 1 TL 、2 TL ……n TL ; Therefore, the overall node set can be expressed as V = V AD ∪V[[ID=4 VI]] HD ∪V TL ; Therefore, for any three types of nodes in this heterogeneous graph, if the spatial Euclidean distance d ijk between any two of them is less than the set distance threshold d threshold = 10m, then connect these two nodes in the heterogeneous graph structure and set the connection edge weight equal to 1; furthermore, obtain the adjacency matrix A and the feature matrix X for the heterogeneous graph. The element A ijk in the adjacency matrix A is as follows:

[0017]

[0018] In the formula, the subscripts ijk represent the i-th autonomous vehicle, the j-th human-driven vehicle, and the k-th signal light, and A ijk represents the connection situation between the i-th autonomous vehicle, the j-th human-driven vehicle, and the k-th signal light; and obtain the heterogeneous feature matrix where is the characteristic information of the autonomous vehicle, and represent the x and y coordinates of the i-th autonomous vehicle; and respectively represent the speed and acceleration of the i-th autonomous vehicle; represents the lane where the i-th autonomous vehicle is located; is the characteristic information of the human-driven vehicle, and represent the x and y coordinates of the j-th human-driven vehicle; and respectively represent the speed and acceleration of the j-th human-driven vehicle; represents the lane where the j-th human-driven vehicle is located; where represents the current phase of the k-th traffic signal, represents the time required for the k-th traffic signal to switch to the next phase;

[0019] As a preferred technical solution of the present invention, the step S2 includes the following steps:

[0020] S2.1. Since there are differences in attributes, observation dimensions, etc. between autonomous vehicles and traffic signals, design a heterogeneous encoder to extract different features for different types of nodes, so that it can uniformly output a high-dimensional representation for the subsequent decision-making network:

[0021] Z i~k = Encoder θ (A, X)

[0022] In the formula, Encoder θ (·) represents the heterogeneous encoder mapping process; θ is the parameter of the encoder; Z i~k is the encoded feature;

[0023] S2.2. For the basic principle of the encoder described in step S2.1, design modules such as multi-head self-attention graph convolution, ordinary graph convolution, heterogeneous graph convolution, and time series convolution as the main structure of the observation encoder for traffic signal type nodes; among them, the calculation process of multi-head self-attention graph convolution can be expressed as:

[0024]

[0025] In the formula, X ijk represents the environmental observation data of the i-th autonomous vehicle or the j-th human-driven vehicle or the k-th traffic signal in the heterogeneous graph; represents the environmental observation data of the intelligent agent with a different node type from X ijk , and the subscript Represents a node type different from ijk; W gat Is a learnable linear mapping matrix for multi-head self-attention graph convolution; a is the attention weight vector; e ijk Represents the unnormalized attention coefficient vector; LeakyReLU(·) represents the LeakyReLU activation function; softmax(·) represents the softmax activation function; ReLU(·) represents the ReLU activation function; α ijk Is the attention coefficient normalized by the softmax activation function; X′ ijk Represents the environmental observation data of the i-th autonomous vehicle or the j-th human-driven vehicle or the k-th traffic signal after multi-head self-attention graph convolution; For ordinary graph convolution, its calculation process can be expressed as:

[0026]

[0027] In the formula, W gcn And b gcn Respectively represent the learnable weight matrix and bias term of ordinary graph convolution, X′ ijk,gcn Is the environmental observation data of the i-th autonomous vehicle or the j-th human-driven vehicle or the k-th traffic signal after ordinary graph convolution; Represents the set of other nodes with existing connections of the node represented by the i-th autonomous vehicle or the j-th human-driven vehicle or the k-th traffic signal in the heterogeneous graph structure;

[0028] Based on the above two forms of graph convolution, the heterogeneous graph convolution process of the present invention is as follows: An edge relationship is constructed between any two node types to represent the action relationship of one node type on another node type; Therefore, for any edge type, in one information aggregation and propagation process, the process of heterogeneous graph convolution can be expressed as:

[0029]

[0030] In the formula, Represents the environmental observation data of the i-th autonomous vehicle or the j-th human-driven vehicle or the k-th traffic signal after heterogeneous graph convolution; ⊙GNN(·) represents the graph convolution operator acting on different edge types, which can be specifically expressed as the above multi-head self-attention graph convolution process and ordinary graph convolution process in the present invention; Since one node type may establish edge relationships with multiple other types of nodes, it is necessary to unify the outputs of different convolution operators to the same dimension and perform aggregation, which can be specifically expressed as:

[0031]

[0032] In the formula, AGGREGATE(·) represents the sum aggregation function; Represents the environmental observation data after global encoding of the signal lights after aggregation;

[0033] Furthermore, input the node features after heterogeneous graph convolution into the GRU network:

[0034]

[0035] In the formula, z t and r t respectively represent the update gate and the reset gate at the t-th moment. The update gate is used to determine how much information at the (t - 1)-th moment needs to be retained at the t-th moment, and the reset gate determines to what extent the hidden state at the (t - 1)-th moment is ignored; W z and b z respectively represent the parameter matrix and the bias vector of the update gate; h t-1 represents the hidden state at the (t - 1)-th moment; X t represents the environmental observation data of each node after heterogeneous graph convolution at the t-th moment; σ(·) represents the sigmoid activation function; W r and b r respectively represent the weight parameter and the bias vector of the reset gate; use the reset gate to control the dependence degree of h t-1 and calculate the candidate hidden state tanh(·) represents the tanh activation function; W and b are the parameter matrix and the bias of the candidate hidden state; after the above steps, the environmental observation data of the signal lights can realize the embedding encoding in the high-dimensional feature space;

[0036] S2.3. For autonomous vehicles, their observations have a strong real-time feedback property, that is, the intelligent agent learning strategy of autonomous vehicles needs to rely on the single-step "state-action-reward" set. Therefore, design the encoder of autonomous vehicles as a static single-step heterogeneous graph encoder, input the environmental observation data of each autonomous vehicle at the current time step, and output the encoded environmental observation data of all autonomous vehicles at the current time step. Specifically, as described in step S2.2, use the same heterogeneous graph convolution operation to perform convolution on the environmental observation data of autonomous vehicles. Denote the environmental observation data of autonomous vehicles after heterogeneous graph convolution as Furthermore, improve the non-linear expression ability through a linear mapping:

[0037]

[0038] In the formula, W1 is the weight parameter of the linear mapping layer; Encoding environmental observation data for an autonomous vehicle after linear mapping; through the two heterogeneous graph encoder structures shown in the above steps S2.1 to S2.3, effective feature abstraction and fusion can be performed on heterogeneous agents such as autonomous vehicles, human-driven vehicles, and traffic lights, providing unified and highly expressive input features for the subsequent action generation network and value evaluation network;

[0039] As a preferred technical solution of the present invention, the step S3 includes the following steps:

[0040] S3.1. Share the defined encoder part between the "action generation network" and the "value evaluation network" to reduce model redundancy and ensure coding consistency. That is, design a heterogeneous graph convolution layer to further fuse the encoded heterogeneous graph features:

[0041] h i~k =HeteroGraphCon(Z i~k )

[0042] In the formula, HeteroGraphCon(·) represents the heterogeneous graph convolution process described in step S2.2, and h i~k represents the shared high-dimensional feature representation after heterogeneous graph convolution;

[0043] S3.2. Design the action generation networks for autonomous vehicles and traffic lights, which can be specifically described as follows: For autonomous vehicles: The actions include the action probability distributions of acceleration and lane change; for traffic lights: The action can be the signal timing for the next stage. Use a multi-layer perceptron as the action generation model as follows:

[0044]

[0045] In the formula, represents the action of the i-th autonomous vehicle, specifically including the lane change action of the i-th autonomous vehicle and the acceleration action of the i-th autonomous vehicle. Among them is a probability distribution quantity, and each element of which represents the probability of each strategy in the strategy set {LCL, LK, LCR}, and LCL, LK, and LCR in the strategy set represent left lane change, no lane change, and right lane change respectively; represents the action generation network mapping process of the i-th autonomous vehicle, and φ ad is the weight parameter of the autonomous vehicle action generation network; is the signal timing for the next stage of the k-th traffic light set; Denotes the action generation network mapping process of k signal lights, φ tl are the weight parameters of the action generation network of the signal lights;

[0046] S3.3. The value evaluation network is used to evaluate the value of the current agent's action decision. For the heterogeneous multi-agent scenario of the present invention, a centralized value evaluation network mode is adopted. Input the actions and state observations of all agents into the value evaluation network to obtain the corresponding action values of each agent:

[0047]

[0048] In the formula, Q i~k represents the value brought by the actions of each autonomous vehicle and the signal lights; Q ψ (·) represents inputting the currently globally shared feature h i~k after secondary encoding, the actions of all autonomous vehicles and the actions of all signal lights to output the mapping process of the action value of each agent; ψ is the weight parameter of this centralized value evaluation network;

[0049] As a preferred technical solution of the present invention, the step S4 includes the following steps:

[0050] S4.1. In multi-agent heterogeneous reinforcement learning, in order to balance the exploration depth and policy stability of multi-dimensional vehicle and signal light control, the present invention proposes a multi-level exploration and exploitation strategy. This strategy switches between the following three modes by defining and dynamically regulating three threshold parameters: 1) Random exploration: Completely random action sampling, indicating that the agent conducts action exploration; 2) Soft exploitation: Sampling actions according to the discrete action probability distribution output by the action generation network; 3) Hard exploitation: Directly selecting the action with the highest output probability of the action generation network; For the lane-changing actions of autonomous vehicles, in order to represent the exploration and exploitation mechanisms of the acceleration actions of autonomous vehicles and the phase timing actions of signal lights, the present invention introduces a noise attenuation strategy, which continuously attenuates a threshold to make the continuous actions closer and closer to the actions generated by the action generation network.

[0051] S4.2. Based on the basic description in step S4.1, detailed threshold definitions and theoretical explanations are carried out; Denote ε as a decay threshold, and its initial value is set to ε start = 0.9, and its final decay value is ε end = 0.05, and the decay steps are steps d,ε = 10000. Therefore, as the agent is trained, the decay of ε can be specifically expressed as:

[0052]

[0053] Where α is a decay coefficient, and the decay of ε is maintained by continuously calculating the decay coefficient; this parameter is mainly used in the exploration and exploitation processes of the acceleration actions of autonomous vehicles and the phase timing actions of traffic lights. Specifically, a random Gaussian noise generator is designed as follows:

[0054] N(x) = μ + σ·Z(x), Z(x) ~ N(0, 1)

[0055] Where N(x) represents the action noise to be applied; μ and σ represent the mean and standard deviation of the noise respectively; Z(x) ~ N(0, 1) represents a standard normal distribution random number; since the noise application objects are the acceleration actions of autonomous vehicles and the phase timing actions of traffic lights, the standard deviation μ = 0 and σ = [3, 10] are taken; where σ = [3, 10] means that the standard deviation of the acceleration action noise of the autonomous vehicle is 3, and the standard deviation of the timing action of the traffic light is 10; in order to represent the smooth transition of the exploration and exploitation processes of the acceleration of the autonomous vehicle and the timing action of the traffic light, ε is applied to the noise to represent the continuously decaying noise value:

[0056]

[0057] Where represents the noise value after the ε decay effect;

[0058] Similarly, let ε h and ε m be two increasing thresholds respectively, and their initial values are set to ε h,start = 0.5 and ε m,start = 0.4 respectively, and after the increasing step number is steps a,ε = 10000, their termination values reach ε h,end = 0.95 and ε m,end = 0.85 respectively; their update follows:

[0059]

[0060] Where both β and γ are increasing coefficients, and the increase of ε h and ε m is maintained by continuously calculating the increasing coefficients;

[0061] S4.3. Based on the description in step S4.2, in each training step, the system obtains a random number r and determines which exploration and exploitation strategy to choose by judging which of the following intervals it falls into: 1) When r > ε hWhen it is time, enter the random exploration mode. At this time, the lane-changing action of the autonomous vehicle will be randomly selected from left lane-changing, no lane-changing, and right lane-changing, and the acceleration action of the autonomous vehicle and the phase timing action of the traffic signal will also be randomly selected; 2) When ε m <r ≤ ε h When it is time, enter the soft exploitation mode. At this time, the lane-changing action of the autonomous vehicle will be sampled and selected according to the probability distribution given by its action generation network. The acceleration action of the autonomous vehicle and the phase timing action of the traffic signal will be added with the noise value after ε decay described in step S4.2; 3) When r ≤ ε m When it is time, enter the hard exploitation mode. At this time, the lane-changing action of the autonomous vehicle will directly select the lane-changing action with the highest probability according to the probability distribution given by its action generation network. The acceleration action of the autonomous vehicle and the phase timing action of the traffic signal will act on the environment directly without any modification; This multi-level exploration and exploitation strategy can not only provide sufficient exploration in the early training stage, but also gradually strengthen exploitation in the later stage to ensure the convergence effect and generalization performance of the learned strategy;

[0062] As a preferred technical solution of the present invention, step S6 includes the following steps:

[0063] S6.1. Based on the experience pool for the autonomous vehicle and the traffic signal in step S5, regularly and uniformly extract a batch of experience samples from it to form two sequences containing a number of samples where b represents the sampling index; B represents the number of experience entries in one sampling;

[0064] S6.2. As described in step S3, in the training stage, it is necessary to input the samples extracted in step S6.1 into the action generation network and the value evaluation network to obtain the corresponding value estimation and policy gradient information; For the i-th autonomous vehicle, given the environmental observation data at the t + 1 moment, its action generation network predicts a lane-changing action probability distribution and an acceleration action; For the k-th traffic signal, given the environmental observation data at the t + 1 moment, its action generation network predicts the duration of the next phase; After obtaining all the actions of the autonomous vehicle and the traffic signal at the t + 1 moment, the environmental observation data at the t + 1 moment is merged with them and input into the value evaluation network to obtain the action value of each agent at the t + 1 moment. At this time, the target action values of the autonomous vehicle and the traffic signal can be calculated according to the Bellman equation as follows:

[0065]

[0066] In the formula, γ is the discount factor, taking 0.95; and respectively represent the target action values of the i-th autonomous driving vehicle and the k-th traffic light; and respectively represent the action values of the i-th autonomous driving vehicle and the k-th traffic light at the (t + 1)-th moment predicted by the value evaluation network;

[0067] S6.3. In step S6.2, the present invention respectively calculates the target action values for traffic lights and autonomous driving vehicles. For the value evaluation network, the present invention defines the loss function for updating its own weight parameters as:

[0068]

[0069] In the formula, represents the loss of the value evaluation network at the t-th moment; and respectively represent the action values of the autonomous driving vehicle and the traffic light predicted by the action value network at the t-th moment; I and K respectively represent the total numbers of autonomous driving vehicles and traffic lights; at this time, the weight parameters ψ of the value evaluation network can be updated according to minimizing this prediction loss, and the Adam optimizer is selected as the update optimizer with a learning rate of 0.0001;

[0070] S6.4. After updating the weight parameters of the value evaluation network in step S6.3, the environmental observation data of the autonomous driving vehicle and the traffic light at the t-th moment are input into their corresponding action generation networks to predict the agent actions in the environment at the t-th moment, and the predicted actions and the environmental observation data at the t-th moment are input into the value evaluation network again to predict the action values at the t-th moment. Therefore, the loss of the action generation network of the autonomous driving vehicle can be obtained as:

[0071]

[0072] In the formula, represents the loss of the action generation network of the autonomous driving vehicle at the t-th moment; represents the lane-changing action probability distribution of the i-th autonomous driving vehicle at the t-th moment; and respectively represent the values of the lane-changing action and the acceleration action of the i-th autonomous driving vehicle predicted by the value evaluation network at the t-th moment; represents the loss of the traffic light action generation network at the t-th moment; represents the value of the phase duration action of the k-th traffic light at the t-th moment; it can be seen that whether it is the action generation network of the autonomous driving vehicle or the action generation network of the traffic light, its loss value is always optimized in the direction of increasing the action value; the Adam optimizer is selected for all action generation networks of the present invention to optimize the weight parameters, and the learning rate is set to 0.0001;

[0073] Compared with the prior art, the present invention provides a complex traffic cooperative control method based on heterogeneous multi-agent reinforcement learning, which has the following beneficial effects:

[0074] Compared with the prior art, the present invention has the following beneficial effects: First, a multi-agent model of autonomous vehicles, human-driven vehicles and traffic lights is built through a heterogeneous graph structure to accurately capture the complex interactions between different types of nodes, and two encoders are combined to take into account both time series and single-step observations, greatly enhancing the dynamic scene perception ability. Second, heterogeneous multi-agent deep reinforcement learning and a three-level exploration and exploitation strategy are introduced to achieve fast convergence and efficient cooperation in large-scale mixed traffic scenarios, ensuring sufficient search depth in the early stage and enhancing the stability of the optimal strategy in the later stage. Finally, by independently storing in a decentralized experience pool and unified value evaluation, the mutual interference between different agents is reduced, and it can be flexibly extended and deployed according to different road scales or traffic structures, thus effectively improving driving safety and road traffic efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0075] Figure 1 is a schematic flow chart of the present invention;

[0076] Figure 2 is a schematic diagram of the multi-level exploration and exploitation strategy of the present invention;

[0077] Figure 3 is a schematic diagram of the heterogeneous multi-agent deep reinforcement learning framework of the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS

[0078] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0079] Please refer to Figures 1 - 3 , a complex traffic cooperative control method based on heterogeneous multi-agent reinforcement learning, characterized in that it includes the following steps:

[0080] S1. Design a heterogeneous multi-agent graph structure including autonomous vehicles, human-driven vehicles and traffic lights.

[0081] The step S1 includes the following steps:

[0082] S1.1. The autonomous vehicle and the traffic light identify the surrounding objects, the human-driven vehicle, and according to the distance threshold d thresholdDetermine whether to link with each other when the distance is equal to 10m. If the distance is less than the distance threshold, it is considered that the autonomous vehicle or signal light should be connected to the human-driven vehicle; otherwise, they are not connected. Since the node types in this structure are not completely consistent, the generated graph structure belongs to a heterogeneous graph;

[0083] S1.2. Convert the heterogeneous graph in step S1.1 into a mathematical expression: Use to represent the set of autonomous vehicle nodes, where represents the autonomous vehicle numbered 1 AD , 2 AD ... n AD , and n AD represents the number of autonomous vehicles; use to represent the set of human-driven vehicle nodes, where represents the human-driven vehicle numbered 1 HD , 2 HD ... n HD , and n HD represents the number of human-driven vehicles; use to represent the set of signal light nodes, where represents the signal light numbered 1 TL , 2 TL ... n TL ; Therefore, the overall node set can be expressed as V = V AD ∪V HD ∪V TL ; Therefore, for any three types of nodes in this heterogeneous graph, if the spatial Euclidean distance d ijk between any two nodes is less than the set distance threshold d threshold = 10m, then connect these two nodes in the heterogeneous graph structure and set the connection edge weight equal to 1; furthermore, obtain the adjacency matrix A and feature matrix X for the heterogeneous graph. The element A ijk in the adjacency matrix A is as follows:

[0084]

[0085] In the formula, the subscripts ijk represent the i-th autonomous vehicle, the j-th human-driven vehicle, and the k-th signal light, and A ijk represents the connection situation between the i-th autonomous vehicle, the j-th human-driven vehicle, and the k-th signal light; and obtain the heterogeneous feature matrix where is the feature information of the autonomous vehicle, and represent the x and y coordinates of the i-th autonomous vehicle; and represent the speed and acceleration of the i-th autonomous vehicle respectively; represents the lane where the i-th autonomous driving vehicle is located; represents the characteristic information of the human-driven vehicle, and represent the x and y coordinates of the j-th human-driven vehicle; and respectively represent the speed and acceleration of the j-th human-driven vehicle; represents the lane where the j-th human-driven vehicle is located; where represents the current phase of the k-th traffic signal, represents the time required for the k-th traffic signal to switch to the next phase;

[0086] S2. Design a heterogeneous observation encoder for autonomous driving vehicles and traffic signals based on the heterogeneous multi-agent graph structure in step S1. Input the environmental observation data of the heterogeneous multi-agent and output the encoded environmental observation data.

[0087] The step S2 includes the following steps:

[0088] S2.1. Since there are differences in attributes, observation dimensions, etc. between autonomous driving vehicles and traffic signals, design a heterogeneous encoder to extract different features for different types of nodes, so that it can uniformly output a high-dimensional representation for the subsequent decision-making network:

[0089] Z i~k = Encoder θ (A, X)

[0090] In the formula, Encoder θ (·) represents the heterogeneous encoder mapping process; θ is the parameter of the encoder; Z i~k is the encoded feature;

[0091] S2.2. According to the basic principle of the encoder described in step S2.1, design modules such as multi-head self-attention graph convolution, ordinary graph convolution, heterogeneous graph convolution, and time series convolution as the main structures of the observation encoder for traffic signal type nodes; among them, the calculation process of the multi-head self-attention graph convolution can be expressed as:

[0092]

[0093] In the formula, X ijk represents the environmental observation data of the i-th autonomous driving vehicle or the j-th human-driven vehicle or the k-th traffic signal in the heterogeneous graph; represents the environmental observation data of the agent with a different node type from X ijk The subscript represents a node type different from ijk; W gatis a learnable linear mapping matrix for multi-head self-attention graph convolution; a is the attention weight vector; e ijk represents the unnormalized attention coefficient vector; LeakyReLU(·) represents the LeakyReLU activation function; softmax(·) represents the softmax activation function; ReLU(·) represents the ReLU activation function; α ijk is the attention coefficient normalized by the softmax activation function; X′ ijk represents the environmental observation data of the i-th autonomous vehicle or the j-th human-driven vehicle or the k-th traffic signal after multi-head self-attention graph convolution; for ordinary graph convolution, its calculation process can be expressed as:

[0094]

[0095] In the formula, W gcn and b gcn respectively represent the learnable weight matrix and bias term of ordinary graph convolution, and X′ ijk,gcn is the environmental observation data of the i-th autonomous vehicle or the j-th human-driven vehicle or the k-th traffic signal after ordinary graph convolution; represents the set of other nodes with existing connections of the node represented by the i-th autonomous vehicle or the j-th human-driven vehicle or the k-th traffic signal in the heterogeneous graph structure;

[0096] Based on the above two forms of graph convolution, the heterogeneous graph convolution of the present invention is constructed as follows: an edge relationship is constructed between any two node types, representing the action relationship of one node type on another node type; therefore, for any edge type, in one information aggregation and propagation process, the process of heterogeneous graph convolution can be expressed as:

[0097]

[0098] In the formula, represents the environmental observation data of the i-th autonomous vehicle or the j-th human-driven vehicle or the k-th traffic signal after heterogeneous graph convolution; ⊙GNN(·) represents the graph convolution operator acting on different edge types, which can be specifically expressed as the above multi-head self-attention graph convolution process and ordinary graph convolution process in the present invention; since one node type may establish edge relationships with multiple other types of nodes, it is necessary to unify the outputs of different convolution operators to the same dimension and perform aggregation, which can be specifically expressed as:

[0099]

[0100] In the formula, AGGREGATE(·) represents the sum aggregation function; Represents the environmental observation data after global encoding of traffic lights after aggregation;

[0101] Furthermore, input the node features after heterogeneous graph convolution into the GRU network:

[0102]

[0103] In the formula, z t and r t respectively represent the update gate and reset gate at the t-th moment. The update gate is used to determine how much information at the (t - 1)-th moment needs to be retained at the t-th moment, and the reset gate determines to what extent the hidden state at the (t - 1)-th moment is ignored; W z and b z respectively represent the parameter matrix and bias vector of the update gate; h t-1 represents the hidden state at the (t - 1)-th moment; X t represents the environmental observation data of each node after heterogeneous graph convolution at the t-th moment; σ(·) represents the sigmoid activation function; W r and b r respectively represent the weight parameter and bias vector of the reset gate; use the reset gate to control the dependence degree of h t-1 and calculate the candidate hidden state tanh(·) represents the tanh activation function; W and b are the parameter matrix and bias of the candidate hidden state; after the environmental observation data of the traffic lights goes through the above steps, the original environmental observation data can be embedded and encoded in a high-dimensional feature space;

[0104] S2.3. For autonomous vehicles, their observations have a strong real-time feedback property, that is, the intelligent agent learning strategy of autonomous vehicles needs to rely on the single-step "state-action-reward" set. Therefore, design the encoder of autonomous vehicles as a static single-step heterogeneous graph encoder, input the environmental observation data of each autonomous vehicle at the current time step, and output the encoded environmental observation data of all autonomous vehicles at the current time step. Specifically, as described in step S2.2, use the same heterogeneous graph convolution operation to perform convolution on the environmental observation data of autonomous vehicles, and denote the environmental observation data of autonomous vehicles after heterogeneous graph convolution as Furthermore, enhance the non-linear expression ability through a linear mapping:

[0105]

[0106] In the formula, W1 is the weight parameter of the linear mapping layer; Encoding environmental observation data for an autonomous vehicle after linear mapping; through the two heterogeneous graph encoder structures shown in steps S2.1 to S2.3 above, effective feature abstraction and fusion can be performed on heterogeneous intelligent agents such as autonomous vehicles, human-driven vehicles, and traffic lights, providing unified and highly expressive input features for the subsequent action generation network and value evaluation network;

[0107] S3. Design an action generation network and a corresponding value evaluation network for heterogeneous multi-intelligent agents, sharing the heterogeneous observation encoder in step S2 between the action generation network and the value evaluation network. That is, the action generation network outputs the actions of heterogeneous multi-intelligent agents based on the encoded environmental observation data, and the value evaluation network predicts the corresponding value based on the actions of heterogeneous multi-intelligent agents.

[0108] The said step S3 includes the following steps:

[0109] S3.1. Share the defined encoder part in step S2 between the "action generation network" and the "value evaluation network" to reduce model redundancy and ensure encoding consistency. That is, design a heterogeneous graph convolutional layer to further fuse the encoded heterogeneous graph features:

[0110] h i~k = HeteroGraphCon(Z i~k )

[0111] In the formula, HeteroGraphCon(·) represents the heterogeneous graph convolution process described in step S2.2, and h i~k represents the shared high-dimensional feature representation after heterogeneous graph convolution;

[0112] S3.2. Design the action generation networks for autonomous vehicles and traffic lights, which can be specifically described as follows: For autonomous vehicles: The actions include the action probability distributions of acceleration and lane change; for traffic lights: The action can be the signal timing for the next stage. Use a multi-layer perceptron as the action generation model as follows:

[0113]

[0114] In the formula, represents the action of the i-th autonomous vehicle, specifically including the lane change action of the i-th autonomous vehicle and the acceleration action of the i-th autonomous vehicle. Among them is a probability distribution quantity, and each element represents the probability of each strategy in the strategy set {LCL, LK, LCR} respectively. The LCL, LK, and LCR in the strategy set represent left lane change, no lane change, and right lane change respectively; Represents the action generation network mapping process of the i-th autonomous vehicle, φ ad is the weight parameter of the autonomous vehicle action generation network; is the signal timing for the next stage of the k-th traffic light set; Represents the action generation network mapping process of k traffic lights, φ tl is the weight parameter of the traffic light action generation network;

[0115] S3.3. The value evaluation network is used to evaluate the value of the current agent's action decision. For the heterogeneous multi-agent scenario of the present invention, a centralized value evaluation network mode is adopted. The actions and state observations of all agents are input into the value evaluation network to obtain the action values corresponding to each agent:

[0116]

[0117] In the formula, Q i~k represents the value brought by the actions of each autonomous vehicle and traffic light; Q ψ (·) represents inputting the current globally shared feature h after secondary encoding i~k , the actions of all autonomous vehicles and the actions of all traffic lights to output the mapping process of the action value of each agent; ψ is the weight parameter of this centralized value evaluation network.

[0118] S4. As Figure 2 shown, a multi-level exploration and exploitation strategy suitable for the actions of heterogeneous multi-agents is constructed, and the output exploration value is used to explore and regularize the actions generated by the action generation network in step S3.

[0119] The said step S4 includes the following steps:

[0120] S4.1. In multi-agent reinforcement learning, in order to balance the exploration depth and policy stability of multi-dimensional vehicle and traffic light control, the present invention proposes a multi-level exploration and exploitation strategy. This strategy switches between the following three modes by defining and dynamically regulating three threshold parameters: 1) Random exploration: completely random action sampling, indicating that the agent conducts action exploration; 2) Soft exploitation: sampling actions according to the discrete action probability distribution output by the action generation network; Hard exploitation: directly selecting the action with the highest output probability of the action generation network; For the lane-changing actions of autonomous vehicles, in order to represent the exploration and exploitation mechanisms of the acceleration actions of autonomous vehicles and the phase timing actions of traffic lights, the present invention introduces a noise attenuation strategy, and continuously attenuates a threshold to make the continuous actions closer and closer to the actions generated by the action generation network.

[0121] S4.2. Based on the basic description in step S4.1, conduct detailed threshold definition and theoretical elaboration; denote ε as a decay threshold, and its initial value is set to ε start = 0.9, and its final decay value is ε end = 0.05, and the number of decay steps is steps d,ε = 10000. Therefore, as the agent is trained, the decay of ε can be specifically expressed as:

[0122]

[0123] In the formula, α is a decay coefficient, and the decay of ε is maintained by continuously calculating the decay coefficient; this parameter is mainly used in the exploration and exploitation processes of the acceleration actions of autonomous vehicles and the phase timing actions of traffic lights. Specifically, design a random Gaussian noise generator as follows:

[0124] N(x) = μ + σ·Z(x), Z(x) ~ N(0, 1)

[0125] In the formula, N(x) represents the action noise to be applied; μ and σ represent the mean and standard deviation of the noise respectively; Z(x) ~ N(0, 1) represents a standard normal distribution random number; since the noise application objects are the acceleration actions of autonomous vehicles and the phase timing actions of traffic lights, take the standard deviation μ = 0, σ = [3, 10]; where σ = [3, 10] means that the standard deviation of the acceleration action noise of the autonomous vehicle is 3, and the standard deviation of the timing action of the traffic light is 10; in order to represent the smooth transition of the exploration and exploitation processes of the acceleration of autonomous vehicles and the timing actions of traffic lights, apply ε to the noise to represent the continuously decaying noise value:

[0126]

[0127] In the formula, represents the noise value after the ε decay effect;

[0128] Similarly, denote ε h and ε m as two increasing thresholds respectively, and set their initial values to ε h,start = 0.5 and ε m,start = 0.4 respectively, and after the increasing number of steps is steps a,ε = 10000, their termination values reach ε h,end = 0.95 and ε m,end = 0.85 respectively; their update follows:

[0129]

[0130] where β and γ are both decay increasing coefficients, and ε is maintained by continuously calculating the increasing coefficients h and ε m increase;

[0131] S4.3. Based on the description in step S4.2, in each training step, the system obtains a random number r and determines which exploration and exploitation strategy to choose by judging which interval it falls into: 1) When r > ε h , enter the random exploration mode. At this time, the lane-changing action of the autonomous vehicle will be randomly selected from left lane-changing, non-lane-changing, and right lane-changing, and the acceleration action of the autonomous vehicle and the phase timing action of the traffic signal will also be randomly selected; 2) When ε m < r ≤ ε h , enter the soft exploitation mode. At this time, the lane-changing action of the autonomous vehicle will be sampled and selected according to the probability distribution given by its action generation network, and the acceleration action of the autonomous vehicle and the phase timing action of the traffic signal will be added with the noise value after the ε decay described in step S4.2; 3) When r ≤ ε m , enter the hard exploitation mode. At this time, the lane-changing action of the autonomous vehicle will directly select the lane-changing action with the highest probability according to the probability distribution given by its action generation network, and the acceleration action of the autonomous vehicle and the phase timing action of the traffic signal will act on the environment without any modification; This multi-level exploration and exploitation strategy can not only provide sufficient exploration in the early training stage, but also gradually strengthen exploitation in the later stage to ensure the convergence effect and generalization performance of the learned strategy.

[0132] S5. Design two experience pools for the autonomous vehicle and the traffic signal, and store the environmental observation data, actions, and rewards of the agent; specifically, each heterogeneous agent (autonomous vehicle, traffic signal) outputs an action according to its action generation network. After the environment executes this action, new observations and corresponding rewards are obtained, and the quadruple information at the t-th moment is stored in the experience pool. For the autonomous vehicle, record the quadruple information at the t-th moment and store it in the experience pool of the autonomous vehicle, where and respectively represent the environmental observation data of the i-th autonomous vehicle at the t-th moment, the action generated by the action generation network, the obtained reward, and the environmental observation data of the i-th autonomous vehicle at the t + 1-th moment; for the traffic signal, similarly record the quadruple information at the t-th moment and store it in the experience pool of the traffic signal, where and respectively represent the environmental observation data of the k-th signal light at time t, the action generated by the action generation network, the obtained reward, and the environmental observation data of the k-th signal light at time t+1; it should be noted that, in order to avoid untimely feedback of the reward signal, the experience storage frequency of the autonomous vehicle in the present invention is once per step, while the experience storage of the signal light is stored once after the end of one phase cycle;

[0133] S6. As Figure 3 shown, establish a heterogeneous multi-agent deep reinforcement learning framework: during the learning process, collect the experiences of heterogeneous multi-agents, input them into the action generation network designed in step S3, and obtain the action value through the corresponding value evaluation network. Update the value evaluation network according to the accuracy of the value evaluation, and update the action generation network according to the magnitude of the value of the action generated by the action generation network, so as to continuously iterate to achieve the update of the network, that is, the action learning of heterogeneous multi-agents.

[0134] The said step S6 includes the following steps:

[0135] S6.1. Based on the experience pools for autonomous vehicles and signal lights in step S5, regularly and uniformly extract a batch of experience samples from them to form two sequences containing a number of samples where b represents the sampling index; B represents the number of experience entries in one sampling;

[0136] S6.2. As described in step S3, during the training phase, it is necessary to input the samples extracted in step S6.1 into the action generation network and the value evaluation network to obtain the corresponding value estimation and policy gradient information; for the i-th autonomous vehicle, given the environmental observation data at time t+1, its action generation network predicts a lane-changing action probability distribution and an acceleration action; for the k-th signal light, given the environmental observation data at time t+1, its action generation network predicts the duration of the next phase; after obtaining all the actions of the autonomous vehicle and the signal light at time t+1, merge the environmental observation data at time t+1 with them and input them into the value evaluation network to obtain the action value of each agent at time t+1. At this time, the target action values of the autonomous vehicle and the signal light can be calculated according to the Bellman equation as follows:

[0137]

[0138] In the formula, γ is the discount factor, taking 0.95; and respectively represent the target action values of the i-th autonomous vehicle and the k-th signal light; and respectively represent the action values of the i-th autonomous driving vehicle and the k-th traffic light predicted by the value evaluation network at the (t + 1)-th moment;

[0139] S6.3. In step S6.2, the present invention respectively calculates the target action values for the traffic light and the autonomous driving vehicle. For the value evaluation network, the present invention defines the loss function for updating its own weight parameters as:

[0140]

[0141] In the formula, represents the loss of the value evaluation network at the t-th moment; and respectively represent the action values of the autonomous driving vehicle and the traffic light predicted by the action value network at the t-th moment; I and K respectively represent the total numbers of autonomous driving vehicles and traffic lights. At this time, the weight parameters ψ of the value evaluation network can be updated according to minimizing this prediction loss. The update optimizer selects the Adam optimizer, and the learning rate is 0.0001;

[0142] S6.4. After updating the weight parameters of the value evaluation network in step S6.3, input the environmental observation data of the autonomous driving vehicle and the traffic light at the t-th moment into their corresponding action generation networks to predict the agent actions in the environment at the t-th moment, and input the predicted actions and the environmental observation data at the t-th moment into the value evaluation network again to predict the action value at the t-th moment. Therefore, the loss of the action generation network of the autonomous driving vehicle can be obtained as:

[0143]

[0144] In the formula, represents the loss of the action generation network of the autonomous driving vehicle at the t-th moment; represents the lane-changing action probability distribution of the i-th autonomous driving vehicle at the t-th moment; and respectively represent the values of the lane-changing action and the acceleration action of the i-th autonomous driving vehicle predicted by the value evaluation network at the t-th moment; represents the loss of the traffic light action generation network at the t-th moment; represents the value of the phase duration action of the k-th traffic light at the t-th moment. It can be seen that whether it is the action generation network of the autonomous driving vehicle or the traffic light action generation network, their loss values are always optimized in the direction of increasing the action value. The present invention selects the Adam optimizer for all action generation networks to optimize the weight parameters, and the learning rate is set to 0.0001.

[0145] S7. Save the model parameters after learning is completed, load and deploy the heterogeneous multi-agent collaborative control model, and input the environmental observation data of each heterogeneous multi-agent to output the control quantities of all controlled objects;

[0146] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A complex traffic collaborative control method based on heterogeneous multi-agent reinforcement learning, characterized in that, Including the following steps: S1. Construct the environmental observation data of heterogeneous multi - agents based on a heterogeneous graph structure including autonomous driving vehicles, human - driven vehicles, and traffic lights; S2. Based on the heterogeneous multi - agent graph structure in step S1, design heterogeneous observation encoders for autonomous driving vehicles and traffic lights. Input the environmental observation data of heterogeneous multi - agents and output the encoded environmental observation data; S3. Design an action generation network and a corresponding value evaluation network for heterogeneous multi - agents. The action generation network outputs the actions of heterogeneous multi - agents according to the encoded environmental observation data, and the value evaluation network predicts the corresponding value based on the actions of heterogeneous multi - agents; The specific process of step S3 is as follows: S3.

1. Design a heterogeneous graph convolutional layer to further fuse the encoded heterogeneous graph features, so that the heterogeneous observation encoder in step S2 can be shared between the action generation network and the value evaluation network. The heterogeneous graph convolution process of the heterogeneous graph convolutional layer is expressed as follows: h i~k = HeteroGraphCon(Z i~k ) In the formula, HeteroGraphCon(·) represents the heterogeneous graph convolution process, and h i~k represents the shared high-dimensional feature representation after heterogeneous graph convolution; S3.

2. Design action generation networks for autonomous driving vehicles and traffic lights. Specifically, it can be described as: for autonomous driving vehicles, the actions include the probability distribution of acceleration and lane - changing actions; for traffic lights, the action can be the signal timing of the next stage. Use a multi - layer perceptron as the action generation model as follows: In the formula, represents the action of the i-th autonomous vehicle, specifically including the lane-changing action of the i-th autonomous vehicle and the acceleration action of the i-th autonomous vehicle where is a probability distribution quantity, and each element respectively represents the probability of each strategy in the strategy set {LCL, LK, LCR}. LCL, LK, and LCR in the strategy set respectively represent left lane change, no lane change, and right lane change; represents the action generation network mapping process of the i-th autonomous vehicle, φ i is the weight parameter of the action generation network of the i-th autonomous vehicle; is the signal timing of the next stage of the k-th traffic light set; represents the action generation network mapping process of the k traffic lights, φ k is the weight parameter of the action generation network of the k-th traffic light; S3.

3. Input the actions and state observations of all agents into the value evaluation network to obtain the action values corresponding to each agent. Its expression is as follows: where Q i~k represents the value brought by the actions of each autonomous vehicle and traffic signal; Q ψ (·) represents the mapping process that takes as input the globally shared feature h i~k after two - stage encoding, the actions of all autonomous vehicles and the actions of all traffic signals and outputs the action value of each agent; ψ is the weight parameter of the value evaluation network; S4. Construct a multi - level exploration and exploitation strategy suitable for the actions of heterogeneous multi - agents. The output exploration value is used to perform exploration regularization on the actions generated by the action generation network in step S3; S5. Design two experience pools for autonomous driving vehicles and traffic lights, and store the environmental observation data, actions, and rewards of the agents. Each heterogeneous agent outputs an action according to its action generation network. After the environment executes this action, new observation data and corresponding rewards are obtained, and the quadruple information at the t - th moment is stored in the experience pool; S6. Establish a heterogeneous multi - agent deep reinforcement learning framework. During the learning process, collect the experience data of heterogeneous multi - agents, input it into the action generation network designed in step S3, and obtain the action value through the corresponding value evaluation network. Update the value evaluation network according to the accuracy of the value evaluation, and update the action generation network according to the value of the actions generated by the action generation network, so as to continuously iterate to realize the update of the network, that is, the action learning of heterogeneous multi - agents; S7. Save the model parameters after learning is completed, load and deploy the heterogeneous multi - agent collaborative control model, input the environmental observation data of each heterogeneous agent, and output the control quantities of all controlled objects to realize the heterogeneous multi - agent collaborative control in the mixed traffic flow scenario.

2. The complex traffic collaborative control method based on heterogeneous multi-agent reinforcement learning according to claim 1, characterized in that Step S1 includes the following steps: S1.

1. The autonomous vehicle and the signal lamp identify the distance between the surrounding objects and the human-driven vehicle, and determine whether they should be linked to each other according to the distance threshold d threshold = 10m. If it is less than the distance threshold, it is considered that the autonomous vehicle or the signal lamp should be connected to the human-driven vehicle; otherwise, they are not connected. S1.

2. Convert the heterogeneous graph in step S1.1 into a mathematical expression: Use to represent the set of autonomous vehicle nodes, where represents the autonomous vehicle numbered 1 AD , 2 AD ... n AD , and n AD represents the number of autonomous vehicles; use to represent the set of human-driven vehicle nodes, where represents the human-driven vehicle numbered 1 HD , 2 HD ... n HD , and n HD represents the number of human-driven vehicles; use to represent the set of signal light nodes, where represents the signal light numbered 1 TL , 2 TL ... n TL ; Therefore, the overall node set can be expressed as V = V AD ∪V HD ∪V TL ; Therefore, for any three types of nodes in this heterogeneous graph, if the spatial Euclidean distance d ijk between any two nodes is less than the set distance threshold d threshold = 10m, then connect these two nodes in the heterogeneous graph structure and set the connection edge weight to 1; furthermore, obtain the adjacency matrix A and the feature matrix X for the heterogeneous graph. The element A ijk in the adjacency matrix A is as follows: In the formula, the subscripts ijk represent the i-th autonomous vehicle, the j-th human-driven vehicle, and the k-th traffic signal, and A ijk represents the connection situation among the i-th autonomous vehicle, the j-th human-driven vehicle, and the k-th traffic signal; and the heterogeneous feature matrix is obtained where is the feature information of the autonomous vehicle, and represent the x and y coordinates of the i-th autonomous vehicle; and represent the speed and acceleration of the i-th autonomous vehicle respectively; represents the lane where the i-th autonomous vehicle is located; is the feature information of the human-driven vehicle, and represent the x and y coordinates of the j-th human-driven vehicle; and represent the speed and acceleration of the j-th human-driven vehicle respectively; represents the lane where the j-th human-driven vehicle is located; where represents the current phase of the k-th traffic signal, represents the time required for the k-th traffic signal to switch to the next phase.

3. A complex traffic collaborative control method based on heterogeneous multi-agent reinforcement learning according to claim 1, characterized in that The specific process of step S2 is as follows: S2.

1. Since there are differences in attributes and observation dimensions between autonomous driving vehicles and traffic lights, design a heterogeneous encoder to extract different features for different types of nodes, so that it can uniformly output a high - dimensional representation for the subsequent decision - making network: Z i~k = Encoder θ (A,X) In the formula, Encoder θ (·) represents the heterogeneous encoder mapping process; θ is the parameter of the encoder; Z i~k is the encoded feature; S2.

2. Design multi-head self-attention graph convolution, ordinary graph convolution, heterogeneous graph convolution, and time series convolution modules as the main structure of the observation encoder for signal light type nodes; S2.

3. For autonomous vehicles, the encoder of autonomous vehicles is designed as a static single-step heterogeneous graph encoder, which inputs the environmental observation data of each autonomous vehicle at the current time step and outputs the encoded environmental observation data of all autonomous vehicles at the current time step.

4. A complex traffic collaborative control method based on heterogeneous multi-agent reinforcement learning according to claim 1, characterized in that For an autonomous vehicle, record the quadruple information at time t and store it in the experience pool of the autonomous vehicle, where and respectively represent the environmental observation data of the i-th autonomous vehicle at time t, the action generated by the action generation network, the obtained reward, and the environmental observation data of the i-th autonomous vehicle at time t+1; For the signal light, the quadruple information at the t-th moment is also recorded and stored in the experience pool of the signal light, where and respectively represent the environmental observation data of the k-th signal light at the t-th moment, the action generated by the action generation network, the obtained reward, and the environmental observation data of the k-th signal light at the (t + 1)-th moment.

5. A complex traffic collaborative control method based on heterogeneous multi-agent reinforcement learning according to claim 1, characterized in that, In step S4, three threshold parameters are defined and dynamically regulated through a multi-level exploration and exploitation strategy, enabling the heterogeneous agent to switch between the following three modes: 1) Random exploration: Completely random action sampling, indicating that the agent conducts action exploration; 2) Soft exploitation: Sampling actions according to the discrete action probability distribution output by the action generation network; 3) Hard exploitation: Directly selecting the action with the highest probability output by the action generation network; the above three modes are for the lane-changing actions of autonomous vehicles, representing the exploration and exploitation mechanisms for the acceleration actions of autonomous vehicles and the phase timing actions of signal lights; Moreover, a threshold is continuously attenuated through a noise attenuation strategy to make the continuous actions increasingly close to the actions generated by the action generation network. During the multi-heterogeneous agent reinforcement learning process, at each training step, the system obtains a random number r: When r > ε h it enters the random exploration mode. At this time, the lane-changing action of the autonomous vehicle will be randomly selected from left lane-changing, no lane-changing, and right lane-changing, and the acceleration action of the autonomous vehicle and the phase timing action of the traffic signal will also be randomly selected; When ε m <r ≤ ε h , enter the soft exploitation mode. At this time, the lane-changing action of the autonomous vehicle will be sampled and selected according to the probability distribution given by its action generation network. The acceleration action of the autonomous vehicle and the phase timing action of the traffic signal will be added with the noise value after the ε attenuation effect; When r ≤ ε m enter the hard utilization mode, at which time the lane-changing action of the autonomous vehicle will directly select the lane-changing action with the highest probability according to the probability distribution given by its action generation network, and the acceleration action of the autonomous vehicle and the phase timing action of the traffic signal will act on the environment directly without any modification; where ε h and ε m are two increasing thresholds respectively.

Citation Information

Patent Citations

  • Traffic signal lamp control method for multi-body heterogeneous road right distribution

    CN117523867A

  • Vehicle formation and signal lamp cooperative control method and system based on multi-agent reinforcement learning

    CN117671978A