Trajectory-based multi-agent cooperative enhanced traffic signal lamp control method
By using historical trajectories to predict future states and optimize the collaboration mechanism, the problem of insufficient adaptability of agents in dynamic environments is solved, and real-time response to dynamic environments and efficiency improvements in traffic signal control are achieved.
Patent Information
- Application Number
- CN202510264993.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-10
AI Technical Summary
The prior art is difficult to capture environmental changes in real time in dynamic environments, resulting in limited adaptability and performance of agents in complex or rapidly changing environments.
By using historical trajectories to predict future states, agents can make more forward-looking decisions based on these predictions, achieve real-time responses to dynamic environments, and optimize traffic signal control through collaborative mechanisms.
It enhances the dynamic adaptability of the decision-making of the agent, improves the efficiency and adaptability of traffic signal control, solves the problems brought about by centralized and distributed control, and improves network throughput and traffic flow scheduling efficiency.
Smart Images

Figure CN120126331A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of communication signal control, and in particular to a method for multi-agent collaborative reinforcement control of traffic lights based on trajectories. Background Art
[0002] With the rapid economic development and continuous improvement of living standards, the number of motor vehicles has shown a significant growth trend. The increase in vehicles has a positive impact on improving personal travel efficiency, promoting regional economic development, and driving industrial growth. However, this growth has also brought a series of traffic problems such as traffic congestion, decreased travel efficiency, and environmental pollution. Especially in the urban central area, the contradiction between limited traffic infrastructure and the increasing number of vehicles is becoming increasingly acute, which poses higher requirements for the carrying capacity and management efficiency of the traffic system. How to optimize the vehicle scheduling strategy to maximize the throughput of the traffic network under the given traffic infrastructure constraints is a challenge. Reasonable traffic signal control can effectively guide vehicles to pass through intersections smoothly, reduce the residence time of vehicles at intersections, and thus improve the operation efficiency of the entire traffic network.
[0003] Artificial intelligence is an interdisciplinary research field that aims to create intelligent systems capable of performing complex tasks that typically require human intelligence, such as learning, reasoning, perception, language understanding, and decision-making. The development of AI is designed to simulate human cognitive processes, enhance the autonomy and adaptability of machines, enabling them to operate effectively in dynamic and uncertain environments. In the research and practice of AI, three main learning paradigms are adopted to build intelligent systems: supervised learning, unsupervised learning, and reinforcement learning. Among them, deep reinforcement learning is an advanced algorithm framework that combines the theories of deep learning and reinforcement learning. This framework endows agents with the ability to learn optimal strategies through trial and error in complex environments by leveraging the powerful function approximation capabilities of deep neural networks. The core advantages of DRL lie in its ability to handle high-dimensional, unstructured data and its adaptability in dynamic environments, which enables it to demonstrate broad application potential in complex tasks such as path optimization. In future vehicle networking scenarios, due to the high requirements for real-time performance and low latency, each traffic intersection will be equipped with a reinforcement learning server, which serves as an intelligent node responsible for collecting and processing real-time traffic data from the roads connected to it. Using the current traffic conditions, the server will control the signals at the current intersection. As urban traffic networks continue to expand, centralized reinforcement learning methods face challenges of a rapidly growing state space and action space. Moreover, in centralized methods, all decisions are processed by a central agent, which may lead to bottlenecks in computing resources. Distributed methods allow each agent to learn and make decisions independently, thus effectively scaling to larger systems. Distributed methods improve computational efficiency by distributing the computational burden among multiple agents. In the scenario of distributed multi-agent reinforcement learning, agents make decisions relying only on local observations, which may result in (1) getting trapped in local optimal solutions rather than global optimal solutions and (2) difficulty in converging to a stable strategy, leading to a slow learning process.
[0004] Collaborative reinforcement learning, as an emerging learning paradigm, its core idea is to achieve more efficient handling of complex tasks and environments through collaboration and knowledge sharing among agents.
[0005] However, previous methods rely on offline data to evaluate trajectory similarity and may show limitations in dynamic environments. Such methods cannot capture the immediate changes in the environment in real time, resulting in a lack of necessary dynamic adaptability in the decision-making process. Due to its dependence on pre-collected trajectory data, this method has inherent deficiencies in responding to environmental changes in real time, which limits the adaptability and performance of agents in complex or rapidly changing environments. Summary of the Invention
[0006] The object of the present invention is to overcome the above-mentioned defects in the prior art and provide a method for collaborative reinforcement control of traffic lights based on trajectories that can adapt to real-time traffic environment changes. By predicting future states based on historical trajectories, the agents can make more forward-looking decisions based on these predictions, achieve real-time response to dynamic environments, and optimize traffic signal control through a collaboration mechanism.
[0007] To achieve the above object, the present invention adopts the following technical solutions:
[0008] A method for collaborative reinforcement control of traffic lights based on trajectories, comprising the following steps:
[0009] S0. Equip a reinforcement learning server at each traffic intersection node. This server serves as an agent and is responsible for collecting and processing real-time traffic data on the roads connected to it.
[0010] S1. Extract the historical trajectory states of each element i in the agent and generate a time series
[0011] S2. Through a multi-layer perceptron, convert the time series into a state embedding
[0012] S3. Through positional encoding, associate the state embedding of each element of the agent with its position in the sequence to obtain a feed-forward neural network. In the feed-forward neural network, the state embeddings of two elements at adjacent positions are respectively denoted as
[0013] S4. Through the self-attention mechanism, calculate the attention weights of each element in the input sequence to all other elements. After normalizing the attention weights of each element, generate a weighted α i,j of the sequence and form a weight matrix W;
[0014] S5. Use a graph neural network model to capture the spatial features of different nodes, and use the graph attention mechanism to map the feature vectors of the nodes to a new feature space. The state embedding of the node in the new feature space is denoted as to obtain the trajectory of the node. By comparing and learning the trajectory similarities between different nodes, trajectory prediction and traffic signal control are carried out.
[0015] In the present invention, historical trajectories are utilized to predict future trajectories, thereby achieving real-time response to a dynamic environment and enhancing the decision-making dynamic adaptability of an agent. By analyzing historical trajectory data, the agent can learn the dynamic characteristics of the environment and, based on this, predict possible future trajectories, and then make more accurate decisions.
[0016] Capturing the temporal characteristics of trajectories is the key to improving prediction accuracy. To this end, this application adopts a Transformer model (including a self-attention mechanism and a feed-forward neural network), which is an architecture with significant advantages in processing sequential data. The Transformer model can efficiently capture long-term dependencies through its self-attention mechanism, which is particularly important for understanding and predicting the temporal dynamic characteristics in trajectories. This method can not only process complex dynamic interaction information but also process data in parallel, significantly improving the computational efficiency. This application uses the Transformer model to encode historical trajectory data, thereby extracting key features in the time series and predicting future trajectories.
[0017] (1) First, the historical trajectory of the agent is defined as a time series of agent states:
[0018]
[0019] where, is the state of agent i at time step t, and T obs is the set observation window size, which determines the length of time to look back when considering the historical trajectory. This definition allows for the analysis of the agent's behavior pattern from the time dimension and provides a basis for predicting its future trajectory.
[0020] (2) Second, the historical trajectory state of the agent is input into a multi-layer perceptron to obtain an embedded representation of the state. The concept of state embedding is widely used to extract and represent the features of a state. State embedding is achieved by mapping the state space into an embedding space so that machine learning models can more effectively process and learn the relationships between states. State embedding is widely used to extract and represent the features of a state. State embedding maps a discrete or high-dimensional state space into a continuous low-dimensional space, enabling more effective processing and learning of the relationships between states. Specifically, this application inputs the historical trajectory state of the agent into a multi-layer perceptron to obtain an embedded representation of the state. This process can be formalized as:
[0021]
[0022] (3) Next, positional encoding is an essential part of the Transformer architecture. By injecting the position information of each element in the sequence into the model, the model can capture the sequential relationship of the elements in the sequence and obtain a feed-forward neural network. Specifically, in this application, the embedding of the historical state is combined with the position information, and each state embedding is associated with its position in the sequence through positional encoding, so that the Transformer model can not only understand the semantic information of the state, but also identify the relative position relationship between states.
[0023] The specific relative position relationship can be expressed by the following formula:
[0024]
[0025] where i is the dimension.
[0026] (4) The core structure of the Transformer consists of a self-attention mechanism and a feed-forward neural network. The self-attention mechanism allows the model to consider all positions in the sequence simultaneously when processing each element in the sequence, thereby capturing long-range dependencies within the sequence. This mechanism is achieved by calculating the attention weights of each element in the sequence to all other elements, and these weights reflect the mutual relationship between different elements. In the self-attention layer, each input sequence is mapped to three vector sets: query, key, and value. The query vector is used to compare with all key vectors to determine the attention weights of each element to other elements. This comparison is usually achieved through a dot product operation, followed by normalization using the softmax function to ensure that the sum of all weights is 1. The normalized weights are used to weighted sum the corresponding value vectors to generate the weighted sequence, expressed as follows:
[0027]
[0028] where Q is the query vector, K is the key vector, V is the value vector, and d k is the dimension of the vector.
[0029] The self-attention mechanism focuses on capturing the mutual relationship between elements within the same sequence, revealing the complex dependency structure within the sequence. The encoder-decoder attention mechanism focuses on modeling the correspondence between the historical sequence (the output of the encoder) and the predicted sequence (the input of the decoder). In sequence-to-sequence tasks, the encoder is responsible for encoding the input sequence into a high-dimensional, compact representation, and the decoder uses this representation to gradually construct the output sequence. The encoder-decoder attention mechanism allows the decoder to dynamically focus on the most relevant part of the input sequence at each step of generating the output sequence, thereby ensuring the semantic consistency and alignment between the output sequence and the input sequence. Specifically, the mechanism uses the output of the encoder as the key and value vectors, which contain rich semantic information of the input sequence. At the same time, the output of the previous self-attention layer in the decoder is used as the query vector, which represents the current state and requirements of the decoder at a specific decoding step. This mechanism not only enhances the model's understanding of the input sequence, but also ensures that the generation of the output sequence can make full use of the contextual information provided by the encoder, thereby improving the performance of sequence-to-sequence tasks.
[0030] After the self-attention mechanism and encoder-decoder attention mechanism of each encoder and decoder, the feed-forward neural network comes into play. It extracts the complex features of the input data through linear transformation and activation function, which helps the model to better understand the data.
[0031] At this point, the analysis on the time dimension is completed.
[0032] (5) Analysis of the temporal dimension is important because it involves the state changes of a single agent over time. However, the mutual influence between intersections cannot be ignored, and capturing the relationship between intersections is crucial for predicting trajectories. This is because traffic is a complex dynamic system in which the state of an individual is affected not only by its own behavior, but also by the surrounding environment and the behavior of other individuals. Therefore, when making trajectory predictions, considering the relationship between intersections helps to more accurately simulate the dynamic characteristics of traffic flow, thereby improving the accuracy and reliability of predictions. This relationship capture can be achieved through methods such as graph neural networks, which are able to capture spatial features. In this way, the model can not only understand the movement patterns of individuals over time, but also take into account the interactions between individuals, which is particularly important for trajectory prediction.
[0033] Graph neural networks are a type of deep learning model specialized in processing graph-structured data. By combining the advantages of graph computing and neural networks, they can effectively capture the graph structure and abstract node features. Graph neural networks allow node features to spread in the graph through a message-passing mechanism, enabling the representation of each node to contain information about its neighbor nodes. The graph attention mechanism assigns different weights to each node, enabling the model to focus on more important nodes in the graph and thus obtain a richer feature representation.
[0034] First, through feature mapping, the feature vector of a node is mapped to a new feature space to prepare for subsequent similarity calculations. Next, in the preliminary calculation step of the attention scores, the model evaluates the preliminary attention scores between each pair of nodes, which is usually based on their feature similarity. Then, through the normalization step, the model adjusts these scores to ensure that the sum of the attention of each node to its neighbors is 1, thus reflecting the relative importance of the neighbor nodes. Finally, in the feature aggregation and update step, the model combines the normalized attention coefficients and the features of the neighbor nodes to update the feature representation of each node to capture local and global structural information in the graph.
[0035]
[0036] sim i,j = a(Wh i,t , Wh j,t );
[0037]
[0038] where W is the weight matrix, a is a single-layer feedforward neural network, N i is the neighbor node of intersection i, and σ is the activation function.
[0039] Next, the prediction of the trajectory is specifically elaborated:
[0040] A trajectory is a sequence of states that are successively visited during the decision-making process. The trajectory is represented as tra i = <s 0 , s 1 ,..., s m >, where s i is the state passing through the intersection.
[0041] First, calculate the state distance. The Markov distance is a measure to quantify the similarity between states in a Markov decision process. Specifically, this distance is defined based on the state transition probability and the immediate reward function, by quantifying the differences in state transitions and rewards between states. This measure not only captures the similarity of the state transition structure but also takes into account the similarity of the immediate rewards associated with states, thus providing a quantitative framework for understanding and comparing different Markovs. The distance between states can be expressed as:
[0042]
[0043] where α ∈ (0, 1) is the weight coefficient, used to balance the relative importance between the reward difference and the state transition probability distribution difference. represents the state transition probability of selecting action a in state s, and T K (d) is the Kantorovich distance.
[0044] Combining the Markov distance and metric learning, this application uses a multi-layer perceptron to achieve the embedded representation of states. Specifically, this application trains an MLP to learn an embedding function that encodes the states in MDPs into low-dimensional vectors, such that the distance between these vectors in the embedding space is equal to their Markov distance in the original MDPs. This embedding technique aims to capture the state transition probability and reward differences between states through the non-linear transformation ability of the neural network, thus maintaining the complex similarity relationships between states in the low-dimensional space.
[0045] Since there are no exactly identical cases between trajectories, local features may be shrunk or enlarged, making it difficult to find the most suitable learning object in the current state. To solve this problem, dynamic time warping (DTW) is introduced for calculating the trajectory similarity. DTW allows for local stretching or compression of trajectories on the time axis through dynamic programming, so that the two trajectories are as consistent as possible in shape. This method can effectively capture the local changes and distortions in the trajectories, providing a more accurate similarity measure. In this way, DTW can better identify and match the similar patterns in the trajectories, even if these patterns are different in time. Therefore, the calculation formulas for DTW and the trajectory distance are respectively expressed as:
[0046] Score(tra i ,tra j ) = D(s i,n , s j,n )
[0047] D(s i,k , s j,l ) = d(s i,k , s j,l) + min(D(s i,k-1 , s j,l ), D(s i,k , s j,l-1 ), D(s i,k-1 , s j,l-1 ))
[0048] where s i,k and s j,l are the states of agent i and agent j at time steps k and l, respectively.
[0049] Through this method, the present application can identify similar trajectories. The long short-term memory network is a special recurrent neural network structure that can effectively capture long-distance dependencies in sequential data. When processing trajectory data, the LSTM can aggregate the state embeddings in the trajectory through its unique gating mechanism to obtain a comprehensive sequential embedding representation. This sequential embedding not only contains the local features of the trajectory but also incorporates global context information, providing a rich feature space for trajectory similarity analysis. Similar to the state embedding, the present application equates the distance of the vector in the embedding space with the distance of the trajectory. In the LSTM cell, three main gating structures, the input gate, the forget gate, and the output gate, work together to determine when to update the cell state, which is the memory core of the LSTM cell. The input gate is responsible for deciding what new information to write, the forget gate determines which old information needs to be forgotten, and the output gate controls the influence of the current state on the next time step. The decisions of these gates are based on the sigmoid function, which outputs a value between 0 and 1 indicating the degree of information passage. The cell state can be expressed as:
[0050]
[0051] where c t represents the cell state at time step t, f t and i t represent the activation vectors of the forget gate and the input gate, respectively, W c and U c represent the weight matrices of the input gate and the hidden state, respectively, x t represents the input vector, b c represents the bias vector, ° represents the Hadamard product, and tanh is the hyperbolic tangent activation function used to introduce non-linearity.
[0052] The hidden state h t of the LSTM cell, which is passed to the next LSTM cell or used for output prediction, is controlled by the output gate and calculated as follows:
[0053]
[0054] Among them, o t is the activation vector of the output gate.
[0055] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0056] (1) The present invention proposes a collaborative control method of collaborative reinforcement learning for large-scale distributed traffic signal control scenarios, effectively solving the problems brought by centralized and distributed control, and improving network throughput and traffic flow scheduling efficiency.
[0057] (2) The present invention first predicts future trajectories based on historical trajectories and spatio-temporal states of neighbor nodes. Then, the states of the trajectories are mapped to the embedding space, and the long short-term memory network is further used to aggregate information. The distance in the embedding space is used as an index to measure the similarity of trajectories.
[0058] (3) Based on similar trajectory pairs, the present invention further proposes a collaborative reinforcement learning algorithm. This algorithm realizes collaborative learning based on knowledge cooperation between agents by efficiently transferring the experiences between similar agents, and improves the training efficiency and generalization ability of the model. Description of the Drawings
[0059] Figure 1 is the traffic network instance diagram of Embodiment 1 of the present invention;
[0060] Figure 2 is the trajectory prediction architecture diagram of Embodiment 1 of the present invention;
[0061] Figure 3 is the trajectory embedding diagram of Embodiment 1 of the present invention. Detailed Embodiments
[0062] To better illustrate the purpose, technical solutions and advantages of the present invention, the present invention will be further described below in conjunction with specific embodiments and drawings, but the embodiments do not impose any form of limitation on the present invention. Unless otherwise specified, the reagents, methods and equipment used in the present invention are conventional reagents, methods and equipment in the technical field. Unless otherwise specified, the reagents and materials used in the present invention are commercially available.
[0063] Embodiment 1
[0064] This embodiment provides a method for multi-agent collaborative reinforcement control of traffic lights based on trajectories, including the following steps:
[0065] As Figure 1 shown: The traffic network is represented by a directed graph G=(V, E), where v∈V represents intersections, and e v,u ∈E represents the roads between intersections. Each intersection v is managed by an agent, and its neighbor set is defined as NBv ={u|e v,u ∈E}.
[0066] In order to obtain real-time dynamic status information of the traffic network, the system deploys advanced traffic sensors at various intersections. These sensors can collect a variety of traffic data, including key indicators such as vehicle waiting time and traffic density. These data provide the agent with real-time feedback on the current traffic conditions, enabling it to dynamically adjust traffic signal control strategies according to actual traffic needs.
[0067] Each intersection is equipped with an entry ramp and an exit ramp, which are channels for vehicles to enter and leave the intersection. The entry ramp allows vehicles to enter the intersection, while the exit ramp is responsible for guiding vehicles to leave. These ramps are composed of multiple lanes, where the set of all entry lanes of an intersection v is represented as Lane[v]. Traffic movement refers to the process of vehicles transferring from an entry lane to an exit lane in an intersection, which is directly affected by traffic signal control. Reasonable traffic signal control can effectively guide vehicles to pass through the intersection smoothly, reduce the time vehicles stay at the intersection, and thus improve the operation efficiency of the entire traffic network. In addition, the phase in traffic signal control refers to the combination of green light releases set for different traffic flows. A typical intersection may contain four phases: north-south straight (NSS), north-south left turn (NSL), east-west straight (EWS) and east-west left turn (EWL). The agent needs to reasonably set the current phase according to the current traffic conditions and traffic signal control strategy.
[0068] Partially Observable Markov Decision Process (Dec-POMDP) is a framework extended from the classical Markov Decision Process (MDP), which is expressed as<N,S,O,A,R,P,π> , applicable to multi-agent systems, in which each agent can only observe part of the environment state and can only control its own actions. In this model, each agent must formulate a strategy based on its local observations and common model knowledge of the environment to maximize its cumulative reward or meet specific performance criteria. Where N is the number of agents, and the state S of the system is the set of local states of all agents. O is the observation space, which represents the set of all possible observations that the agent can observe. Since the environment is partially observable, the agent cannot directly obtain the current true state, but infers the state through observation. Each agent i can only observe its corresponding local state o i , through its action a iAffect the environment. The action set A of the agent is the set of all agent actions, and the choice of action directly affects the state transition and the result of the observation. The reward function is represented by R = S × A, indicating the immediate reward obtained by taking an action in a state. The reward function is the goal of the agent's decision optimization and is usually used to measure the quality of the decision. The transition probability P(S t+1 = s'|S t = s,A t = a) describes the probability that the system transitions to a new state s' given the current state s and action a. The goal of each agent is to find a policy π that can maximize its expected cumulative reward given its local observation.
[0069] In this embodiment, each intersection is a DRL agent that makes decisions based on its local observation. The observation vector of the intersection consists of multiple components, denoted as where, represents the current phase of intersection i, represents the queue length of the incoming lane j. These observation vectors provide the agent with detailed information about the current traffic conditions, enabling it to make decisions based on real-time data. According to the current observation of the intersection, the intersection can perform actions that is, select different traffic signal control phases to adjust the traffic flow. Traffic signal management aims to improve the operating efficiency of the traffic network. Therefore, in this application, the reward function is set to that is, to motivate the agent by reducing the queue length of vehicles at the intersection. This reward mechanism encourages the agent to adopt strategies that can effectively alleviate traffic congestion and improve road capacity. The value function is a function that evaluates the expected future reward of a state or state-action pair. Its core role is to help the agent evaluate the long-term benefits of being in a specific state or taking a specific action, thereby guiding the agent to choose the optimal strategy. The action value function can be expressed as:
[0070]
[0071] where γ is the discount factor, which weights future rewards according to the distance in time, enabling the agent to balance current and future rewards when making decisions. The update of the action value function is achieved through the following iterative process:
[0072] Q(s,a) = Q(s,a) + α[r + γmax a' Q(s',a') - Q(s,a)]
[0073] where α is the learning rate, which determines the magnitude of each update.
[0074] Next, the historical trajectory is used to predict the future trajectory, so as to achieve real-time response to the dynamic environment and enhance the decision-making dynamic adaptability of the agent. By analyzing the historical trajectory data, the agent can learn the dynamic characteristics of the environment, predict the possible future trajectories based on this, and then make more accurate decisions.
[0075] (1) In the time dimension
[0076] Capturing the time characteristics of the trajectory is the key to improving the prediction accuracy. For this purpose, this application adopts the Transformer model, which is an architecture with significant advantages in processing sequence data. The Transformer model can efficiently capture long-term dependencies through its self-attention mechanism, which is particularly important for understanding and predicting the time dynamic characteristics in the trajectory. This method can not only process complex dynamic interaction information but also process data in parallel, significantly improving the computational efficiency. This application uses the Transformer model to encode the historical trajectory data, extract the key features in the time series, and predict the future trajectory.
[0077] First, this application defines the historical trajectory of the agent as a time series of agent states:
[0078]
[0079] where is the state of agent i at time step t, and T obs is the set observation window size, which determines the time length of the historical trajectory we consider when looking back. This definition analyzes the behavior pattern of the agent from the time dimension and provides a basis for predicting its future trajectory.
[0080] This application inputs the historical trajectory state of the agent into a multi-layer perceptron to obtain an embedded representation of the state. The concept of state embedding is widely used to extract and represent the features of the state. State embedding is achieved by mapping the state space to the embedding space so that the machine learning model can more effectively process and learn the relationships between states.
[0081] State embedding is widely used to extract and represent the features of the state. State embedding maps the discrete or high-dimensional state space to a continuous low-dimensional space, which can more effectively process and learn the relationships between states. Specifically, this application inputs the historical trajectory state of the agent into a multi-layer perceptron to obtain an embedded representation of the state. This process can be formalized as:
[0082]
[0083] Positional encoding is an essential part of the Transformer architecture. By injecting the position information of each element in the sequence into the model, the model can capture the sequential relationship of the elements in the sequence and obtain a feed-forward neural network. Specifically, the embedding of the historical state is combined with the position information, and each state embedding is associated with its position in the sequence through positional encoding, so that the Transformer model can not only understand the semantic information of the state, but also identify the relative position relationship between states.
[0084] The specific relative position relationship can be expressed by the following formula:
[0085]
[0086] where i is the dimension.
[0087] The core structure of Transformer consists of a self-attention mechanism and a feed-forward neural network. The self-attention mechanism allows the model to consider all positions in the sequence simultaneously when processing each element in the sequence, so as to capture the long-range dependencies within the sequence. This mechanism is achieved by calculating the attention weights of each element in the sequence to all other elements, and these weights reflect the mutual relationships between different elements. In the self-attention layer, each input sequence is mapped to three vector sets: query, key, and value. The query vector is used to compare with all key vectors to determine the attention weights of each element to other elements. This comparison is usually achieved through a dot product operation, and then normalized by the softmax function to ensure that the sum of all weights is 1. The normalized weights are used to weighted sum the corresponding value vectors to generate the weighted sequence, which is expressed as follows:
[0088]
[0089] where Q is the query vector, K is the key vector, V is the value vector, and d k is the dimension of the vector.
[0090] The self-attention mechanism focuses on capturing the mutual relationship between elements within the same sequence, revealing the complex dependency structure within the sequence. The encoder-decoder attention mechanism focuses on modeling the correspondence between the historical sequence (the output of the encoder) and the predicted sequence (the input of the decoder). In sequence-to-sequence tasks, the encoder is responsible for encoding the input sequence into a high-dimensional, compact representation, and the decoder uses this representation to gradually construct the output sequence. The encoder-decoder attention mechanism allows the decoder to dynamically focus on the most relevant part of the input sequence at each step of generating the output sequence, thereby ensuring the semantic consistency and alignment between the output sequence and the input sequence. Specifically, the mechanism uses the output of the encoder as the key and value vectors, which contain rich semantic information of the input sequence. At the same time, the output of the previous self-attention layer in the decoder is used as the query vector, which represents the current state and requirements of the decoder at a specific decoding step. This mechanism not only enhances the model's understanding of the input sequence, but also ensures that the generation of the output sequence can make full use of the contextual information provided by the encoder, thereby improving the performance of sequence-to-sequence tasks.
[0091] After the self-attention mechanism and encoder-decoder attention mechanism of each encoder and decoder, the feed-forward neural network comes into play. It extracts the complex features of the input data through linear transformation and activation function, which helps the model to better understand the data.
[0092] (II) Spatial Dimension
[0093] The analysis of the time dimension is certainly important because it involves the state changes of a single agent over time. However, the mutual influence between intersections cannot be ignored, and capturing the relationship between intersections is crucial for predicting trajectories. This is because traffic is a complex dynamic system in which the state of an individual is affected not only by its own behavior, but also by the surrounding environment and the behavior of other individuals. Therefore, when making trajectory predictions, considering the relationship between intersections helps to more accurately simulate the dynamic characteristics of traffic flow, thereby improving the accuracy and reliability of predictions. This relationship capture can be achieved through methods such as graph neural networks, which are able to capture spatial features. In this way, the model is able to not only understand the movement patterns of individuals over time, but also take into account the interactions between individuals, which is particularly important for trajectory prediction.
[0094] Graph neural networks are a type of deep learning model specialized in processing graph-structured data. By combining the advantages of graph computing and neural networks, they can effectively capture the graph structure and abstract node features. Graph neural networks allow node features to propagate in the graph through a message-passing mechanism, enabling the representation of each node to contain information about its neighbor nodes. The graph attention mechanism assigns different weights to each node, enabling the model to focus on more important nodes in the graph and thus obtain richer feature representations.
[0095] First, through feature mapping, the feature vectors of nodes are mapped to a new feature space to prepare for subsequent similarity calculations. Next, in the preliminary calculation step of attention scores, the model evaluates the preliminary attention scores between each pair of nodes, which is usually based on their feature similarity. Then, through the normalization step, the model adjusts these scores to ensure that the sum of the attention of each node to its neighbors is 1, thus reflecting the relative importance of neighbor nodes. Finally, in the feature aggregation and update step, the model combines the normalized attention coefficients and the features of neighbor nodes to update the feature representation of each node to capture local and global structural information in the graph.
[0096]
[0097] sim i,j = a(Wh i,t , Wh j,t );
[0098]
[0099] where W is the weight matrix, a is a single-layer feedforward neural network, N i is the neighbor node of intersection i, and σ is the activation function.
[0100] Next, the prediction of the trajectory is elaborated specifically:
[0101] (III) Trajectory similarity analysis (see the schematic diagram in Figure 3 )
[0102] A trajectory is a sequence of states that are accessed sequentially during the decision-making process. The trajectory is represented as tra i = <s 0 , s 1 ,..., s m >>, where s i is the state passing through the intersection.
[0103] First, calculate the state distance. The Markov distance is a measure to quantify the similarity between states in a Markov decision process. Specifically, this distance is defined based on the state transition probability and the immediate reward function, by quantifying the differences in state transitions and rewards between states. This measure not only captures the similarity of the state transition structure but also takes into account the similarity of the immediate rewards associated with states, thus providing a quantitative framework for understanding and comparing different Markovs. The distance between states can be expressed as:
[0104]
[0105] where α ∈ (0, 1) is the weight coefficient, used to balance the relative importance between the reward difference and the state transition probability distribution difference. represents the state transition probability of selecting action a in state s, and T K (d) is the Kantorovich distance.
[0106] Combining the Markov distance and metric learning, this application uses a multi-layer perceptron to achieve the embedded representation of states. Specifically, an MLP is trained to learn an embedding function that encodes the states in MDPs as low-dimensional vectors, such that the distance between these vectors in the embedding space is equal to their Markov distance in the original MDPs. This embedding technique aims to capture the state transition probability and reward differences between states through the non-linear transformation ability of the neural network, thus maintaining the complex similarity relationships between states in the low-dimensional space.
[0107] Since there are no exactly identical cases between trajectories, local features may be shrunk or enlarged, making it difficult to find the most suitable learning object in the current state. To solve this problem, dynamic time warping (DTW) is introduced for calculating the trajectory similarity. DTW allows for local stretching or compression of trajectories on the time axis through dynamic programming methods, so that the two trajectories are as consistent as possible in shape. This method can effectively capture the local changes and distortions in trajectories, providing a more accurate similarity measure. In this way, DTW can better identify and match similar patterns in trajectories, even if these patterns differ in time. Therefore, the calculation formulas for DTW and the trajectory distance are respectively expressed as:
[0108] Score(tra i ,tra j ) = D(s i,n , s j,n )
[0109] D(s i,k , s j,l ) = d(s i,k , s j,l) + min(D(s i,k-1 , s j,l ), D(s i,k , s j,l-1 ), D(s i,k-1 , s j,l-1 ))
[0110] where s i,k and s j,l are the states of agent i and agent j at time steps k and l, respectively.
[0111] In this way, the present application can identify similar trajectories. The long short-term memory network is a special type of recurrent neural network structure that can effectively capture long-range dependencies in sequential data. When processing trajectory data, the LSTM can aggregate the state embeddings in the trajectory through its unique gating mechanism to obtain a comprehensive sequential embedding representation. This sequential embedding not only contains the local features of the trajectory but also incorporates the global context information, providing a rich feature space for trajectory similarity analysis. Similar to the state embedding, we equate the distance of the vector in the embedding space with the distance of the trajectory. In the LSTM cell, three main gating structures, the input gate, the forget gate, and the output gate, work together to determine when to update the cell state, which is the memory core of the LSTM cell. The input gate is responsible for deciding what new information to write, the forget gate determines which old information needs to be forgotten, and the output gate controls the influence of the current state on the next time step. The decisions of these gates are based on the sigmoid function, which outputs a value between 0 and 1 indicating the degree of information passage. The cell state can be expressed as:
[0112]
[0113] where c t represents the cell state at time step t, f t and i t represent the activation vectors of the forget gate and the input gate, respectively, W c and U c represent the weight matrices of the input gate and the hidden state, respectively, x t represents the input vector, b c represents the bias vector, denotes the Hadamard product, and tanh is the hyperbolic tangent activation function used to introduce non-linearity.
[0114] The hidden state h t of the LSTM cell, which is passed to the next LSTM cell or used for output prediction, is controlled by the output gate and calculated as follows:
[0115]
[0116] Among them, o t is the activation vector of the output gate.
[0117] (IV) Signal Transmission
[0118] First, preprocess the trajectories in the source domain, and split the continuous trajectory dataset into sequence segments of a fixed length L. This step aims to extract local patterns and dynamic changes within a specific time window from the original trajectories. Subsequently, these trajectory sequences of length L are input into the LSTM to obtain the embedded representation of the trajectories.
[0119] Historical trajectories may not be sufficient to provide enough information for making accurate decisions. Predictive trajectories provide an estimate of future states, enabling the agent to make more forward-looking decisions based on these predictions. This forward-looking nature cannot be provided solely by relying on historical trajectories. Predictive trajectories in reinforcement learning can help the agent with long-term planning. Also, predictive trajectories can help the agent quickly adapt to new environments. To predict the future development of the current trajectory, the last L - K states in the sequence are used as the input sequence to predict the subsequent K states. Similarly, this trajectory sequence of length L is input into the LSTM to obtain the embedded representation of the trajectories. Further, by calculating the distance between the source domain trajectory embedding and the predictive trajectory embedding, the trajectory embedding with a smaller distance is selected as the learning object, and the same action as in the source domain is selected at the current state.
[0120] The experiences generated during the interaction between the agent and the environment provide the necessary feedback information for the agent, enabling it to adjust its strategy according to the immediate state of the environment. Therefore, the agent not only needs to learn and optimize its action strategy in the current state but also should effectively transfer the experiences accumulated in the source domain to the target domain. This cross-domain experience transfer can significantly improve the adaptability and learning efficiency of the agent in the new environment because it allows the agent to utilize existing knowledge to guide the exploration and decision-making processes in the new environment.
[0121] Trajectories in the source domain are usually formed according to the optimal policy in that environment. Therefore, when conducting knowledge collaboration, methods that use the magnitude of action rewards or action value functions as the basis for sampling probabilities may not be applicable. Instead, the similarity between the source domain and the target domain should be considered as the main basis for sampling probabilities. By evaluating the similarity between the experiences in the source domain and the target domain, samples that are helpful for learning the target domain policy can be selected more effectively, thereby improving the effect of transfer learning. Specifically, compare the similarities of trajectories of length L in the source domain and the target domain, where the probability of each sample being sampled is proportional to its similarity. Such a method helps to quickly converge to an effective policy in the target domain. Especially when there are significant differences between the source domain and the target domain, similarity becomes a key transfer learning metric. In this way, the reinforcement learning algorithm can utilize the experiences in the source domain more precisely to guide the policy optimization process in the target domain.
[0122]
[0123] Among them, d(i,tar) is the similarity of the current trajectory, and tra sim is the trajectory similar to the target domain trajectory.
[0124] The temporal difference error is also an important basis for evaluating the sampling probability. TD Error quantifies the difference between the current state value prediction and the actually observed return, and its formula is
[0125] δ = R t + γV(S t+1 ) - V(S t )
[0126] Among them, V(S t ) is the value function of the state, and γ is the discount function. Compared with the transitions with smaller TD Error, the transitions with larger TD Error provide more information about the environment. Therefore, they should be sampled more frequently to update the agent's policy. Therefore, the sampling probability based on the temporal difference error is
[0127]
[0128] To achieve a more effective experience sampling strategy, comprehensively consider the similarity of trajectories and the temporal difference error of states as the basis for calculating the sampling probability. Therefore, the sampling probability of the experience is
[0129] p(i) = p sim (i) * p TD (i)
[0130] To ensure that the sum of the sampling probabilities of all experiences in experience replay is 1, that is, to satisfy the basic properties of probability distribution, the softmax function is used to normalize the sampling probabilities.
[0131] As the number of training episodes increases, the model gradually shifts from relying on source domain data to making more use of target domain data for training. This process helps the model better capture the inherent distribution and structural characteristics of the target domain data. Especially when there are significant differences between the source domain and the target domain, gradually reducing the learning from the source domain can effectively improve its performance on the target task. The probability of learning from the source domain is set to
[0132]
[0133] where x is the current training round and tot is the total number of training rounds.
[0134] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the protection scope of the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the essence and scope of the technical solutions of the present invention.
Claims
1. A method for enhancing control of traffic lights by multi-agent collaboration based on trajectories, characterized in that: The following steps are involved: S0. Each traffic intersection node is equipped with a reinforcement learning server, which acts as an intelligent agent and is responsible for collecting and processing real-time traffic data of the roads connected to it; S1. Extract the historical trajectory state of each element i in the agent And generate time series S2. Through multi-layer perceptron, the time series Convert to state embedding S3. Embed the state of each element of the agent through position encoding Associated with their positions in the sequence, a feedforward neural network is obtained, in which the state embeddings of two elements at adjacent positions are recorded as S4. Through the self-attention mechanism, the attention weight of each element in the input sequence to all other elements is calculated. After the attention weight of each element is normalized, the weighted α of the sequence is generated. i,j , and form a weight matrix W; S5. Use the graph neural network model to capture the spatial features of different nodes, use the graph attention mechanism to map the feature vector of the node to the new feature space, and the state embedding of the node in the new feature space is expressed as The trajectory of the node can be obtained, and the trajectory prediction and traffic light control can be performed by comparing and learning the trajectory similarity between different nodes.
2. The method for controlling traffic lights by multi-agent collaboration based on trajectories according to claim 1 is characterized in that: The time series of the historical trajectory of the agent in step S1. is expressed as: In the formula, is the state of agent i at time step t, T obs is the set observation window size.
3. The method for controlling traffic lights by multi-agent collaboration based on trajectories according to claim 1 is characterized in that: In step S2., the time series is converted into state embedding through a multi-layer perceptron. The conversion process is expressed as follows: In the formula, Represents state embedding.
4. The method for controlling traffic lights by multi-agent collaboration based on trajectories according to claim 1 is characterized in that: The relative position relationship of the state embedding in the sequence described in step S3. is: Here, i is the dimension.
5. The method for controlling traffic lights by multi-agent collaboration based on trajectories according to claim 1 is characterized in that: In step S4., Form the input sequence of the self-attention mechanism.
6. The method for controlling traffic lights by multi-agent collaboration based on trajectories according to claim 5 is characterized in that: The input sequence is mapped to obtain a query vector Q, a key vector K and a value vector V.
7. The method for controlling traffic lights by multi-agent collaboration based on trajectories according to claim 6 is characterized in that: The weights of the nodes are obtained through the self-attention mechanism, and the relationship between the processing is expressed as follows: Where Q is the query vector, K is the key vector, V is the value vector, and d k is the dimension of the vector.
8. The method for controlling traffic lights by multi-agent collaboration based on trajectories according to claim 1 is characterized in that: The function used for normalization described in step S4. is the softmax function.
9. The method for controlling traffic lights by multi-agent collaboration based on trajectories according to claim 1, characterized in that: In step S5., the state embedding of the new node is updated according to the following relationship: sim i,j =a(Wh i,t ,Wh j,t ); Where W is the weight matrix, a is a single-layer feedforward neural network, N i is the neighbor node of intersection i, and σ is the activation function.
10. The method for controlling traffic lights by multi-agent collaboration based on trajectories according to claim 9 is characterized in that: In step S5., the trajectory similarity is calculated as follows: Among them, α∈(0,1) is a weight coefficient used to balance the relative importance between the reward difference and the difference in state transition probability distribution. represents the state transition probability of selecting action a in state s, T K (d) is the Kantorovich distance.
Citation Information
Cited By
Traffic signal lamp control method and device, electronic equipment and computer readable medium
CN121260008A