Traffic signal control method and system based on large language model agent
Through an agent based on a large language model, the traffic data of multiple intersections and the road network spatial relationship diagram are used for coordinated decision-making, which solves the problem that traffic signal control methods in the prior art are difficult to optimize globally in complex road networks, and improves the flexibility of traffic signal control and overall traffic efficiency.
Patent Information
- Application Number
- CN202510566269.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-08-15
AI Technical Summary
The existing traffic signal control methods are difficult to achieve global optimization in complex and changeable traffic environments, resulting in a decline in overall traffic efficiency, especially in a closely connected road network, single-intersection decision-making model is difficult to cope with complex and changeable traffic conditions.
Agents based on large language models are used to control traffic signals. By obtaining traffic data and road network spatial relationship diagrams at multiple intersections, text prompts are constructed and input to the agent, so that they generate traffic signal configurations at the next moment, taking into account the real-time and historical traffic conditions of multiple intersections, and making collaborative decisions.
It improves the flexibility of traffic signal control and the overall traffic efficiency of the road network, can timely adjust traffic congestion status, solves the limitations of traditional single-junction decisions, and realizes global traffic signal optimization.
Smart Images

Figure CN120496339A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the fields of artificial intelligence and traffic signal control technology, and in particular to a traffic signal control method and system based on a large language model intelligent agent. Background Art
[0002] Traffic signal control (TSC) is a system that manages road traffic flow through traffic lights and related equipment, ensuring the safe, efficient, and orderly passage of vehicles and pedestrians. Its core components include traffic lights, traffic signal controllers, traffic detectors, and communication networks. Traffic lights typically use red, yellow, and green, representing stop, ready to go, and go, respectively. The traffic signal controller, the core device of the system, adjusts the lights based on traffic flow or preset times. Traffic detectors monitor the presence and movement of vehicles and pedestrians, providing real-time data to the controller. The communication network connects the controllers, enabling information transmission and coordination to ensure the efficient operation of the entire system. Traffic signal control systems dynamically adjust signal timing based on road traffic conditions to optimize traffic efficiency, reduce congestion, and ensure safety. Traditional systems mostly use timed control, while modern technologies have introduced adaptive control, which uses real-time traffic flow monitoring to dynamically adjust signal strategies to better meet actual needs. Furthermore, with the development of artificial intelligence and big data technologies, intelligent traffic signal control systems can predict traffic flow changes and utilize machine learning algorithms to optimize signal timing, further improving management efficiency.
[0003] Existing traffic signal control methods primarily include those based on traffic engineering, reinforcement learning, and large language models. The algorithm design of traffic signal control methods based on traffic engineering relies heavily on specialized knowledge, requiring the involvement of experts in the field in their development and tuning. Furthermore, because the algorithms are primarily based on preset rules and heuristics, their flexibility and adaptability are limited, making them incapable of coping with complex and changing traffic environments. Traffic signal control methods based on reinforcement learning, however, often have limited generalization capabilities due to their training data typically covering only a limited number of traffic scenarios. They perform poorly in larger traffic networks or under extreme traffic conditions (such as extremely high traffic volumes). Traffic signal control methods based on large language models are often limited to a single intersection and fail to fully consider the interactions and synergies between intersections. In complex road networks with dense intersection connections and high traffic density, this single-intersection decision-making model struggles to achieve global optimization, potentially leading to a decrease in overall traffic efficiency. Summary of the Invention
[0004] In response to the above technical problems, the present application provides a traffic signal control method and system based on a large language model intelligent agent, which fully considers the overall traffic conditions of each intersection, improves the flexibility of traffic signal control and the overall traffic efficiency of the road network.
[0005] In a first aspect, an embodiment of the present application provides a traffic signal control method based on a large language model agent, comprising:
[0006] Acquire traffic data of several preset intersections at the current moment, the traffic data including real-time observation data, historical observation data of a preset time period, and traffic light configuration data corresponding to the historical observation data;
[0007] Constructing a plurality of text prompts based on a preset road network spatial relationship diagram and each of the traffic data, wherein the text prompts include the road network spatial relationship diagram, traffic data of any target intersection, traffic data of adjacent intersections of the target intersection, and a preset task instruction for the target intersection, wherein each of the target intersections corresponds to each of the intelligent agents on a one-to-one basis;
[0008] Inputting each of the text prompts to the corresponding intelligent agents, so that each of the intelligent agents generates the traffic signal configuration at the next moment according to the text prompt;
[0009] Controlling the traffic signal of each corresponding target intersection at the next moment according to each of the traffic signal configurations;
[0010] The intelligent agent is constructed based on a large language model and trained through supervised fine-tuning.
[0011] The embodiment of the present application provides a traffic signal control method based on a large language model agent, which constructs a number of text prompts based on the traffic data of each intersection and the road network spatial relationship diagram, and inputs them into each preset agent, so that it makes a decision to generate the traffic signal configuration of each target intersection, and finally controls the traffic signal of each target intersection at the next moment according to each traffic signal configuration. The embodiment of the present application realizes global traffic signal optimization by setting multiple agents, combining the large language model agent with multi-intersection collaborative control. In the process of traffic signal control, each agent considers the traffic conditions of the target intersection and the adjacent intersections at the same time, and integrates real-time and historical traffic data to generate the signal traffic signal configuration, so that each agent can fully consider the real-time traffic conditions and historical traffic flow change trends of multiple intersections, and perform real-time data collection and real-time decision-making at each time step, so as to adjust the traffic congestion status of the intersection in a timely manner, solve the limitations of traditional single-intersection decision-making, and improve the flexibility of traffic signal control and the overall traffic efficiency of the road network.
[0012] In one possible implementation, when any agent generates a traffic signal configuration for the next moment according to the text prompt, each of the agents generates a traffic signal configuration for the next moment according to the text prompt, including:
[0013] Determining the traffic complexity of the target intersection according to the text prompt;
[0014] determining a current reasoning strategy from among a plurality of reasoning strategies according to the traffic complexity;
[0015] The traffic signal configuration at the next moment is generated according to the current reasoning strategy and the text prompt.
[0016] This embodiment of the application provides a method for generating traffic signal configurations based on text prompts. By assessing the traffic complexity of a target intersection and then determining the most appropriate reasoning strategy from a number of reasoning strategies, the agent can flexibly adjust its reasoning logic based on the actual congestion level at the intersection. For example, a simple strategy can be used for rapid response in mild congestion, while multi-step reasoning optimization can be enabled in more complex congestion, improving the flexibility of traffic signal control and the agent's decision-making efficiency.
[0017] Furthermore, determining the traffic complexity of the target intersection according to the text prompt includes:
[0018] Determining whether each adjacent lane is in a congested state based on the queued vehicle data in the text prompt, wherein the adjacent lane is a lane connected to the target intersection in an adjacent intersection of the target intersection;
[0019] The traffic complexity of the target intersection is determined according to the number of adjacent lanes in a congested state.
[0020] In this embodiment, the congestion status of adjacent lanes is determined by queuing vehicle data. Through quantitative analysis of the congestion status of adjacent lanes, the traffic complexity of the target intersection is accurately assessed. This data-based judgment method avoids subjective bias, provides a reliable basis for subsequent reasoning strategy selection, and enhances the scientific nature of the agent's decision-making.
[0021] Furthermore, if the current reasoning strategy is the first reasoning strategy, then generating the traffic signal configuration at the next moment according to the current reasoning strategy and the text prompt includes:
[0022] With the goal of releasing the most queued vehicles at the target intersection, determining the optimal traffic signal configuration from a plurality of preset candidate traffic signal configurations;
[0023] The optimal traffic signal configuration is used as the traffic signal configuration at the next moment.
[0024] In this embodiment of the present application, according to the first reasoning strategy, the agent prioritizes resolving the current vehicle backlog at the target intersection, directly reducing local congestion pressure. This strategy is applicable to scenarios where adjacent lanes are not congested. In this scenario, since the adjacent lanes at the target intersection are not congested, the agent can focus solely on the traffic conditions at the target intersection, aiming to free up the most queued vehicles and rapidly improve traffic efficiency at the target intersection.
[0025] Furthermore, if the current reasoning strategy is the second reasoning strategy, then generating the traffic signal configuration at the next moment according to the current reasoning strategy and the text prompt includes:
[0026] Determining the current traffic status of the target intersection and adjacent intersections based on the text prompt analysis;
[0027] Determining an optimal traffic signal configuration from among a plurality of preset candidate traffic signal configurations based on the current traffic state with the goal of relieving congestion in adjacent lanes;
[0028] The optimal traffic signal configuration is used as the traffic signal configuration at the next moment.
[0029] In this embodiment of the present application, according to the second inference strategy, the agent will perform collaborative optimization based on the traffic conditions of adjacent intersections, prioritizing the alleviation of congestion in adjacent lanes. This strategy is applicable to scenarios where there are a small number of congested adjacent lanes. In this scenario, the agent first determines the current traffic conditions of the target intersection and adjacent intersections based on text prompts, and then prioritizes the optimization of congested adjacent lanes based on the current traffic conditions. This local collaboration avoids the problem of "resolving congestion at one intersection but worsening congestion at adjacent intersections," thereby improving the balance of the regional road network.
[0030] Furthermore, if the current reasoning strategy is the third reasoning strategy, then generating the traffic signal configuration at the next moment according to the current reasoning strategy and the text prompt includes:
[0031] Determining the current traffic status of the target intersection and adjacent intersections based on the text prompt analysis;
[0032] Predicting and generating the corresponding traffic states at the next moment according to the current traffic state and each preset candidate traffic signal configuration;
[0033] With the goal of maximizing the overall traffic efficiency of the target intersection and adjacent intersections, the traffic conditions at each next moment are analyzed and evaluated, and the candidate traffic signal configuration corresponding to the optimal traffic condition at the next moment is used as the traffic signal configuration at the next moment.
[0034] In this embodiment of the present application, using the third reasoning strategy, the agent not only analyzes the current traffic state but also predicts the future traffic state of the target intersection and adjacent intersections based on different candidate traffic signal configurations. Finally, based on these future traffic states, the agent makes a decision to maximize overall network efficiency, determining the optimal candidate traffic signal configuration. This third reasoning strategy enables the agent to perform a comprehensive reasoning decision, globally optimizing the target intersection and adjacent intersections, thereby improving the overall traffic efficiency of the network.
[0035] In one possible implementation, the intelligent agent is constructed based on a large language model and trained by supervised fine-tuning, including:
[0036] Building an initial agent based on the large language model and a preset reasoning chain, where the reasoning chain is a multi-step reasoning process in which the large language model generates final output data based on input data;
[0037] Establishing a simulation environment based on the road network spatial relationship diagram;
[0038] In the simulation environment, a number of traffic scenarios are generated using a traffic simulator and a preset traffic flow dataset;
[0039] In the simulation environment, according to each traffic scenario and a preset optimization goal, a traffic simulator is used to determine a simulated traffic signal configuration corresponding to each traffic scenario;
[0040] According to each of the traffic scenarios, generating traffic complexity and traffic status corresponding to each traffic scenario through the large language model;
[0041] Constructing a training dataset containing multi-step reasoning based on each of the traffic scenarios, traffic complexity, traffic status, and simulated traffic signal configurations;
[0042] The initial intelligent agent is supervised and fine-tuned according to the training data set and a preset loss function to obtain the intelligent agent.
[0043] An embodiment of the present application provides a method for training an intelligent agent. First, a simulation environment and an initial intelligent agent are constructed. Then, a training dataset containing multi-step reasoning is generated using the simulation environment and the initial intelligent agent. The initial intelligent agent is then supervised and fine-tuned using the training dataset and a preset loss function. Because the training dataset contains both artificially generated traffic scenes and simulated traffic signal configurations, as well as traffic complexity and traffic states generated by the initial intelligent agent, the intelligent agent is able to learn the dynamic interaction patterns in complex road networks based on the reasoning chain, while avoiding overfitting of the intelligent agent during training. This allows for supervised fine-tuning of the intelligent agent and improves its generalization and robustness in real-world scenarios.
[0044] Furthermore, after obtaining the intelligent agent, the intelligent agent is used to generate a plurality of traffic signal configurations, and the intelligent agent is optimized and adjusted according to the plurality of traffic signal configurations based on a preset environmental feedback function.
[0045] The embodiment of the present application introduces an environmental feedback mechanism to continuously optimize the intelligent agent, ensuring that the model can adapt to long-term changes in traffic flow (such as seasonal fluctuations or road reconstruction), improving the accuracy and rationality of the intelligent agent's decision-making, and thereby enhancing the sustainability and long-term effectiveness of traffic signal control based on the intelligent agent.
[0046] In a second aspect, an embodiment of the present application provides a traffic signal control system based on a large language model agent, comprising an acquisition module, a text prompt construction module, a signal generation module, and a control module;
[0047] The acquisition module is used to acquire traffic data of a number of preset intersections at the current moment, wherein the traffic data includes real-time observation data, historical observation data of a preset time period, and traffic light configuration data corresponding to the historical observation data;
[0048] The text prompt construction module is used to construct a plurality of text prompts based on a preset road network spatial relationship diagram and each of the traffic data, wherein the text prompts include the road network spatial relationship diagram, traffic data of any target intersection, traffic data of adjacent intersections of the target intersection, and a preset task instruction for the target intersection, and each of the target intersections corresponds to each of the intelligent agents on a one-to-one basis;
[0049] The signal generation module is used to input each of the text prompts to the corresponding intelligent agents, so that each of the intelligent agents generates the traffic signal configuration at the next moment according to the text prompts;
[0050] The control module is used to control the traffic signal of each corresponding target intersection at the next moment according to each traffic signal configuration;
[0051] The intelligent agent is constructed based on a large language model and trained through supervised fine-tuning.
[0052] In one possible implementation, when any agent generates a traffic signal configuration for the next moment according to the text prompt, each of the agents generates a traffic signal configuration for the next moment according to the text prompt, including:
[0053] Determining the traffic complexity of the target intersection according to the text prompt;
[0054] determining a current reasoning strategy from among a plurality of reasoning strategies according to the traffic complexity;
[0055] The traffic signal configuration at the next moment is generated according to the current reasoning strategy and the text prompt. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 The figure is a flow chart of a traffic signal control method based on a large language model agent;
[0057] Figure 2 A schematic diagram of traffic complexity corresponding to different traffic scenarios in a traffic signal control method based on a large language model agent;
[0058] Figure 3 A flow chart showing the steps from acquiring data to generating traffic signal configurations in a traffic signal control method based on a large language model agent;
[0059] Figure 4 A schematic diagram of the training process of an agent in a traffic signal control method based on a large language model agent;
[0060] Figure 5 This is a structural diagram of a traffic signal control system based on a large language model agent. DETAILED DESCRIPTION
[0061] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0062] It should be noted that the step numbers herein are for convenience of explanation of the specific embodiments and do not serve to define the order in which the steps are to be performed. In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature designated "first" or "second" may explicitly or implicitly include one or more of such features.
[0063] Throughout the specification, some professional terms that appear in this specification are defined as follows:
[0064] Large language models: Large language models represent a significant breakthrough in the field of artificial intelligence in recent years. Their core approach is to leverage deep learning techniques, trained on massive amounts of text data, to understand and generate natural language. Their development has evolved from statistical language models to neural network models, and finally to pre-trained models based on the Transformer architecture. The Transformer architecture's self-attention mechanism enables the model to efficiently capture long-range dependencies, significantly improving language processing performance. As model size increases, large language models have demonstrated exceptional generalization capabilities in tasks such as text generation, translation, and question-answering. For example, models such as the GPT series and BERT have achieved significant results in various fields. Large language models have a wide range of applications, encompassing intelligent customer service, content creation, information retrieval, and other fields. Their technological breakthroughs have driven the development of artificial intelligence and created significant economic value for businesses and society. In the future, as technology improves, large language models will demonstrate even greater capabilities in more complex tasks, becoming a vital force in driving the development of an intelligent society.
[0065] Road network spatial relationship diagram: The road network spatial relationship diagram is a directed graph consisting of intersection nodes V and lanes L. Lanes are divided into straight lanes L according to the direction in which they can be driven. go , left turn lane L right , and right turn lane L right These lanes connect various intersections to form the road network spatial relationship diagram.
[0066] Traffic signal control: At each signal switching time step, the agent controlling the intersection switches from a predefined signal set A = {a1,…,a m}Select a signal. Traffic signal is represented by a=set(L allow ), where L allow is a set of lanes that allow traffic and there is no conflicting movement between these lanes (i.e. L allow The traffic light is green for the first lane and red for other lanes).
[0067] Traffic signal control based on Large Language Model (LLM) agents: Consider a road network consisting of multiple intersections, each of which is controlled by a LLM agent with a control policy π. At each signal switching time step t, each agent receives the following information: 1) Traffic observation data O from the assigned intersection and adjacent intersections t ; 2) Spatial relationship G between intersections; 3) Historical traffic interaction data T t ; 4) Relative task description D. Based on these inputs, the agent infers the optimal traffic signal action a from the action space A t , whose goal is to maximize the traffic efficiency of the entire road network:
[0068] a t =π([Ot ,G,T t ],D,A),
[0069] Example 1:
[0070] like Figure 1 As shown, embodiment 1 provides a traffic signal control method based on a large language model agent, including steps S1 to S4:
[0071] Step S1: acquiring traffic data of a plurality of preset intersections at the current moment, wherein the traffic data includes real-time observation data, historical observation data of a preset time period, and traffic light configuration data corresponding to the historical observation data;
[0072] Step S2: constructing a plurality of text prompts based on a preset road network spatial relationship diagram and each of the traffic data, wherein the text prompts include the road network spatial relationship diagram, traffic data of any target intersection, traffic data of adjacent intersections of the target intersection, and a preset task instruction for the target intersection, wherein each target intersection corresponds to each agent one-to-one;
[0073] Step S3: inputting each of the text prompts to the corresponding agents, so that each of the agents generates a traffic signal configuration at the next moment according to the text prompts;
[0074] Step S4: controlling the traffic signals of the corresponding target intersections at the next moment according to the traffic signal configurations;
[0075] The intelligent agent is constructed based on a large language model and trained through supervised fine-tuning.
[0076] The embodiment of the present application provides a traffic signal control method based on a large language model agent, which constructs a number of text prompts based on the traffic data of each intersection and the road network spatial relationship diagram, and inputs them into each preset agent, so that it makes a decision to generate the traffic signal configuration of each target intersection, and finally controls the traffic signal of each target intersection at the next moment according to each traffic signal configuration. The embodiment of the present application realizes global traffic signal optimization by setting multiple agents, combining the large language model agent with multi-intersection collaborative control. In the process of traffic signal control, each agent considers the traffic conditions of the target intersection and the adjacent intersections at the same time, and integrates real-time and historical traffic data to generate the signal traffic signal configuration, so that each agent can fully consider the real-time traffic conditions and historical traffic flow change trends of multiple intersections, and perform real-time data collection and real-time decision-making at each time step, and can adjust the traffic congestion status of the intersection in a timely manner, solving the limitations of traditional single-intersection decision-making, and improving the flexibility of traffic signal control and the overall traffic efficiency of the road network.
[0077] In a preferred embodiment, in step S1, the real-time observation data includes real-time observation data of each lane at each preset intersection:
[0078] o t ={n queue ,n move ,τ,ρ}
[0079] Among them, n queue is the number of vehicles in the queue, n move is the number of vehicles traveling, τ is the average waiting time, and ρ∈[0,1] represents the lane traffic occupancy rate. These lane-level observation data are aggregated to form the real-time observation data of the target intersection and its adjacent intersections corresponding to the intelligent agent, that is, Where L is the lane associated with the intersection where the agent is located and the adjacent intersection. In order to capture the temporal pattern of traffic congestion propagation, this embodiment collects historical observation data and the traffic light configuration data corresponding to the historical observation data within a fixed time window Δt:
[0080]
[0081] in Indicates that at time t i In addition, to capture the spatial relationship between intersections, a road network spatial relationship graph G = (V, L) is constructed in advance, where V represents the intersection set and L represents the lane set.
[0082] Therefore, for each target intersection, this embodiment can collect the current traffic observation data. t , spatial relationship G, historical traffic data T t The LLM agent responsible for each target intersection will make independent decisions based on this information to maximize the traffic efficiency of itself and surrounding intersections.
[0083] In a preferred embodiment, in step S2, since each agent is constructed based on a large language model, it is necessary to convert the collected traffic data of each intersection into text prompts corresponding to each agent so that the agent can make inference decisions. Each agent needs to make decisions based on the traffic conditions of the target intersection and adjacent intersections at the same time, so the corresponding text prompts include the road network spatial relationship graph G, the traffic data of the target intersection and adjacent intersections. t and T tAnd a pre-set task instruction D for the target intersection. Task instruction D focuses on traffic signal control and includes the following information: a background description of the intersection control environment, the meaning of different actions (signal phases), and basic coordination knowledge with neighboring intersections, such as "avoid sending vehicles to congested sections downstream" or "provide clearance upstream." The text prompt X is represented as follows:
[0084] X=Prompt( t ,G,T t ,D)
[0085] Among them, Prompt(·) represents the process of constructing a text prompt.
[0086] In one possible implementation, in step S3, when any agent generates the traffic signal configuration at the next moment according to the text prompt, each of the agents generates the traffic signal configuration at the next moment according to the text prompt, including:
[0087] Determining the traffic complexity of the target intersection according to the text prompt;
[0088] determining a current reasoning strategy from among a plurality of reasoning strategies according to the traffic complexity;
[0089] The traffic signal configuration at the next moment is generated according to the current reasoning strategy and the text prompt.
[0090] This embodiment of the application provides a method for generating traffic signal configurations based on text prompts. By assessing the traffic complexity of a target intersection and then determining the most appropriate reasoning strategy from a number of reasoning strategies, the agent can flexibly adjust its reasoning logic based on the actual congestion level at the intersection. For example, a simple strategy can be used for rapid response in mild congestion, while multi-step reasoning optimization can be enabled in more complex congestion, improving the flexibility of traffic signal control and the agent's decision-making efficiency.
[0091] Furthermore, determining the traffic complexity of the target intersection according to the text prompt includes:
[0092] Determining whether each adjacent lane is in a congested state based on the queued vehicle data in the text prompt, wherein the adjacent lane is a lane connected to the target intersection in an adjacent intersection of the target intersection;
[0093] The traffic complexity of the target intersection is determined according to the number of adjacent lanes in a congested state.
[0094] In this embodiment, the congestion status of adjacent lanes is determined by queuing vehicle data. Through quantitative analysis of the congestion status of adjacent lanes, the traffic complexity of the target intersection is accurately assessed. This data-based judgment method avoids subjective bias, provides a reliable basis for subsequent reasoning strategy selection, and enhances the scientific nature of the agent's decision-making.
[0095] In a preferred embodiment, the agent can follow a three-step decision-making process to perform traffic signal control:
[0096] 1. Analyze the traffic conditions of the current intersection and its adjacent intersections to obtain Y ana ;
[0097] 2. Predict the traffic status at the next moment under different signal configurations a∈A
[0098] 3. Choose the signal configuration that is most conducive to the overall traffic flow as the final decision
[0099] The decision-making process is formulated as follows:
[0100]
[0101] Y ana Indicates the detailed analysis results of the current traffic conditions; It represents the expected traffic state of the area at the next moment after the target intersection activates signal a under the prediction of LLM; represents the optimal signal configuration based on the analysis.
[0102] However, although the above reasoning strategy can effectively handle the complex traffic dependencies between intersections and improve decision accuracy, its cumbersome reasoning steps (especially the prediction of future traffic status) will result in a high time cost. Therefore, the embodiment of the present application further introduces the concept of traffic complexity, so that the intelligent agent first evaluates the traffic complexity of the target intersection based on the text prompt, and adaptively selects the most cost-effective reasoning strategy without affecting performance. Specifically, Figure 2 As shown, this embodiment is based on the traffic complexity n c Grading the decision complexity of traffic scenarios, n c Refers to the number of congested adjacent lanes, and adjacent lanes refer to lanes connected to the own intersection controlled by adjacent intersections. c The larger the value, the less error tolerance the agent has in its decision-making process and the more factors it needs to consider carefully. In the actual reasoning process of the agent, the agent first identifies n c Then select the corresponding reasoning strategy, and finally generate the traffic signal configuration at the next moment based on the reasoning strategy and text prompts.
[0103] Furthermore, if the current reasoning strategy is the first reasoning strategy, then generating the traffic signal configuration at the next moment according to the current reasoning strategy and the text prompt includes:
[0104] With the goal of releasing the most queued vehicles at the target intersection, determining the optimal traffic signal configuration from a plurality of preset candidate traffic signal configurations;
[0105] The optimal traffic signal configuration is used as the traffic signal configuration at the next moment.
[0106] In this embodiment of the present application, according to the first reasoning strategy, the agent prioritizes resolving the current vehicle backlog at the target intersection, directly reducing local congestion pressure. This strategy is applicable to scenarios where adjacent lanes are not congested. In this scenario, since the adjacent lanes at the target intersection are not congested, the agent can focus solely on the traffic conditions at the target intersection, aiming to free up the most queued vehicles and rapidly improve traffic efficiency at the target intersection.
[0107] Furthermore, if the current reasoning strategy is the second reasoning strategy, then generating the traffic signal configuration at the next moment according to the current reasoning strategy and the text prompt includes:
[0108] Determining the current traffic status of the target intersection and adjacent intersections based on the text prompt analysis;
[0109] Determining an optimal traffic signal configuration from among a plurality of preset candidate traffic signal configurations based on the current traffic state with the goal of relieving congestion in adjacent lanes;
[0110] The optimal traffic signal configuration is used as the traffic signal configuration at the next moment.
[0111] In this embodiment of the present application, according to the second inference strategy, the agent will perform collaborative optimization based on the traffic conditions of adjacent intersections, prioritizing the alleviation of congestion in adjacent lanes. This strategy is applicable to scenarios where there are a small number of congested adjacent lanes. In this scenario, the agent first determines the current traffic conditions of the target intersection and adjacent intersections based on text prompts, and then prioritizes the optimization of congested adjacent lanes based on the current traffic conditions. This local collaboration avoids the problem of "resolving congestion at one intersection but worsening congestion at adjacent intersections," thereby improving the balance of the regional road network.
[0112] Furthermore, if the current reasoning strategy is the third reasoning strategy, then generating the traffic signal configuration at the next moment according to the current reasoning strategy and the text prompt includes:
[0113] Determining the current traffic status of the target intersection and adjacent intersections based on the text prompt analysis;
[0114] Predicting and generating the corresponding traffic states at the next moment according to the current traffic state and each preset candidate traffic signal configuration;
[0115] With the goal of maximizing the overall traffic efficiency of the target intersection and adjacent intersections, the traffic conditions at each next moment are analyzed and evaluated, and the candidate traffic signal configuration corresponding to the optimal traffic condition at the next moment is used as the traffic signal configuration at the next moment.
[0116] In this embodiment of the present application, using the third reasoning strategy, the agent not only analyzes the current traffic state but also predicts the future traffic state of the target intersection and adjacent intersections based on different candidate traffic signal configurations. Finally, based on these future traffic states, the agent makes a decision to maximize overall network efficiency, determining the optimal candidate traffic signal configuration. This third reasoning strategy enables the agent to perform a comprehensive reasoning decision, globally optimizing the target intersection and adjacent intersections, thereby improving the overall traffic efficiency of the network.
[0117] In a preferred embodiment, Figure 2 As shown, according to n c The inference strategy can be set to three types, including:
[0118] 1) If Figure 2 As shown in (a), when n c = 0 (no collaboration required), corresponding to the first reasoning strategy: If all adjacent lanes are unobstructed, it means that the traffic pressure at the current intersection is relatively small, and the decision does not need to consider the influence of neighbors. The agent can directly choose the signal configuration that releases the most queued vehicles. In this case, the complex analysis and prediction process is skipped, and the formula is as follows:
[0119]
[0120] 2) If Figure 2 As shown in (b), when n c = 1 (simple collaboration), corresponding to the second reasoning strategy: if there is only one adjacent lane with high traffic occupancy, the agent performs simple reasoning about the current situation, considering the current state and collaboration relationship between the intersection and the adjacent lane, but does not need to perform future traffic state prediction:
[0121]
[0122] 3) If Figure 2 As shown in (c), when n cWhen >1 (complex collaboration), the third reasoning strategy corresponds: when multiple adjacent lanes are congested, the agent needs to comprehensively consider the spatiotemporal dependencies between multiple lanes, including the multi-hop propagation effects between multiple intersections. In this case, the agent needs to perform a complete spatiotemporal awareness collaborative decision-making process to ensure the optimization of global traffic efficiency:
[0123]
[0124] At the beginning of the reasoning process, each intersection agent will automatically analyze n based on the current observation information. c , and selects the appropriate reasoning path accordingly. As fine-tuning progresses, the agent gradually learns to more accurately distinguish which adjacent lanes are in a "congested" or "high occupancy" state, thereby dynamically adjusting its reasoning strategy to achieve efficient collaborative control. The steps from the agent's acquisition of data to the generation of traffic signal configurations are as follows: Figure 3 shown.
[0125] In one possible implementation, the intelligent agent is constructed based on a large language model and trained by supervised fine-tuning, including:
[0126] Building an initial agent based on the large language model and a preset reasoning chain, where the reasoning chain is a multi-step reasoning process in which the large language model generates final output data based on input data;
[0127] Establishing a simulation environment based on the road network spatial relationship diagram;
[0128] In the simulation environment, a number of traffic scenarios are generated using a traffic simulator and a preset traffic flow dataset;
[0129] In the simulation environment, according to each traffic scenario and a preset optimization goal, a traffic simulator is used to determine a simulated traffic signal configuration corresponding to each traffic scenario;
[0130] According to each of the traffic scenarios, generating traffic complexity and traffic status corresponding to each traffic scenario through the large language model;
[0131] Constructing a training dataset containing multi-step reasoning based on each of the traffic scenarios, traffic complexity, traffic status, and simulated traffic signal configurations;
[0132] The initial intelligent agent is supervised and fine-tuned according to the training data set and a preset loss function to obtain the intelligent agent.
[0133] An embodiment of the present application provides a method for training an intelligent agent. First, a simulation environment and an initial intelligent agent are constructed. Then, a training dataset containing multi-step reasoning is generated using the simulation environment and the initial intelligent agent. The initial intelligent agent is then supervised and fine-tuned using the training dataset and a preset loss function. Because the training dataset contains both artificially generated traffic scenes and simulated traffic signal configurations, as well as traffic complexity and traffic states generated by the initial intelligent agent, the intelligent agent is able to learn the dynamic interaction patterns in complex road networks based on the reasoning chain, while avoiding overfitting of the intelligent agent during training. This allows for supervised fine-tuning of the intelligent agent and improves its generalization and robustness in real-world scenarios.
[0134] In a preferred embodiment, the process of constructing the agent based on a large language model and training it through supervised fine-tuning is as follows:
[0135] First, a traffic simulator and a synthetic traffic flow dataset are used to generate diverse traffic scenarios covering different congestion levels. t Then, GPT-4o (the initial agent) is used to analyze the traffic complexity and traffic conditions of the designated intersection and its neighboring intersections:
[0136] n c ,Y ana =f GPT-4o (O t ),
[0137] Among them, n c represents traffic complexity, Y ana Represents the traffic analysis results.
[0138] In order to enable the model to learn the traffic impact relationship between intersections, the simulator is run to obtain the next traffic state after activating different signal configurations in the current state:
[0139]
[0140] where f sim Represents the state transition function of the simulator.
[0141] Typically, during reinforcement learning, a large amount of policy sampling is required to interact with the environment, thereby calculating the long-term cumulative reward for training and bringing the model's policy close to the global optimum. However, the high inference overhead of LLM agents makes large-scale, long-term sampling difficult. Therefore, to help LLM agents learn faster, this example uses a five-step simulation to determine a pseudo-golden signal.
[0142] Definition of pseudo-golden signal: A pseudo-golden signal refers to a signal configuration that is selected during the simulation process by evaluating the impact of different signal configurations on traffic flow and is able to minimize the total number of vehicles queuing at an intersection and its adjacent intersections within a specific time.
[0143] Specifically, in order to determine the pseudo golden signal a * We simulate multiple signal operations and evaluate their impact on traffic flow, and select the signal that minimizes the total number of vehicles queued at its own and neighboring intersections within five time steps as a * :
[0144]
[0145] where f queue represents the queue length measurement function of the simulator. This pseudo golden signal is used as the optimized supervision signal for training, which is consistent with the n generated by GPT-4o. c ,Y ana , and the future traffic status obtained through simulation Together, they are synthesized into a training dataset containing multi-step reasoning for fine-tuning the initial agent.
[0146] To optimize the LLM, this example uses a supervised fine-tuning method to minimize the negative log-likelihood of generating the target inference chain:
[0147]
[0148] Where Y is the synthetic reasoning chain, P π (y w |X,Y <w ) represents the preceding token Y in a given prompt X and inference chain <w Generate a marker y w probability.
[0149] Furthermore, after obtaining the intelligent agent, the intelligent agent is used to generate a plurality of traffic signal configurations, and the intelligent agent is optimized and adjusted according to the plurality of traffic signal configurations based on a preset environmental feedback function.
[0150] The embodiment of the present application introduces an environmental feedback mechanism to continuously optimize the intelligent agent, ensuring that the model can adapt to long-term changes in traffic flow (such as seasonal fluctuations or road reconstruction), improving the accuracy and rationality of the intelligent agent's decision-making, and thereby enhancing the sustainability and long-term effectiveness of traffic signal control based on the intelligent agent.
[0151] In a preferred embodiment, after obtaining the intelligent agent, although the intelligent agent can effectively follow the reasoning process we constructed, its decision-making ability is still insufficient. To solve this problem, this embodiment proposes an iterative optimization method combined with environmental feedback to further improve the model performance. Specifically, we simulate different traffic signal configurations based on the agent's decision-making and evaluate their effectiveness through the environmental feedback function Q. The higher the feedback value, the better the signal configuration under the given traffic conditions. Subsequently, we identify the reasoning chain corresponding to the signal configuration that can generate the highest environmental feedback value, and use it as a "pseudo-golden reasoning chain" for further fine-tuning:
[0152]
[0153] Where T represents the time window for collecting environmental feedback. The environmental feedback function Q is defined as the inverse of the total queue length of the adjacent intersection after five time steps, ensuring that lower congestion levels can obtain higher feedback values, thereby optimizing the decision-making ability of the model. The fine-tuning process of the agent and the optimization process based on environmental feedback are as follows: Figure 4 shown.
[0154] Example 2:
[0155] like Figure 5 As shown, embodiment 2 provides a traffic signal control system based on a large language model agent, including an acquisition module 10, a text prompt construction module 20, a signal generation module 30 and a control module 40;
[0156] The acquisition module 10 is used to acquire traffic data of a plurality of preset intersections at the current moment, wherein the traffic data includes real-time observation data, historical observation data of a preset time period, and traffic light configuration data corresponding to the historical observation data;
[0157] The text prompt construction module 20 is used to construct a plurality of text prompts based on a preset road network spatial relationship diagram and each of the traffic data, wherein the text prompts include the road network spatial relationship diagram, traffic data of any target intersection, traffic data of adjacent intersections of the target intersection, and a preset task instruction for the target intersection, and each of the target intersections corresponds to each of the intelligent agents on a one-to-one basis;
[0158] The signal generation module 30 is used to input each of the text prompts to the corresponding agents, so that each of the agents generates the traffic signal configuration at the next moment according to the text prompts;
[0159] The control module 40 is used to control the traffic signal of each corresponding target intersection at the next moment according to each traffic signal configuration;
[0160] The intelligent agent is constructed based on a large language model and trained through supervised fine-tuning.
[0161] In one possible implementation, when any agent generates a traffic signal configuration for the next moment according to the text prompt, each of the agents generates a traffic signal configuration for the next moment according to the text prompt, including:
[0162] Determining the traffic complexity of the target intersection according to the text prompt;
[0163] determining a current reasoning strategy from among a plurality of reasoning strategies according to the traffic complexity;
[0164] The traffic signal configuration at the next moment is generated according to the current reasoning strategy and the text prompt.
[0165] Furthermore, determining the traffic complexity of the target intersection according to the text prompt includes:
[0166] Determining whether each adjacent lane is in a congested state based on the queued vehicle data in the text prompt, wherein the adjacent lane is a lane connected to the target intersection in an adjacent intersection of the target intersection;
[0167] The traffic complexity of the target intersection is determined according to the number of adjacent lanes in a congested state.
[0168] Furthermore, if the current reasoning strategy is the first reasoning strategy, then generating the traffic signal configuration at the next moment according to the current reasoning strategy and the text prompt includes:
[0169] With the goal of releasing the most queued vehicles at the target intersection, determining the optimal traffic signal configuration from a plurality of preset candidate traffic signal configurations;
[0170] The optimal traffic signal configuration is used as the traffic signal configuration at the next moment.
[0171] Furthermore, if the current reasoning strategy is the second reasoning strategy, then generating the traffic signal configuration at the next moment according to the current reasoning strategy and the text prompt includes:
[0172] Determining the current traffic status of the target intersection and adjacent intersections based on the text prompt analysis;
[0173] Determining an optimal traffic signal configuration from among a plurality of preset candidate traffic signal configurations based on the current traffic state with the goal of relieving congestion in adjacent lanes;
[0174] The optimal traffic signal configuration is used as the traffic signal configuration at the next moment.
[0175] Furthermore, if the current reasoning strategy is the third reasoning strategy, then generating the traffic signal configuration at the next moment according to the current reasoning strategy and the text prompt includes:
[0176] Determining the current traffic status of the target intersection and adjacent intersections based on the text prompt analysis;
[0177] Predicting and generating the corresponding traffic states at the next moment according to the current traffic state and each preset candidate traffic signal configuration;
[0178] With the goal of maximizing the overall traffic efficiency of the target intersection and adjacent intersections, the traffic conditions at each next moment are analyzed and evaluated, and the candidate traffic signal configuration corresponding to the optimal traffic condition at the next moment is used as the traffic signal configuration at the next moment.
[0179] In one possible implementation, the traffic signal control further includes a model training module, which is used to construct the intelligent agent based on the large language model and train it through supervised fine-tuning, including a model construction unit, an environment establishment unit, a first simulation unit, a second simulation unit, a data set construction unit, and a training unit;
[0180] The model building unit is used to build an initial intelligent agent based on the large language model and a preset reasoning chain, where the reasoning chain is a multi-step reasoning process in which the large language model generates final output data based on input data;
[0181] The environment establishment unit is used to establish a simulation environment based on the road network spatial relationship graph;
[0182] The first simulation unit is used to generate a plurality of traffic scenarios in the simulation environment using a traffic simulator and a preset traffic flow data set;
[0183] The second simulation unit is used to determine, in the simulation environment, according to each traffic scenario and a preset optimization target, a simulated traffic signal configuration corresponding to each traffic scenario through a traffic simulator;
[0184] The data set construction unit is used to generate traffic complexity and traffic status corresponding to each traffic scene through the large language model according to each traffic scene;
[0185] The training unit is used to construct a training data set containing multi-step reasoning according to each of the traffic scenarios, traffic complexity, traffic status and simulated traffic signal configuration;
[0186] The initial intelligent agent is supervised and fine-tuned according to the training data set and a preset loss function to obtain the intelligent agent.
[0187] Furthermore, after obtaining the intelligent agent, the intelligent agent is used to generate a plurality of traffic signal configurations, and the intelligent agent is optimized and adjusted according to the plurality of traffic signal configurations based on a preset environmental feedback function.
[0188] The embodiment of the present application provides a traffic signal control system based on a large language model agent. According to the traffic data of each intersection and the road network spatial relationship diagram, a number of text prompts are constructed and input into each preset agent, so that it makes a decision to generate the traffic signal configuration of each target intersection, and finally controls the traffic signal of each target intersection at the next moment according to each traffic signal configuration. The embodiment of the present application realizes global traffic signal optimization by setting multiple agents, combining the large language model agent with multi-intersection collaborative control. In the process of traffic signal control, each agent considers the traffic conditions of the target intersection and the adjacent intersections at the same time, and integrates real-time and historical traffic data to generate the signal traffic signal configuration, so that each agent can fully consider the real-time traffic conditions and historical traffic flow change trends of multiple intersections, and perform real-time data collection and real-time decision-making at each time step, so as to adjust the traffic congestion status of the intersection in a timely manner, solve the limitations of traditional single-intersection decision-making, and improve the flexibility of traffic signal control and the overall traffic efficiency of the road network.
[0189] The more detailed working principle and process flow of this embodiment can be referred to, but not limited to, the relevant records of the first embodiment.
[0190] The specific embodiments described above further illustrate the objectives, technical solutions, and beneficial effects of this application. It should be understood that the above descriptions are merely specific embodiments of this application and are not intended to limit the scope of protection of this application. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of this application by those skilled in the art should be included within the scope of protection of this application.
Claims
1. A traffic signal control method based on a large language model agent, characterized in that: include: Acquire traffic data of several preset intersections at the current moment, the traffic data including real-time observation data, historical observation data of a preset time period, and traffic light configuration data corresponding to the historical observation data; Constructing a plurality of text prompts based on a preset road network spatial relationship diagram and each of the traffic data, wherein the text prompts include the road network spatial relationship diagram, traffic data of any target intersection, traffic data of adjacent intersections of the target intersection, and a preset task instruction for the target intersection, wherein each of the target intersections corresponds to each of the intelligent agents on a one-to-one basis; Inputting each of the text prompts to the corresponding intelligent agents, so that each of the intelligent agents generates the traffic signal configuration at the next moment according to the text prompt; Controlling the traffic signal of each corresponding target intersection at the next moment according to each of the traffic signal configurations; The intelligent agent is constructed based on a large language model and trained through supervised fine-tuning.
2. The traffic signal control method based on a large language model agent according to claim 1, characterized in that: When any agent generates the traffic signal configuration at the next moment according to the text prompt, each of the agents generates the traffic signal configuration at the next moment according to the text prompt, including: Determining the traffic complexity of the target intersection according to the text prompt; determining a current reasoning strategy from among a plurality of reasoning strategies according to the traffic complexity; The traffic signal configuration at the next moment is generated according to the current reasoning strategy and the text prompt.
3. The traffic signal control method based on a large language model agent according to claim 2, characterized in that: Determining the traffic complexity of the target intersection according to the text prompt includes: Determining whether each adjacent lane is in a congested state based on the queued vehicle data in the text prompt, wherein the adjacent lane is a lane connected to the target intersection in an adjacent intersection of the target intersection; The traffic complexity of the target intersection is determined according to the number of adjacent lanes in a congested state.
4. The traffic signal control method based on a large language model agent according to claim 2, characterized in that: If the current reasoning strategy is the first reasoning strategy, then generating the traffic signal configuration at the next moment according to the current reasoning strategy and the text prompt includes: With the goal of releasing the most queued vehicles at the target intersection, determining the optimal traffic signal configuration from a plurality of preset candidate traffic signal configurations; The optimal traffic signal configuration is used as the traffic signal configuration at the next moment.
5. The traffic signal control method based on a large language model agent according to claim 2, characterized in that: If the current reasoning strategy is the second reasoning strategy, then generating the traffic signal configuration at the next moment according to the current reasoning strategy and the text prompt includes: Determining the current traffic status of the target intersection and adjacent intersections based on the text prompt analysis; Determining an optimal traffic signal configuration from among a plurality of preset candidate traffic signal configurations based on the current traffic state with the goal of relieving congestion in adjacent lanes; The optimal traffic signal configuration is used as the traffic signal configuration at the next moment.
6. The traffic signal control method based on a large language model agent according to claim 2, characterized in that: If the current reasoning strategy is the third reasoning strategy, then generating the traffic signal configuration at the next moment according to the current reasoning strategy and the text prompt includes: Determining the current traffic status of the target intersection and adjacent intersections based on the text prompt analysis; Predicting and generating the corresponding traffic states at the next moment according to the current traffic state and each preset candidate traffic signal configuration; With the goal of maximizing the overall traffic efficiency of the target intersection and adjacent intersections, the traffic conditions at each next moment are analyzed and evaluated, and the candidate traffic signal configuration corresponding to the optimal traffic condition at the next moment is used as the traffic signal configuration at the next moment.
7. A traffic signal control method based on a large language model agent according to any one of claims 1 to 6, characterized in that: The intelligent agent is constructed based on a large language model and trained by supervised fine-tuning, including: Building an initial agent based on the large language model and a preset reasoning chain, where the reasoning chain is a multi-step reasoning process in which the large language model generates final output data based on input data; Establishing a simulation environment based on the road network spatial relationship diagram; In the simulation environment, a number of traffic scenarios are generated using a traffic simulator and a preset traffic flow dataset; In the simulation environment, according to each traffic scenario and a preset optimization goal, a traffic simulator is used to determine a simulated traffic signal configuration corresponding to each traffic scenario; According to each of the traffic scenarios, generating traffic complexity and traffic status corresponding to each traffic scenario through the large language model; Constructing a training dataset containing multi-step reasoning based on each of the traffic scenarios, traffic complexity, traffic status, and simulated traffic signal configurations; The initial intelligent agent is supervised and fine-tuned according to the training data set and a preset loss function to obtain the intelligent agent.
8. The traffic signal control method based on a large language model agent according to claim 7, characterized in that: After obtaining the intelligent agent, a plurality of traffic signal configurations are generated using the intelligent agent, and the intelligent agent is optimized and adjusted according to the plurality of traffic signal configurations based on a preset environmental feedback function.
9. A traffic signal control system based on a large language model agent, characterized in that: It includes an acquisition module, a text prompt building module, a signal generation module and a control module; The acquisition module is used to acquire traffic data of a number of preset intersections at the current moment, wherein the traffic data includes real-time observation data, historical observation data of a preset time period, and traffic light configuration data corresponding to the historical observation data; The text prompt construction module is used to construct a plurality of text prompts based on a preset road network spatial relationship diagram and each of the traffic data, wherein the text prompts include the road network spatial relationship diagram, traffic data of any target intersection, traffic data of adjacent intersections of the target intersection, and a preset task instruction for the target intersection, and each of the target intersections corresponds to each of the intelligent agents on a one-to-one basis; The signal generation module is used to input each of the text prompts to the corresponding intelligent agents, so that each of the intelligent agents generates the traffic signal configuration at the next moment according to the text prompts; The control module is used to control the traffic signal of each corresponding target intersection at the next moment according to each traffic signal configuration; The intelligent agent is constructed based on a large language model and trained through supervised fine-tuning.
10. The traffic signal control system based on a large language model agent according to claim 9, characterized in that: When any agent generates the traffic signal configuration at the next moment according to the text prompt, each of the agents generates the traffic signal configuration at the next moment according to the text prompt, including: Determining the traffic complexity of the target intersection according to the text prompt; determining a current reasoning strategy from among a plurality of reasoning strategies according to the traffic complexity; The traffic signal configuration at the next moment is generated according to the current reasoning strategy and the text prompt.
Citation Information
Cited By
Credit control scheme generation method, related device, equipment and storage medium
CN120808617A