Traffic decision collaborative optimization system based on generative AI

The traffic decision-making collaborative optimization system based on generative AI solves the problems of high cost and high risk of traditional methods in extreme situations and emergencies. It realizes efficient and realistic traffic data generation and optimization decision-making, thereby improving urban traffic management.

CN120932450APending Publication Date: 2025-11-11XI AN JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511129119.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Traditional urban traffic management and emergency management are costly and risky when facing extreme situations and emergencies, and are difficult to cover all kinds of extreme scenarios and emergencies.

Method used

A traffic decision-making collaborative optimization system based on generative AI is adopted, including natural language input, prompt management, natural language understanding and task planning, traffic basic model execution, result output and intermediate response, task evaluation and continuation, final response generation and dialogue memory storage. Combined with diffusion communication mechanism and spatiotemporal transformer, spatiotemporal dependencies are modeled through hierarchical attention module to generate highly realistic and efficient virtual traffic data.

Benefits of technology

It enables efficient simulation and optimization of traffic plans in a virtual environment, reduces reliance on real-world data collection, improves data generation efficiency, rapidly evaluates multiple plans, optimizes urban traffic decisions, and solves the problem of local decision-making bias caused by data gaps in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120932450A_ABST
    Figure CN120932450A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent traffic and artificial intelligence, in particular to a traffic decision collaborative optimization system based on generative AI (artificial intelligence), and the traffic decision collaborative optimization system based on the generative AI comprises a traffic decision collaborative optimization module, a traffic decision collaborative optimization module and a traffic decision collaborative optimization module, s1, natural language input; s2, prompting management; s3, natural language understanding and task planning; and S4, executing the traffic basic model. The traffic decision collaborative optimization system based on the generative AI has remarkable advantages in the aspects of simulation authenticity, data generation efficiency and decision optimization, the AIGC can generate highly real virtual data, and the data can fully simulate complex interaction in traffic flow, such as influence of traffic signals on vehicle flow and change of pedestrian behaviors. And secondly, the AIGC can greatly improve the efficiency of data generation and reduce the dependence on real data acquisition, and most importantly, the AIGC can help urban managers to quickly evaluate the effects of various traffic schemes by combining with an optimization algorithm, so that the optimal decision scheme is selected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of intelligent transportation and artificial intelligence technology, specifically a traffic decision-making collaborative optimization system based on generative AI. Background Technology

[0002] Smart cities refer to the efficient management and optimization of urban resources through various information technologies and artificial intelligence. In this process, the combination of AIGC and digital twin technology plays a crucial role. Digital twins refer to the real-time mapping of virtual models with physical entities, thereby simulating and predicting the real world in a virtual environment. The combination of AIGC and digital twins can not only reflect the real-time operation status of urban transportation systems, but also conduct rehearsals by generating virtual scenarios to predict various scenarios such as traffic flow and emergency evacuation, helping city managers to better plan traffic routes, optimize signal timing, and formulate emergency response strategies.

[0003] Traditional urban traffic management and emergency response typically rely on real-world data collection and complex field tests. While these tests are crucial, their high cost and unpredictable risks make it difficult to cover extreme scenarios and various emergencies. For example, testing scenarios such as "traffic congestion caused by a severe rainstorm" or "evacuation of 300,000 people around a nuclear power plant" is not only extremely costly but also carries significant risks. However, by generating extreme traffic scenarios through AIGC, managers can conduct simulation tests in a virtual environment, optimize solutions, and make decision adjustments. To this end, we propose a traffic decision-making collaborative optimization system based on generative AI. Summary of the Invention

[0004] The purpose of this invention is to provide a traffic decision-making collaborative optimization system based on generative AI, to address the problem mentioned in the background art that its high cost and unpredictable risks make it difficult to cover extreme situations and various emergencies. To achieve the above objective, this invention provides the following technical solution: a traffic decision-making collaborative optimization system based on generative AI, wherein the traffic decision-making collaborative optimization system includes;

[0005] S1: Natural language input;

[0006] S2: Notification Management;

[0007] S3: Natural Language Understanding and Task Planning;

[0008] S4: Traffic infrastructure model execution;

[0009] S5: Results Output and Intermediate Responses;

[0010] S6: Mission Assessment and Continuation;

[0011] S7: Final response generation;

[0012] S8: Dialogue memory storage;

[0013] Preferably, the user inputs specific simulation task requirements using natural language through the traffic simulation intelligent agent front-end interface. This input text is then passed as a prompt to the next step for prompt management.

[0014] Preferably, as a foundational step, the prompt management needs to define the agent's working mechanism, clarify key considerations, and convey information about the available toolset. Furthermore, this step can integrate historical dialogue context to support multi-turn interactions. The components of this integrated prompt include: user task requests, system prefixes, available tools, inference history, and dialogue history. By integrating these elements into a coherent prompt, the agent obtains the necessary context and instructions, effectively facilitating task decomposition and execution.

[0015] Preferably, with the help of LLM capabilities, the agent can understand natural language prompts and perform deductive reasoning by integrating task requests, available toolsets and reasoning history databases, ultimately forming recognizable and actionable insights.

[0016] Preferably, based on established insights, the agent invokes a pre-defined traffic infrastructure model and strictly adheres to the prerequisites specified in the tool definition when defining parameters. The traffic infrastructure model then executes specific tasks based on the parameters defined by the agent, including database retrieval and analysis, data visualization, and system optimization, ultimately generating the required output.

[0017] Preferably, after the tool is executed, the agent obtains the output of the traffic infrastructure model through the API interface. The agent integrates the tool output into an intermediate response in natural language form for LLM to further plan. In scenarios where multimodal output is required as supplementary information, structured content (such as tables) will be generated in Markdown format, while components such as visualization images and data files will be provided in the form of file paths.

[0018] Preferably, the intelligent agent compares and analyzes the user's task request with the current intermediate response to assess the task completion status. If the task is not yet resolved, the process will return to steps 2 to 5 to ensure the iterative continuation of the execution process.

[0019] Preferably, after confirming the completion of the task in step 6, the agent utilizes the powerful capabilities of LLM to integrate the output content generated by the tool, form a conclusive response, and transmit it to the user through the front-end interface.

[0020] Preferably, user input and LLM output are stored to preserve ongoing dialogue. These records are aggregated into dialogue history and used as part of the prompting management input in subsequent interactions, providing dialogue context for large language models and enabling them to have memory capabilities.

[0021] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0022] In this invention, AIGC generates virtual traffic data, which has significant advantages in terms of simulation realism, data generation efficiency, and decision optimization. First, AIGC can generate highly realistic virtual data that can fully simulate the complex interactions in traffic flow, such as the impact of traffic signals on vehicle flow and changes in pedestrian behavior. Second, AIGC can greatly improve the efficiency of data generation and reduce the dependence on real-world data collection. Most importantly, AIGC can help city managers quickly evaluate the effects of various traffic schemes by combining with optimization algorithms, thereby selecting the optimal decision scheme.

[0023] In this invention, the generated intersection observation information can be dynamically propagated to adjacent nodes during the inference phase through a diffusion communication mechanism, while the spatiotemporal transformer explicitly models the spatiotemporal dependencies between intersections through a hierarchical attention module. This collaborative design solves the problem of local decision bias caused by data gaps in traditional methods. Diffusion generation plays a dual role in this process: on the one hand, partial reward condition diffusion is guided by classifier independence, relying only on observable rewards to generate high-reward trajectories, avoiding noise interference introduced by traditional filling methods; on the other hand, the diffusion communication mechanism uses the reverse process of the diffusion model to generate virtual observations of adjacent intersections, and integrates this information into local decision-making through the communication cross-attention module of the spatiotemporal transformer. Attached Figure Description

[0024] Figure 1 This is a schematic diagram of the process structure of the present invention;

[0025] Figure 2 This is a schematic diagram of the traffic simulation intelligent agent structure of the present invention. Detailed Implementation

[0026] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0027] Please see Figures 1-2 The present invention provides a technical solution: a traffic decision-making collaborative optimization system based on generative AI, the traffic decision-making collaborative optimization system comprising;

[0028] S1: Natural language input;

[0029] S2: Prompt Management;

[0030] S3: Natural Language Understanding and Task Planning;

[0031] S4: Traffic infrastructure model execution;

[0032] S5: Results Output and Intermediate Responses;

[0033] S6: Mission Assessment and Continuation;

[0034] S7: Final response generation;

[0035] S8: Dialogue memory storage;

[0036] Step one: The user inputs specific simulation task requirements using natural language through the traffic simulation intelligent agent front-end interface. This input text is passed to the next step as a prompt for prompt management.

[0037] Step two, as a foundational step, prompts management to define the agent's working mechanism, clarify key considerations, and convey information about available toolsets. Furthermore, this step integrates historical dialogue context to support multi-turn interactions. This integrated prompt comprises: user task requests, system prefixes, available tools, inference history, and dialogue history. By integrating these elements into a coherent prompt, the agent obtains the necessary context and instructions, effectively facilitating task decomposition and execution.

[0038] Step 3: With the help of LLM, the agent can understand natural language prompts and perform deductive reasoning by integrating task requests, available toolsets and reasoning history databases, ultimately forming identifiable and actionable insights.

[0039] Step four: Based on the established insights, the agent invokes a pre-configured traffic infrastructure model and strictly adheres to the prerequisites specified in the tool definition when setting parameters. The traffic infrastructure model then executes specific tasks based on the parameters set by the agent, including database retrieval and analysis, data visualization, and system optimization, ultimately generating the required output.

[0040] Step 5: After the tool is executed, the agent obtains the output of the traffic basic model through the API interface. The agent integrates the tool output into an intermediate response in natural language form for LLM to further plan. In scenarios that require multimodal output as supplementary information, structured content (such as tables) will be generated in Markdown format, while components such as visualization images and data files will be provided in the form of file paths.

[0041] Step 6: The agent compares and analyzes the user's task request with the current intermediate response to assess the task completion status. If the task is not yet resolved, the process will return to steps 2 to 5 to ensure the iterative continuation of the execution process.

[0042] Step 7: After confirming the completion of the task in Step 6, the agent uses the powerful capabilities of LLM to integrate the output generated by the tool, form a conclusive response, and transmit it to the user through the front-end interface.

[0043] Step eight involves storing user input and LLM output to preserve the ongoing dialogue. These records are aggregated into a dialogue history and used as part of the input management for subsequent interactions, providing dialogue context for the large language model and enabling it to have memory capabilities. The traffic simulation agent is an intelligent system combining a large language model and a traffic infrastructure model. Its core architecture consists of a natural language interaction module, a large language model task module, and a traffic infrastructure model library module. The core task of the system prefix is ​​to clarify the agent's role as an AI assistant in the traffic domain, define its capabilities by listing the core functions it supports, avoid illusions or out-of-bounds responses from the large model, and declare the rules for invoking the traffic infrastructure model. Specific invoking rules include an autonomous selection mechanism and priority rules. The autonomous selection mechanism allows the LLM to automatically match traffic foundation models (TFMs) based on task requirements. Priority rules assign default priorities to functionally similar TFMs (e.g., classic models are prioritized when no algorithm is specified). The traffic simulation agent can integrate various TFMs, including but not limited to modules for database access, traffic flow counting, trajectory extraction, traffic performance evaluation, traffic resilience quantification, data visualization, traffic signal optimization, and traffic simulation. However, different TFMs may have similar functionalities. For example, for signal optimization tasks, traffic foundation models may include Webster's algorithm, linear programming methods, and reinforcement learning techniques. These tools have their own applicable scenarios and different input modes. Therefore, ensuring the LLM can clearly identify and skillfully apply TFMs can be achieved by creating prompts for each TFM and inputting them into the prompt management system.

[0044] The method of use and advantages of this invention: The traffic decision-making collaborative optimization system based on generative AI operates as follows:

[0045] like Figure 1 and Figure 2As shown, this illustrates vehicle path planning in a normal scenario. Vehicle path planning is a classic combinatorial optimization problem. Traditional research methods (exact algorithms, approximate algorithms, heuristic algorithms) fail to fully utilize the similar internal structures within the problem. Therefore, applying machine learning to combinatorial optimization has become a research hotspot in recent years. However, mainstream reinforcement learning methods suffer from sparse reward problems and insufficient exploration space. Variational autoencoders (VAEs), on the other hand, can apply neural networks to reasoning problems, solving the problem of continuous data generation. They offer advantages such as fast training speed in both reasoning and generation problems. The solution (path sequence) to the vehicle path problem is essentially a latent variable in a discrete combinatorial space. Directly modeling its true posterior distribution is infeasible (NP-hard). By approximating the true posterior through a learnable variational distribution, the reasoning problem is transformed into optimization. The problem (minimizing KL divergence) can therefore be modeled as a variational probability problem. Path probabilities are modeled using a variational autoencoder, and combined with the graph structure processing capabilities of GNNs and policy optimization through reinforcement learning, efficient path planning can be achieved. The path planning algorithm based on variational inference and reinforcement learning includes four core stages: graph encoding → subgraph decomposition → variational inference → reinforcement learning optimization. A simple modeling example of the vehicle path planning problem will be given below, briefly explaining the optimization principle. The vehicle path planning problem is a combinatorial optimization problem on a fully directed graph G = (V, E). Assume there is one central node (represented by 0) and n customer nodes. The goal of the planning problem is to minimize the total travel distance of the vehicles while satisfying the capacity constraint. The solution π of the vehicle path planning problem is considered as a latent variable, and the posterior distribution is approximated through variational inference. To represent the probability of path π given a problem instance s, use Let represent the probability of selecting node vi in ​​path π. This allows us to obtain the variational probability generation model and its objective function.

[0046] In recent years, with the increasing complexity of urban buildings and the density of people, the safe evacuation of large public places has faced severe challenges. In the event of emergencies such as fires and terrorist attacks, traditional evacuation systems often suffer from low evacuation efficiency and even secondary injuries due to problems such as dynamic environmental changes, panic behavior among the crowd, and insufficient real-time path information. Traditional evacuation path planning methods mainly include A*, APF, DQN, and combinations of these methods. Generative AI technology can be used to quickly generate multiple candidate paths in dynamic fire environments, while optimizing potentially conflicting objectives such as "path length," "safe distance," and "exit capacity," and can be generated in a lightweight manner. Models replace some numerical calculations to accelerate decision-making. The following section will present the optimizations of generative AI technology over traditional methods in the specific scenario of indoor evacuation during a fire. From the overall algorithm module perspective, path generation is the first step, followed by multi-objective optimization and adaptation. In the path generation part, probability distribution maps of paths can be generated using methods such as conditional diffusion models. Path connectivity constraints can also be added, and gradient correction can be used during reverse denoising to force obstacle avoidance for improvement. In the multi-objective optimization and adaptation part, Pareto solutions can be generated through variational autoencoders (VAEs) and reinforcement learning fine-tuning. VAEs can compress the multi-objective space into the latent space z ~ Encoder(L). i D min ,i,C i ), The fine-tuning part of reinforcement learning requires using the PPO algorithm to optimize the decoder output and maximize the composite reward. Additionally, dynamic weight adjustments can be made; for example, by combining an LSTM predictor to forecast the fire spread rate in real time, the target weights w can be adjusted. safe =Sigmoid(f LSTM (I fireWith the exponential growth of motor vehicle ownership, urban transportation systems are facing unprecedented operational pressure. Against this backdrop, intersections, as key nodes in the road network, have become hotspots for traffic congestion. Controlling traffic lights to reduce intersection congestion is crucial. Traffic signal control problems in the transportation field can be divided into single-intersection and multi-intersection problems. For simplicity, this paper only discusses single-intersection traffic signal control. Traditional traffic signal control methods (such as fixed-time, SCOOT, etc.) are based on rules or static optimization, making it difficult to effectively respond to dynamic traffic demands. Reinforcement learning-based methods can dynamically adjust traffic signals according to traffic conditions, but they often assume that traffic data from the intersection is complete and continuous. However, in reality, this assumption is often challenged due to a lack of sensors or errors. Therefore, generative AI technology in traffic signal control mainly falls into two categories: one is improving intersection traffic signals by recovering missing data; the other is optimizing the corresponding decision-making process using generative AI. The traffic signal control problem can be modeled as a partially observable Markov decision process (POMDP), observing the number of vehicles and queue length in each entrance lane of the local intersection (observation data may be partially missing), the action being the phase of the traffic signal, and the reward being defined as the sum of queue lengths for all entrance lanes, used to optimize the signal control strategy (reward data may also be partially missing). To transform the POMDP problem into a generative task of a diffusion model, historical observations, actions, and reward sequences are combined into a trajectory τ, and Gaussian noise is gradually added to the trajectory τ through a Markov chain. The simulation of the diffusion process of missing data allows for traffic data imputation and decision-making during the reverse generation process. By conditionalizing only partially observable rewards, it avoids confusion between imputed values ​​and actual rewards, thereby improving decision accuracy. The core formula used is p. θ (x 0 (τ)|y(τ))=p θ (x 0 (τ obs )|r(τ),y'(τ))·p θ (x 0 (τ miss )|y'(τ)),τ obs It is observable data, τ miss The missing data in traffic signal control's multi-intersection cooperation problem can be deeply integrated with communication mechanisms and spatiotemporal dependency modeling. The core lies in leveraging the generative capabilities of diffusion models and structured attention mechanisms to jointly capture road network dynamics. The diffusion communication mechanism is primarily responsible for observation propagation; during inference, it incorporates locally generated future observations. Send to adjacent intersection Adjacent intersections utilize the received The optimization of the generation process, specifically the spatiotemporal dependency modeling part, can be divided into three modules. The communication cross-attention module is dedicated to analyzing the interaction patterns between the local intersection and adjacent intersections. If upstream intersections lack data due to sensor malfunctions, it reconstructs the possible traffic flow state by combining historical observations with current local features. Simultaneously, the spatial self-attention module focuses on the correlation between lanes within the intersection, while the temporal self-attention module learns the periodic patterns of traffic flow. This multi-granularity modeling ensures that the generated observation data conforms to physical constraints and can adapt to dynamically missing scenarios. Specifically, the generated intersection observation information can be dynamically propagated to adjacent nodes during the inference phase through a diffusion communication mechanism. The spatiotemporal transformer explicitly models the spatiotemporal dependencies between intersections through a hierarchical attention module. This collaborative design solves the problem of local decision-making bias caused by data gaps in traditional methods. Diffusion generation plays a dual role in this process: on the one hand, partial reward condition diffusion is guided by classifier independence, relying only on observable rewards to generate high-reward trajectories, avoiding noise interference introduced by traditional filling methods; on the other hand, the diffusion communication mechanism uses the reverse process of the diffusion model to generate virtual observations of adjacent intersections, and integrates this information into local decision-making through the communication cross-attention module of the spatiotemporal transformer.

[0047] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A traffic decision-making collaborative optimization system based on generative AI, characterized in that: The traffic decision-making collaborative optimization system includes: S1: Natural language input; S2: Notification Management; S3: Natural Language Understanding and Task Planning; S4: Traffic infrastructure model execution; S5: Results Output and Intermediate Responses; S6: Mission Assessment and Continuation; S7: Final response generation; S8: Dialogue memory storage; In S1, users input specific simulation task requirements using natural language through the traffic simulation intelligent agent front-end interface. This input text is passed as a prompt to the next step for prompt management.

2. The traffic decision-making collaborative optimization system based on generative AI according to claim 1, characterized in that: In S2, as a foundational step, the prompt management needs to define the agent's working mechanism, clarify key considerations, and convey information about the available toolset. Furthermore, this step integrates historical dialogue context to support multi-turn interactions. The components of this integrated prompt include: user task requests, system prefixes, available tools, inference history, and dialogue history. By integrating these elements into a coherent prompt, the agent obtains the necessary context and instructions, effectively facilitating task decomposition and execution.

3. The traffic decision-making collaborative optimization system based on generative AI according to claim 2, characterized in that: In S3, leveraging the capabilities of LLM, agents can understand natural language cues and perform deductive reasoning by integrating task requests, available toolsets, and reasoning history libraries, ultimately forming recognizable and actionable insights.

4. The traffic decision-making collaborative optimization system based on generative AI according to claim 3, characterized in that: In S4, based on established insights, the agent invokes a pre-defined traffic infrastructure model and strictly adheres to the prerequisites specified in the tool definition when setting parameters. The traffic infrastructure model then executes specific tasks based on the parameters set by the agent, including database retrieval and analysis, data visualization, and system optimization, ultimately generating the required output.

5. The traffic decision-making collaborative optimization system based on generative AI according to claim 3, characterized in that: In S4, based on established insights, the agent invokes a pre-defined traffic infrastructure model and strictly adheres to the prerequisites specified in the tool definition when setting parameters. The traffic infrastructure model then executes specific tasks based on the parameters set by the agent, including database retrieval and analysis, data visualization, and system optimization, ultimately generating the required output.

6. The traffic decision-making collaborative optimization system based on generative AI according to claim 5, characterized in that: In S6, the agent compares and analyzes the user's task request with the current intermediate response to assess the task completion status. If the task is not yet resolved, the process will return to steps 2 to 5 to ensure the iterative continuation of the execution process.

7. The traffic decision-making collaborative optimization system based on generative AI according to claim 6, characterized in that: In S7, after confirming the completion of the task in step 6, the agent uses the powerful capabilities of LLM to integrate the output content generated by the tool, form a conclusive response, and transmit it to the user through the front-end interface.

8. The traffic decision-making collaborative optimization system based on generative AI according to claim 7, characterized in that: In S8, user input and LLM output are stored to preserve the ongoing dialogue. These records are aggregated into dialogue history and used as part of the prompting management input in subsequent interactions, providing dialogue context for large language models and enabling them to have memory capabilities.