Extensible multi-strategy collaborative optimization method for highway traffic flow management and control
By strengthening the learning agent method, defining the action space, state space and reward functions, building a simulation training environment, and optimizing the agent strategy, the problem of multi-strategy collaborative optimization in highway traffic flow control is solved, and the safety and efficiency of traffic flow control is improved.
Patent Information
- Application Number
- CN202510303391.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-03-14
AI Technical Summary
In the existing technology, in the highway traffic flow control, there is a gap in the coordinated optimization of multiple control strategies, which is difficult to effectively improve traffic safety and efficiency.
The reinforcement learning agent method is adopted to define action space, state space and reward functions, build a simulation training environment, optimize the agent strategy, and obtain the optimal traffic flow control strategy, which is suitable for a variety of control scenarios.
Multi-strategy collaborative optimization has been achieved, the effect of smart highway traffic flow control has been improved, and it is adapted to different traffic flow scenarios, with strong scalability, and equipment updates do not require changes to the overall architecture.
Smart Images

Figure CN120260267A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of highway traffic control, and particularly relates to a multi-strategy collaborative optimization method for extensible highway traffic flow control. Background Art
[0002] With the development of intelligent transportation, its application in highway traffic flow control has received increasing attention. Highway perception devices collect traffic states and meteorological states through various sensors, and process the collected data through information devices and software platforms. Finally, at the application layer, it is presented in different ways to improve traffic service quality. As an important part of the application layer, the traffic flow control system makes decisions on control strategies according to traffic flow states, and publishes them through information release devices to effectively regulate traffic flow states. With the development and application of traffic flow control technologies, traffic flow control strategies have been continuously enriched. How to effectively coordinate various control strategies and further improve highway traffic safety and efficiency is of great significance.
[0003] There are also some existing studies on traffic flow control. For example, the patent with the application number 2022114724056 discloses a highway traffic control method based on reinforcement learning, which mainly uses the data-driven method of reinforcement learning to optimize a single traffic control strategy. The patent with the application number 2024105407291 discloses a highway ramp control method based on reinforcement learning and event-triggered prediction, which mainly uses reinforcement learning and model predictive control to optimize a single ramp control strategy. Although the above existing technologies can all achieve certain results, the relevant content only focuses on a single traffic control strategy, and there is a gap from the co-application of multiple control strategies in the actual application process.
[0004] In view of the above technical problems, the present invention proposes a multi-strategy collaborative optimization method for extensible highway traffic flow control. Summary of the Invention
[0005] The purpose of the present invention is to provide a multi-strategy collaborative optimization method for extensible highway traffic flow control in view of the deficiencies of the prior art.
[0006] To achieve the above purpose, the present invention adopts the following technical solutions:
[0007] A multi-strategy collaborative optimization method for extensible highway traffic flow control includes:
[0008] S1. Define a corresponding reinforcement learning action space based on different types of traffic flow control strategies;
[0009] S2. Define a corresponding reinforcement learning state space based on traffic flow observation states;
[0010] S3. Define the reward function of reinforcement learning based on different target coefficients;
[0011] S4. Construct a simulation training environment;
[0012] S5. Construct a reinforcement learning agent policy based on the action space, state space, reward function, and simulation training environment;
[0013] S6. Optimize the constructed reinforcement learning agent policy in the simulation training environment to obtain the optimal reinforcement learning agent policy;
[0014] S7. Obtain the actual environmental state, and calculate the optimal traffic flow control policy for the actual environmental state according to the optimal reinforcement learning agent policy.
[0015] Further, the step S1 includes:
[0016] S11. Obtain the parameters corresponding to different types of control policies;
[0017] S12. Obtain the types of traffic flow control policies that can be implemented on the road section;
[0018] S13. Define the reinforcement learning action space according to the types of traffic flow control policies that can be implemented on the road section and the corresponding parameters.
[0019] Further, the reinforcement learning action space in the step S13 is expressed as:
[0020]
[0021] Among them, represents the action space; a 限速 represents the parameter of the dynamic speed limit control policy; a 硬路肩 represents the parameter of the hard shoulder opening control policy; a 动态车道 represents the parameter of the dynamic lane management control policy; a 匝道 represents the parameter of the ramp flow regulation control policy; a 动态路径 represents the parameter of the dynamic path guidance control policy; a 其他 represents the parameter of other control policies.
[0022] Further, the step S2 includes:
[0023] S21. Obtain the state parameters corresponding to the traffic flow observation state;
[0024] S22. Obtain the state observation points of the road section;
[0025] S23. Define the reinforcement learning state space according to the state observation points of the road section and the corresponding state parameters.
[0026] Further, in step S23, the reinforcement learning state space is expressed as:
[0027]
[0028] Among them, represents the state space; s1 represents the state parameter of the first observation point; s n represents the state parameter of the nth observation point; s N represents the state parameter of the Nth observation point.
[0029] Further, in step S3, the different objective coefficients include the total travel time of vehicles on the road section and the ramp queue length penalty factor;
[0030] The total travel time of vehicles on the road section is expressed as:
[0031]
[0032] Among them, L T represents the total travel time of vehicles on the road section; T represents the total travel time of vehicles; i represents a certain road section; I all represents the set of all road sections; L i represents the length of road section i; λ i represents the number of lanes of road section i; ρ i (k) represents the traffic flow density of road section i in time period k; I on represents the set of road sections connected to the entrance ramp; w i (k) represents the queue length of the entrance ramp of road section i in time period k;
[0033] The ramp queue length penalty factor is expressed as:
[0034]
[0035] Among them, L W represents the queue length penalty factor; σ i (k) represents the queue length penalty factor of road section i in time period k;
[0036] The reward function of the reinforcement learning in step S3 is expressed as:
[0037]
[0038] Among them, represents the reward function; L O represents other objective function coefficients.
[0039] Further, in step S4, the construction of the simulation training environment is expressed as:
[0040]
[0041] Among them, represents the state space of the next time period k + 1; f tra represents the calculation function of the state space; represents the state space of the current time period k; represents the action space of the current time period k; represents the reward function of the current time period k; g tra represents the calculation function of the reward function.
[0042] Furthermore, the step S5 includes:
[0043] S51. Initialize the policy parameters of the reinforcement learning algorithm and initialize the number of training times;
[0044] S52. Initialize the state space and the reward function, and initialize the number of iterations;
[0045] S53. Obtain the action space of the current time period based on the parameters of the traffic flow control policy type;
[0046] S54. Calculate the state space of the next time period and the reward function of the current time period according to the simulation training environment;
[0047] S55. Update the policy parameters of the reinforcement learning algorithm based on the action space of the current time period, the state space of the current time period, the state space of the next time period, and the reward function of the current time period;
[0048] S56. Determine whether the number of iterations is less than the total number of iterations. If so, execute step S53; if not, execute step S57;
[0049] S57. Determine whether the number of training times is less than the total number of training times. If so, execute step S52; if so, execute the end to obtain the policy of the reinforcement learning agent.
[0050] Furthermore, obtaining the action space of the current time period based on the parameters of the traffic flow control policy type in the step S53 is expressed as:
[0051]
[0052] Among them, represents the action space of the current time period k; π θ represents the policy of the reinforcement learning agent; represents the state space of the current time period k;
[0053] Updating the policy parameters of the reinforcement learning algorithm based on the action space of the current time period, the state space of the current time period, the state space of the next time period, and the reward function of the current time period in the step S55 is expressed as:
[0054]
[0055] Among them, θ represents the algorithm policy parameter of the reinforcement learning agent; represents the state space at the next time period k + 1; represents the reward function at the current time period k; Q represents the algorithm policy parameter update function of the reinforcement learning agent.
[0056] Further, the calculation of the optimal traffic flow control strategy for the actual environmental state in step S7 is expressed as:
[0057] a * = π * (s real )
[0058] Among them, a* represents the optimal traffic flow control strategy; π* represents the optimal reinforcement learning agent strategy; s real represents the actual environmental state.
[0059] Compared with the prior art, the present invention proposes a multi-strategy collaborative optimization method for highway traffic flow control, uses a reinforcement learning agent to make intelligent decisions to determine the strategy type and parameters, constructs a simulation environment to continuously train the agent during this process, and designs a reasonable reward and action space to improve the decision-making effect of the reinforcement learning agent. The trained reinforcement agent can be deployed to the traffic flow control system to effectively improve the intelligent highway traffic flow control effect. Thanks to the reasonable construction of the action space and state space of the reinforcement learning agent, the present invention can be applied to various traffic flow control scenarios. Only by adding the strategies and obtainable states of the control scenario to the corresponding action space and state space and training the agent, the collaborative optimization algorithm can be obtained. There is no need to adjust the optimization algorithm architecture and content function, and the overall scalability is relatively strong. Especially after the update of the section control equipment, a large number of sensing and information publishing devices are added. The technical method provided by the present invention can expand new actions and states without changing the overall architecture. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 is a flowchart of a multi-strategy collaborative optimization method for scalable highway traffic flow control provided in Embodiment 1;
[0061] Figure 2 is a schematic diagram of a highway section with multi-control strategy device deployment provided in Embodiment 2;
[0062] Figure 3 is a schematic diagram of a highway section with multi-control strategy device deployment provided in Embodiment 3;
[0063] Figure 4It is a schematic diagram of the expressway section where the multi-control strategy device provided in the fourth embodiment is deployed. Detailed implementation manners
[0064] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0065] The purpose of the present invention is to provide a multi-strategy collaborative optimization method for expandable expressway traffic flow control in view of the defects of the prior art.
[0066] Embodiment 1
[0067] This embodiment provides a multi-strategy collaborative optimization method for expandable expressway traffic flow control, as Figure 1 shown, including:
[0068] S1. Define a corresponding reinforcement learning action space based on different types of traffic flow control strategies;
[0069] S2. Define a corresponding reinforcement learning state space based on traffic flow observation states;
[0070] S3. Define a reward function for reinforcement learning based on different objective coefficients;
[0071] S4. Construct a simulation training environment;
[0072] S5. Construct a reinforcement learning agent strategy based on the action space, state space, reward function, and simulation training environment;
[0073] S6. Optimize the constructed reinforcement learning agent strategy in the simulation training environment to obtain an optimal reinforcement learning agent strategy;
[0074] S7. Obtain the actual environmental state, and calculate the optimal traffic flow control strategy for the actual environmental state according to the optimal reinforcement learning agent strategy.
[0075] In step S1, a corresponding reinforcement learning action space is defined based on different types of traffic flow control strategies; specifically including:
[0076] S11. Obtain parameters corresponding to different types of control strategies; among them, the control strategies include dynamic speed limit, hard shoulder opening, dynamic lane management, ramp flow regulation, dynamic path guidance, and other control strategies; the action space parameters corresponding to the control strategies are shown in Table 1 below:
[0077] Table 1 Action space parameters involved in different types of control strategies
[0078] Control strategy type Action space parameter Dynamic speed limit Speed limit value, control range (start and end positions), control period (start and end times) Hard shoulder opening Open / close instruction, control range (start and end positions), control period (start and end times) Dynamic lane management Lane open / close instruction, control range (start and end positions), control period (start and end times) Ramp flow regulation Regulation rate, controlled ramp position, control period (start and end times) Dynamic route guidance Route selection, controlled guidance position, control period (start and end times) Other control strategies Main parameters of control strategy, action period of control strategy, action position of control strategy
[0079] S12. Obtain the types of traffic flow control strategies that can be implemented on the road section;
[0080] S13. Define the reinforcement learning action space according to the types of traffic flow control strategies that can be implemented on the road section and the corresponding parameters; where the reinforcement learning action space is expressed as:
[0081]
[0082] Among them, represents the action space; a 限速 represents the parameters of the dynamic speed limit control strategy; a 硬路肩 represents the parameters of the hard shoulder opening control strategy; a 动态车道 represents the parameters of the dynamic lane management control strategy; a 匝道 represents the parameters of the ramp flow regulation control strategy; a 动态路径 represents the parameters of the dynamic route guidance control strategy; a 其他 represents the parameters of other control strategies.
[0083] In step S2, define the corresponding reinforcement learning state space based on the traffic flow observation state; specifically including:
[0084] S21. Obtain the state space parameters corresponding to the traffic flow observation state; where the state space parameters mainly include the cross-sectional average speed, cross-sectional average flow, road section vehicle density, road section queue length, weather parameters, occupation positions, etc.
[0085] S22. Obtain the state observation points in the road section, and the state space parameters that can be obtained at each observation point;
[0086] S23. Define the reinforcement learning state space according to the state observation points in the road section and the state parameters that can be obtained at each observation point; where the reinforcement learning state space is expressed as:
[0087]
[0088] Among them, represents the state space; s1 represents the state parameters of the first observation point; s n represents the state parameters of the nth observation point; s N represents the state parameters of the Nth observation point.
[0089] In step S3, define the reward function of the reinforcement learning based on different objective coefficients; specifically including:
[0090] S31. Define the calculation methods for different objective coefficients; among them, different objective coefficients include the total travel time of vehicles on the road section and the ramp queue length penalty factor; the objective coefficients sometimes also include other objective function coefficients;
[0091] The total travel time of vehicles on the road section is expressed as:
[0092]
[0093] where, L T represents the total travel time of vehicles on the road section; T represents the total travel time of vehicles; i represents a certain road section; I all represents the set of all road sections; L i represents the length of road section i; λ i represents the number of lanes of road section i; ρ i (k) represents the traffic flow density of road section i in time period k; I on represents the set of road sections connected to the entrance ramp; w i (k) represents the queue length of the entrance ramp of road section i in time period k;
[0094] The ramp queue length penalty factor is expressed as:
[0095]
[0096] where, L W represents the queue length penalty factor; σ i (k) represents the queue length penalty factor of road section i in time period k;
[0097] S32. Define the reinforcement learning reward function according to different objective coefficients, expressed as:
[0098]
[0099] where, represents the reward function; L O represents other objective function coefficients.
[0100] In this embodiment, the control strategy in the action space is linked to the traffic flow observation state in the state space through the set reinforcement learning reward function, providing an incentive signal for decision-making for the intelligent agent. The reward function not only evaluates the behavior of the intelligent agent based on the action space parameters, but also adjusts the reward value according to the changes in the state space parameters, so as to guide the intelligent agent to select the optimal control strategy in the dynamic traffic environment to achieve the optimization goal.
[0101] In step S4, a simulation training environment is constructed.
[0102] The simulation training environment is the basis for training reinforcement learning agents. It simulates the real operation of highway traffic flow, providing the agents with interaction scenarios similar to the actual environment. Through the simulation environment, the agents can try different control strategies in the virtual environment and learn and optimize according to the feedback of the reward function.
[0103] The construction method of the simulation training environment includes:
[0104] S41. Divide the simulation section according to the state observation points and the implementation scope of the control strategy;
[0105] State observation points: According to the state space parameters defined in step S2 (such as average cross-section speed, average cross-section flow, road section vehicle density, etc.), determine the positions of the observation points on the simulation section. These observation points should cover key areas, such as entrance ramps, exit ramps, main roads, and hard shoulders.
[0106] Implementation scope of the control strategy: According to the type of control strategy defined in step S1 (such as dynamic speed limit, hard shoulder opening, ramp flow regulation, etc.), determine the implementation scope of each strategy. For example, the dynamic speed limit strategy may be applied to specific sections of the main road, and the ramp flow regulation strategy is applied to the entrance ramp.
[0107] Divide the simulation section: Divide the simulation section into multiple sub-regions, and each sub-region corresponds to one or more observation points and the implementation scope of the control strategy. This division method is convenient for simulating the impact of different strategies on traffic flow in the simulation environment.
[0108] S42. Based on the state space and action space, construct the simulation training environment using a traffic flow mathematical model or simulation software.
[0109] S421. Select a suitable traffic flow model or simulation software.
[0110] The traffic flow mathematical model is based on classical traffic flow theories (such as the Lighthill-Whitham-Richards model, car-following model, etc.) to construct a mathematical model to describe the dynamic characteristics of traffic flow. These models can be used to calculate parameters such as traffic flow speed, flow, and density.
[0111] Simulation software uses professional traffic simulation software (such as VISSIM, AIMSUN, Synchro, etc.). These software provide rich traffic flow modeling functions and can simulate real traffic scenarios, including vehicle driving, traffic signal control, ramp flow regulation, etc.
[0112] S422. Construct the simulation training environment, where the simulation training environment includes road section modeling, traffic flow generation, observation point setting, and control strategy implementation.
[0113] Section Modeling: Construct a simulation section based on the geometric characteristics of the actual highway (such as the number of lanes, length, slope, etc.) and traffic facilities (such as ramps, hard shoulders, etc.).
[0114] Traffic Flow Generation: Generate traffic flow according to actual traffic flow data or prediction models. Different traffic flow scenarios can be set, such as traffic flow during peak hours, off-peak hours, and under adverse weather conditions.
[0115] Observation Point Setting: Set virtual sensors corresponding to the state space observation points defined in step 2 on the simulation section to obtain the state parameters of traffic flow (such as speed, flow, density, etc.) in real time.
[0116] Implementation of Control Strategies: Integrate the control strategies (such as dynamic speed limit, hard shoulder opening, etc.) defined in step S1 into the simulation environment. Through programming interfaces or the control functions of simulation software, dynamic adjustment of these strategies can be achieved.
[0117] In this embodiment, the constructed simulation training environment is represented as:
[0118]
[0119] Among them, represents the state space at the next time period k + 1; f tra represents the calculation function of the state space; represents the state space at the current time period k; represents the action space at the current time period k; represents the reward function at the current time period k; g tra represents the calculation function of the reward function.
[0120] Constructing the simulation training environment in this embodiment is a key link in realizing the training of reinforcement learning agents. By reasonably dividing the simulation section, selecting appropriate traffic flow models or simulation software, and integrating the state space and action space into the simulation environment, a training platform similar to the actual traffic environment can be provided for the agents. The simulation environment can not only simulate the dynamic characteristics of traffic flow but also provide real-time state feedback and reward signals for the agents, thereby realizing the optimization and learning of control strategies.
[0121] In step S5, construct a reinforcement learning agent policy based on the action space, state space, reward function, and simulation training environment; specifically including:
[0122] S51. Initialize the policy parameters θ of the algorithm according to the selected reinforcement learning algorithm, initialize the number of training times e = 0, and set the total number of training times E;
[0123] The reinforcement learning algorithms can be Q-learning, Deep Q-Networks [DQN], Policy Gradient, etc.; the policy parameter θ determines the probability of the agent choosing an action in a given state or the direct mapping; the number of training times determines the total number of rounds of learning of the agent, and the number of iteration times determines the number of steps for the agent to interact with the environment in each round of training.
[0124] S52. Initialize the state space and the reward function, initialize the iteration number k = 0, and set the total number of iterations K;
[0125] Obtain the initial state space from the simulation environment according to the state space parameters (such as cross-sectional average velocity, flow rate, vehicle density, etc.) defined in step S2.
[0126] Initialize the reward function value according to the reward function (such as reducing the total vehicle travel time, reducing the ramp queue length, etc.) defined in step S3.
[0127] S53. Obtain the action space at the current time period based on the parameters of the traffic flow control strategy type.
[0128] According to the current state and the policy parameter π θ , the agent selects an action where the way of selecting the action depends on the reinforcement learning algorithm used, and it is necessary to ensure that the selected action is within the action space defined in step S1 (such as dynamic speed limit value, ramp flow rate adjustment rate, etc.), expressed as:
[0129]
[0130] where, represents the action space at the current time period k; π θ represents the reinforcement learning agent policy; represents the state space at the current time period k;
[0131] S54. Calculate the state space at the next time period and the reward function at the current time period
[0132] Apply the selected action to the simulation training environment, and the simulation training environment calculates the state at the next stage according to the traffic flow model or simulation software
[0133] According to the reward function defined in step S3 calculate the reward value brought by the current action The reward value reflects this action Contribution to the optimization objective (such as reducing travel time, alleviating congestion, etc.).
[0134] S55. Action space based on the current time period State space of the current time period State space of the next time period Reward function of the current time period Update algorithm policy, the parameters of the reinforcement learning algorithm policy, expressed as:
[0135]
[0136] Among them, θ represents the parameters of the reinforcement learning agent algorithm policy; represents the state space of the next time period k + 1; represents the reward function of the current time period k; Q represents the update function of the reinforcement learning agent algorithm policy parameters;
[0137] S56. Determine whether the number of iterations is less than the total number of iterations. If so, execute step S53; if not, execute step S57;
[0138] In this embodiment, the policy parameters are updated according to the feedback to optimize the behavior of the agent;
[0139] S57. Determine whether the number of training times is less than the total number of training times. If so, execute step S52; if so, end and obtain the reinforcement learning agent policy.
[0140] In this embodiment, through multiple trainings and iterations, the optimal policy is gradually approximated.
[0141] According to the above operations, the agent can learn how to select the optimal control strategy combination in a complex traffic flow environment, so as to realize the optimal management of highway traffic flow.
[0142] In step S6, the constructed reinforcement learning agent policy is optimized in the simulation training environment to obtain the optimal reinforcement learning agent policy.
[0143] In this embodiment, the reinforcement learning algorithm designed in step S5 is applied to the simulation training environment in step S4, and through training, the agent can learn the optimal traffic flow control strategy; specifically including:
[0144] S61. Deploy the simulation training environment;
[0145] Deploy the simulation training environment constructed in step S4 to the training system. Ensure that the simulation environment can real-time feedback the state information of the traffic flow (such as speed, flow, vehicle density, etc.), and support the action execution of the agent (such as dynamic speed limit, ramp flow regulation, etc.).
[0146] Adjust the parameters of the simulation environment according to actual needs, such as traffic flow generation rate, road segment length, observation point location, etc., to simulate different traffic scenarios (such as peak hours, bad weather, etc.).
[0147] Ensure that the reinforcement learning algorithm can be seamlessly connected to the interface of the simulation environment to obtain state information and execute actions.
[0148] S62. Select different types of reinforcement learning agent algorithm strategy parameter update functions Q for training;
[0149] According to the complexity of the problem and actual needs, select a suitable reinforcement learning algorithm and initialize relevant parameters according to the selected algorithm. For example:
[0150] Learning rate (α): Controls the speed of parameter update.
[0151] Discount factor (γ): Determines the weight of future rewards.
[0152] Exploration rate (ε): Controls the balance between exploration and exploitation of the agent.
[0153] Neural network weights: If using DQN or Policy Gradient, initialize the weights of the neural network.
[0154] Other hyperparameters: Such as the update frequency of the target network (DQN), batch size (for training the neural network), etc.
[0155] S63. After training, select the best-performing reinforcement learning agent strategy π*.
[0156] According to the algorithm logic designed in step S5, conduct multiple rounds of training to obtain the reinforcement learning agent strategy.
[0157] This embodiment also includes recording key metrics during the training process, such as changes in reward values, convergence of strategies, training time, etc.; dynamically adjusting hyperparameters such as learning rate and exploration rate according to the performance during training to improve the training effect; regularly saving the strategy parameters during the training process so that training can be resumed in case of training interruption.
[0158] According to the performance evaluation results during training, select the best-performing reinforcement learning strategy. The following criteria can be considered:
[0159] Reward maximization: Select the strategy with the highest average reward value.
[0160] Traffic flow optimization: Select the strategy that best optimizes traffic flow metrics (such as congestion mitigation, travel time reduction).
[0161] Stability: Select the strategy with stable convergence of strategy parameters and good robustness.
[0162] Save the optimal policy: Save the parameters of the optimal policy for subsequent deployment and practical applications.
[0163] Through the above operations in this embodiment, the agent can learn how to select the optimal control strategy under complex and changeable traffic flow conditions in the virtual environment, thereby providing effective decision-making support for the actual highway traffic flow control.
[0164] In step S7, obtain the actual environmental state, and calculate the optimal traffic flow control strategy for the actual environmental state according to the optimal reinforcement learning agent policy.
[0165] The deployment of the reinforcement learning agent is a key link in applying the trained and verified optimal reinforcement learning agent policy to the actual traffic flow control system. Its goal is to deploy the optimal policy learned by the agent in the simulation environment to the actual highway informatization system, enabling it to optimize the traffic flow control strategy in real time in the real environment, specifically including:
[0166] S71. Deploy the reinforcement learning agent policy π* with the best effect to the local informatization system;
[0167] Ensure that the actual highway informatization system has sufficient hardware resources (such as servers, sensor networks, communication devices, etc.) to support the operation of the agent. At the same time, the software environment needs to be compatible with the simulation training environment, capable of receiving the policy output of the agent and executing corresponding control actions.
[0168] Establish data interfaces between the agent and the actual traffic flow monitoring system. These interfaces need to be able to obtain traffic flow state data (such as vehicle speed, flow rate, vehicle density, etc.) in real time, and transmit the decision-making instructions of the agent to traffic control devices (such as speed limit signs, lane indicators, ramp signal lights, etc.).
[0169] Deploy the agent to the local informatization system to ensure that it can be seamlessly integrated with the existing traffic management system (such as traffic signal control system, information release system, etc.).
[0170] S72. According to the collected actual environmental state, calculate the optimal traffic flow control strategy for the actual environmental state, expressed as:
[0171] a * =π * (s real )
[0172] where a* represents the optimal traffic flow control strategy; π* represents the optimal reinforcement learning agent policy; s real represents the actual environmental state.
[0173] In the actual operating environment, the policy parameters of the agent are calibrated. Since there may be differences between the actual environment and the simulation environment (such as traffic flow, road conditions, driving behavior, etc.), short-term test runs are required to adjust the policy parameters to ensure their adaptation to the actual environment. Initial state data (such as section vehicle density, average speed, ramp queue length, etc.) is obtained from the actual traffic flow monitoring system and used as the input for the agent's decision-making. Small-scale tests are conducted in the actual environment to verify the effectiveness of the agent's policy and the stability of the system. Manual intervention can be carried out during the test process to ensure the safety and reliability of the system.
[0174] According to the current traffic flow state, the agent uses the trained policy to calculate the optimal control strategy (such as dynamic speed limit value, ramp flow regulation rate, lane opening / closing instruction, etc.). The control strategy output by the agent is transmitted to the traffic control equipment (such as variable speed limit signs, ramp signal lights, lane indicators, etc.) through the communication interface, and these devices perform the corresponding control actions. According to the changes in the real-time traffic flow, the agent can dynamically adjust the control strategy. For example, when congestion is detected on a certain section, the agent can adjust the speed limit value or open the hard shoulder in real time to relieve the congestion.
[0175] This embodiment also includes monitoring and feedback on the agent, specifically:
[0176] The operating state of the agent and the changes in the traffic flow are monitored in real time. Key traffic indicators (such as average vehicle speed, queue length, congestion index, etc.) are displayed through a visualization interface so that operators can understand the system operation in a timely manner. The performance of the agent's policy is evaluated regularly to check whether the expected optimization goals (such as reducing the total vehicle travel time, reducing the ramp queue length, etc.) are achieved. If a performance decline or abnormal situation is found, the policy parameters are adjusted in a timely manner or system maintenance is carried out. A feedback mechanism is established to allow operators to manually intervene in the agent's decision-making according to the actual situation. For example, in the event of special events (such as traffic accidents, bad weather), operators can adjust the agent's policy or directly take over the control equipment.
[0177] According to the above method, the agent can dynamically optimize the traffic flow control strategy in the actual highway environment, improving the operation efficiency and safety of the traffic system.
[0178] Compared with the prior art, a multi-strategy collaborative optimization method is proposed for highway traffic flow control in this embodiment. An RL agent is used to make intelligent decisions to determine the strategy type and parameters. During this process, a simulation environment is constructed to continuously train the agent, and a reasonable reward and action space are designed to improve the decision-making effect of the RL agent. The trained RL agent can be deployed to the traffic flow control system to effectively improve the intelligent highway traffic flow control effect. Thanks to the reasonable construction of the action space and state space of the RL agent, the present invention can be applied to various traffic flow control scenarios. Only by adding the strategies and observable states of the control scenario to the corresponding action space and state space and training the agent, the collaborative optimization algorithm can be obtained. There is no need to adjust the optimization algorithm architecture and content function, and the overall scalability is strong. Especially after the update of the section control equipment, a large number of sensing and information publishing devices are added. The technical method provided by the present invention can expand new actions and states without changing the overall architecture.
[0179] Embodiment 2
[0180] The difference between the multi-strategy collaborative optimization method for scalable highway traffic flow control provided in this embodiment and Embodiment 1 lies in:
[0181] As Figure 2 shown, the control strategy types in this embodiment include dynamic speed limit, hard shoulder opening, and ramp flow regulation. Therefore, a corresponding dynamic speed limit information board, a hard shoulder opening information board, and a ramp regulation information board are set on the section.
[0182] S1. Define the RL action space based on the control strategy type, specifically:
[0183] S11. Define the parameter action space parameters involved in different types of control strategies as shown in Table 2 below:
[0184] Table 2 Parameter action space parameters involved in different types of control strategies
[0185] Control strategy type Action space parameter Dynamic speed limit Speed limit value, control range (start and end positions), control period (start and end times) Hard shoulder opening Open / close instruction, control range (start and end positions), control period (start and end times) Ramp flow regulation Regulation rate, controlled ramp position, control period (start and end times)
[0186] S12. Sort out the traffic flow control strategy types that can be implemented on the section. In this embodiment, they include dynamic speed limit, hard shoulder opening, and ramp flow regulation;
[0187] S13. Define the action space according to the traffic flow control strategy types that can be implemented on the section and different control strategy parameters, expressed as:
[0188]
[0189] S2. Define the RL state space based on the traffic flow observation state, specifically:
[0190] S21. Define common state space parameters, mainly including: cross-sectional average speed, cross-sectional average flow rate, road section vehicle density, and road section queue length;
[0191] S22. Sort out the state observation points of the road section and the state parameters that can be obtained at each observation point. This embodiment includes a total of three observation points, and the cross-sectional average speed and cross-sectional average flow rate can be obtained at all three observation points;
[0192] S23. Define the state space according to the state observation points of the road section and the state parameters that can be obtained at each observation point, expressed as:
[0193]
[0194] S3. Define the reinforcement learning reward based on the optimization objective, specifically:
[0195] S31. Define the calculation methods of different objective coefficients;
[0196] a) Calculation formula for the total travel time of vehicles on the road section:
[0197]
[0198] b) Calculation formula for the ramp queue length penalty factor:
[0199]
[0200] S33. In this embodiment, L O represents the action space cost, expressed as:
[0201] L O = ∑a(k)
[0202] Define the reinforcement learning reward according to various objective coefficients, expressed as:
[0203]
[0204] This reward includes three parts: the total travel time of vehicles, the ramp queue length penalty factor, and the action space cost.
[0205] S4. Construct a simulation training environment, specifically:
[0206] S41. Divide the simulation road section according to the state space observation points and the implementation scope of the control strategy.
[0207] S42. For the state space and action space, this embodiment uses simulation software to construct a simulation training environment.
[0208]
[0209] Among them, f tra and g tra can be obtained through simulation software.
[0210] S5. Design of the reinforcement learning agent algorithm, specifically:
[0211] S51. Initialize the policy parameters θ of the reinforcement learning algorithm, initialize the number of training times e = 0, and the total number of training times E = 1000.
[0212] S52. Initialize the state space and the reward function, and initialize the number of iteration times k = 0, and the total number of iteration times K = 200;
[0213] S53. Obtain the action space of the current time period based on the parameters of the traffic flow control policy type, expressed as:
[0214]
[0215] S54. Calculate the state space of the next time period according to the simulation training environment and the reward function of the current time period
[0216] S55. Based on the action space of the current time period the state space of the current time period the state space of the next time period the reward function of the current time period Update the policy parameters of the reinforcement learning algorithm of the algorithm policy, expressed as:
[0217]
[0218] S56. Determine whether the number of iteration times is less than the total number of iteration times K. If so, execute step S53; if not, execute step S57;
[0219] S57. Determine whether the number of training times is less than the total number of training times E. If so, execute step S52; if so, execute the end to obtain the reinforcement learning agent policy.
[0220] In step S6, the training of the reinforcement learning agent is as follows:
[0221] S61. Deploy the simulation training environment constructed in step S4.
[0222] S62. Select different types of reinforcement learning agent algorithm policy parameter update functions Q for training. In this embodiment, three algorithms, namely Q-learning, DQN, and PPO, are used for separate training.
[0223] S63. After the training is completed, select the best-performing reinforcement learning agent policy π*.
[0224] In step S7, the reinforcement learning agent is deployed.
[0225] S71. Deploy the best-performing reinforcement learning agent policy π* to the local information system.
[0226] S72. Calculate the optimal traffic flow control strategy for the actual environmental state based on the collected actual environmental state, expressed as:
[0227] a * = π * (s real )
[0228] where a* includes the parameters of three strategies: dynamic speed limit, hard shoulder opening, and ramp flow regulation; s real includes the state quantities of three observation points.
[0229] Embodiment III
[0230] The difference between the multi-strategy collaborative optimization method for scalable highway traffic flow control provided in this embodiment and Embodiment I is that:
[0231] As Figure 3 shown, the control strategy type of this embodiment includes dynamic lane management, so a corresponding dynamic lane management information board is set on the road section.
[0232] S1. Define the reinforcement learning action space based on the control strategy type, specifically:
[0233] S11. Define the parameter action space parameters involved in different types of control strategies as shown in Table 3 below:
[0234] Table 3 Parameter action space parameters involved in different types of control strategies
[0235]
[0236] S12. Sort out the types of traffic flow control strategies that can be implemented on the road section. In this embodiment, they include dynamic speed limit, hard shoulder opening, and ramp flow regulation;
[0237] S13. Define the action space according to the types of traffic flow control strategies that can be implemented on the road section and different control strategy parameters, expressed as:
[0238]
[0239] S2. Define the reinforcement learning state space based on the traffic flow observation state, specifically:
[0240] S21. Define the common state space parameters, mainly including: cross-section average speed, cross-section average flow, road section vehicle density, road section queue length;
[0241] S22. Sort out the status observation points of the road section and the status parameters that can be obtained at each observation point. In this embodiment, there are a total of two observation points, and both observation points can obtain the cross-section average speed and the cross-section average flow rate;
[0242] S23. Define the state space according to the status observation points of the road section and the status parameters that can be obtained at each observation point, which is expressed as:
[0243]
[0244] S3. Define the reinforcement learning reward based on the optimization objective, specifically:
[0245] S31. Define the calculation methods of different objective coefficients;
[0246] a) The calculation formula for the total travel time of vehicles on the road section:
[0247]
[0248] b) The calculation formula for the ramp queue length penalty factor:
[0249]
[0250] S33. In this embodiment, L is not defined O , then define the reinforcement learning reward according to various objective coefficients, which is expressed as:
[0251]
[0252] This reward includes two parts: the total travel time of vehicles and the ramp queue length penalty factor.
[0253] S4. Construct a simulation training environment, specifically:
[0254] S41. Divide the simulation road section according to the status space observation points and the implementation scope of the control strategy.
[0255] S42. For the state space and the action space, in this embodiment, a traffic flow mathematical model is used to construct the simulation training environment, which is expressed as:
[0256]
[0257] Among them, f tra and g tra can be obtained through simulation software.
[0258] S5. Design the reinforcement learning agent algorithm, specifically:
[0259] S51. Initialize the policy parameters θ of the reinforcement learning algorithm, initialize the number of training times e = 0, and the total number of training times E = 500.
[0260] S52. Initialize the state space and the reward function, and initialize the number of iteration times k = 0, and the total number of iteration times K = 200;
[0261] S53. Obtain the action space of the current time period based on the parameters of the traffic flow control policy type, which is expressed as:
[0262]
[0263] S54. Calculate the state space of the next time period according to the simulation training environment and the reward function of the current time period
[0264] S55. Based on the action space of the current time period the state space of the current time period the state space of the next time period the reward function of the current time period Update the policy parameters of the algorithm policy of the reinforcement learning algorithm, which is expressed as:
[0265]
[0266] S56. Determine whether the number of iteration times is less than the total number of iteration times K. If so, execute step S53; if not, execute step S57;
[0267] S57. Determine whether the number of training times is less than the total number of training times E. If so, execute step S52; if so, execute the end to obtain the policy of the reinforcement learning agent.
[0268] In step S6, the reinforcement learning agent is trained. Specifically:
[0269] S61. Deploy the simulation training environment constructed in step S4.
[0270] S62. Select different types of reinforcement learning agent algorithm policy parameter update functions Q for training. In this embodiment, three algorithms, Q-learning, DDPG, and A3C, are used for separate training.
[0271] S63. After the training is completed, select the best-performing reinforcement learning agent policy π*.
[0272] In step S7, the reinforcement learning agent is deployed.
[0273] S71. Deploy the best-performing reinforcement learning agent policy π* to the local information system.
[0274] S72. Calculate the optimal traffic flow control strategy for the actual environmental state based on the collected actual environmental state, expressed as:
[0275] a * = π * (s real )
[0276] where a* includes the parameters of the dynamic lane management strategy; s real includes the state variables of two observation points.
[0277] Example 4
[0278] The difference between the multi-strategy collaborative optimization method for scalable highway traffic flow control provided in this example and Example 3 lies in:
[0279] As Figure 3 shown, the control strategy types in this example include dynamic lane management and newly added dynamic route guidance. Therefore, a corresponding dynamic lane management information board and a newly added dynamic route guidance information board are set on the road section.
[0280] S1. Define the reinforcement learning action space based on the control strategy type, specifically:
[0281] S11. Define the parameter action space parameters involved in different types of control strategies as shown in Table 4 below:
[0282] Table 4 Parameter action space parameters involved in different types of control strategies
[0283]
[0284] S12. Sort out the types of traffic flow control strategies that can be implemented on the road section. In this example, they include dynamic lane management and newly added dynamic route guidance;
[0285] S13. Define the action space according to the types of traffic flow control strategies that can be implemented on the road section and the parameters of different control strategies, expressed as:
[0286]
[0287] S2. Define the reinforcement learning state space based on the traffic flow observation state, specifically:
[0288] S21. Define the common state space parameters, mainly including: average cross-section speed, average cross-section flow, vehicle density on the road section, and queue length on the road section;
[0289] S22. Sort out the state observation points on the road section and the state parameters that can be obtained at each observation point. In this example, there are a total of four observation points, and all four observation points can obtain the average cross-section speed and average cross-section flow;
[0290] S23. Define the state space based on the status observation points of the road section and the status parameters that can be obtained at each observation point. Add two observation points, denoted as:
[0291]
[0292] S3. Define the reinforcement learning reward based on the optimization objective, specifically:
[0293] S31. Define the calculation methods for different objective coefficients;
[0294] a) The calculation formula for the total travel time of vehicles on the road section:
[0295]
[0296] b) The calculation formula for the ramp queue length penalty factor:
[0297]
[0298] S33. In this embodiment, L is not defined O , then define the reinforcement learning reward according to various objective coefficients, denoted as:
[0299]
[0300] This reward includes two parts: the total travel time of vehicles and the ramp queue length penalty factor.
[0301] S4. Construct a simulation training environment, specifically:
[0302] S41. Divide the simulation road section according to the status space observation points and the implementation scope of the control strategy.
[0303] S42. For the state space and action space, in this embodiment, a traffic flow mathematical model is used to construct the simulation training environment, denoted as:
[0304]
[0305] Among them, f tra and g tra can be obtained through simulation software.
[0306] S5. Design the reinforcement learning agent algorithm, specifically:
[0307] S51. Initialize the policy parameters θ of the reinforcement learning algorithm, initialize the training times e = 0, and the total training times E = 500.
[0308] S52. Initialize the state space and the reward function, and initialize the iteration times k = 0, and the total iteration times K = 200;
[0309] S53. Obtain the action space for the current time period based on the traffic flow control strategy type, expressed as:
[0310]
[0311] S54. Calculate the state space for the next time period according to the simulation training environment and the reward function for the current time period
[0312] S55. Based on the action space for the current time period the state space for the current time period the state space for the next time period the reward function for the current time period Update the algorithm policy reinforcement learning algorithm policy parameters, expressed as:
[0313]
[0314] S56. Determine whether the number of iterations is less than the total number of iterations K. If so, execute step S53; if not, execute step S57;
[0315] S57. Determine whether the number of training times is less than the total number of training times E. If so, execute step S52; if so, end and obtain the reinforcement learning agent policy.
[0316] In step S6, the reinforcement learning agent is trained, specifically:
[0317] S61. Deploy the simulation training environment constructed in step S4.
[0318] S62. Select different types of reinforcement learning agent algorithm policy parameter update functions Q for training. In this embodiment, three algorithms, Q-learning, DDPG, and A3C, are used for separate training.
[0319] S63. After the training is completed, select the best-performing reinforcement learning agent policy π*.
[0320] In step S7, the reinforcement learning agent is deployed.
[0321] S71. Deploy the best-performing reinforcement learning agent policy π* to the local information system.
[0322] S72. Calculate the optimal traffic flow control strategy for the actual environment state based on the collected actual environment state, expressed as:
[0323] a * = π * (s real )
[0324] where a* includes the parameters of the dynamic lane management strategy; s real includes the state quantities of two observation points.
[0325] Note that the above is only the preferred embodiment of the present invention and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein. Various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, more other equivalent embodiments can be included, and the scope of the present invention is determined by the scope of the appended claims.
Claims
1. A multi-strategy collaborative optimization method for extensible expressway traffic flow control, characterized in that, Including: S1. Define the corresponding reinforcement learning action space based on different types of traffic flow control strategies; S2. Define the corresponding reinforcement learning state space based on the traffic flow observation state; S3. Define the reward function of reinforcement learning based on different objective coefficients; S4. Construct a simulation training environment; S5. Construct a reinforcement learning agent policy based on the action space, state space, reward function, and simulation training environment; S6. Optimize the constructed reinforcement learning agent policy in the simulation training environment to obtain the optimal reinforcement learning agent policy; S7. Obtain the actual environment state, and calculate the optimal traffic flow control strategy for the actual environment state according to the optimal reinforcement learning agent policy.
2. The multi-strategy collaborative optimization method for scalable highway traffic flow control according to claim 1, characterized in that, The step S1 includes: S11. Obtain the parameters corresponding to different types of control strategies; S12. Obtain the types of traffic flow control strategies that can be implemented on the road section; S13. Define the reinforcement learning action space according to the types of traffic flow control strategies that can be implemented on the road section and the corresponding parameters.
3. The multi-strategy collaborative optimization method for scalable highway traffic flow control according to claim 2, characterized in that, The reinforcement learning action space in the step S13 is expressed as: Among them, represents the action space; a 限速 represents the dynamic speed limit control strategy parameter; a 硬路肩 represents the hard shoulder opening control strategy parameter; a 动态车道 represents the dynamic lane management control strategy parameter; a 匝道 represents the ramp flow regulation control strategy parameter; a 动态路径 represents the dynamic route guidance control strategy parameter; a 其他 represents other control strategy parameters.
4. The multi-strategy collaborative optimization method for scalable highway traffic flow control according to claim 1, characterized in that The step S2 includes: S21. Obtain the state parameters corresponding to the traffic flow observation state; S22. Obtain the state observation points of the road section; S23. Define the reinforcement learning state space according to the state observation points of the road section and the corresponding state parameters.
5. The multi-strategy collaborative optimization method for extensible highway traffic flow control according to claim 4, wherein The reinforcement learning state space in the step S23 is expressed as: Among them, represents the state space; s1 represents the state parameter of the first observation point; s n represents the state parameter of the nth observation point; s N represents the state parameter of the Nth observation point.
6. The multi-strategy collaborative optimization method for scalable expressway traffic flow control according to claim 1, wherein, The different objective coefficients in the step S3 include the total travel time of vehicles on the road section and the ramp queue length penalty factor; The total travel time of vehicles on the road section is expressed as: Among them, L T represents the total travel time of vehicles on a section; T represents the total travel time of vehicles; i represents a certain section; I all represents the set of all sections; L i represents the length of section i; λ i represents the number of lanes of section i; ρ i ρ(k) represents the traffic flow density of section i in time period k; I on represents the set of sections connected to the on-ramp; w i w(k) represents the queue length of the on-ramp of section i in time period k; The ramp queue length penalty factor is expressed as: Among them, L W represents the queue length penalty factor; σ i (k) represents the queue length penalty factor of section i at time period k; The reward function of reinforcement learning in the step S3 is expressed as: Among them, represents the reward function; L O represents other objective function coefficients.
7. A multi-strategy collaborative optimization method for scalable highway traffic flow control according to claim 1, characterized in that, The construction of the simulation training environment in the step S4 is expressed as: Among them, represents the state space of the next time period k + 1; f tra represents the calculation function of the state space; represents the state space of the current time period k; represents the action space of the current time period k; represents the reward function of the current time period k; g tra represents the calculation function of the reward function.
8. The multi-strategy collaborative optimization method for scalable highway traffic flow control according to claim 2, characterized in that, The step S5 includes: S51. Initialize the policy parameters of the reinforcement learning algorithm and initialize the number of training times; S52. Initialize the state space and reward function, and initialize the number of iterations; S53. Obtain the action space at the current time based on the parameters of the traffic flow control strategy type; S54. Calculate the state space at the next time and the reward function at the current time according to the simulation training environment; S55. Update the policy parameters of the reinforcement learning algorithm based on the action space at the current time, the state space at the current time, the state space at the next time, and the reward function at the current time; S56. Determine whether the number of iterations is less than the total number of iterations. If so, execute step S53; if not, execute step S57; S57. Determine whether the number of training times is less than the total number of training times. If so, execute step S52; if so, execute the end to obtain the reinforcement learning agent policy.
9. The multi-strategy collaborative optimization method for extensible highway traffic flow control according to claim 8, characterized in that, The obtaining of the action space at the current time based on the parameters of the traffic flow control strategy type in the step S53 is expressed as: Among them, represents the action space at the current time period k; π θ represents the reinforcement learning agent policy; represents the state space at the current time period k; The updating of the policy parameters of the reinforcement learning algorithm based on the action space at the current time, the state space at the current time, the state space at the next time, and the reward function at the current time in the step S55 is expressed as: Among them, θ represents the algorithm policy parameters of the reinforcement learning agent; represents the state space at the next time period k + 1; represents the reward function at the current time period k; Q represents the update function of the algorithm policy parameters of the reinforcement learning agent.
10. A multi-strategy collaborative optimization method for scalable highway traffic flow control according to claim 1, characterized in that The calculation of the optimal traffic flow control strategy for the actual environment state in the step S7 is expressed as: a * = π * (s real ) where a* represents the optimal traffic flow control strategy; π* represents the optimal reinforcement learning agent strategy; s real represents the actual environmental state.
Citation Information
Patent Citations
Distributed traffic signal control method based on generative adversarial network and reinforcement learning
CN113436443A
Expressway traffic control method based on reinforcement learning
CN115713860A
Integrated control method for confluence bottleneck area based on dynamic division of lane-level cells
CN117671955A
Expressway single-ramp management and control method considering hard shoulder opening based on reinforcement learning
CN118411834A
Expressway ramp management and control method based on reinforcement learning and event triggering prediction
CN118430252A
Cited By
Intelligent cooperative control method, system and equipment for highway ramps and medium
CN121171044A