A method and system for controlling ramp traffic on a highway based on physical information reinforcement learning

By combining the METANET physical model and the TD3 algorithm, and introducing a physical loss term and a hybrid experience replay mechanism, the problems of model error and reinforcement learning instability in highway ramp control are solved, realizing collaborative optimization control of ramp traffic flow and improving the operational efficiency and stability of the traffic system.

CN122416756APending Publication Date: 2026-07-17CHANGSHA UNIVERSITY OF SCIENCE AND TECHNOLOGY +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-27
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing technologies for highway ramp control suffer from model errors and parameter mismatches, making it difficult to balance real-time performance and robustness in complex scenarios. Traditional reinforcement learning, on the other hand, faces problems such as low sample efficiency, unstable training, and insufficient policy interpretability.

Method used

By combining the METANET macroscopic traffic flow physical model with the TD3 algorithm, and by introducing a physical loss term and a hybrid experience replay mechanism, the reinforcement learning model is trained by integrating real and virtual experiences, ensuring the consistency and interpretability of the strategy.

Benefits of technology

It improves sample efficiency and training stability, generates stable and efficient ramp control strategies in complex traffic scenarios, reduces total time delay, and improves the operational efficiency and stability of the traffic system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122416756A_ABST
    Figure CN122416756A_ABST
Patent Text Reader

Abstract

本发明公开了一种基于物理信息强化学习的高速公路匝道车流控制方法及系统,建立高速公路交通控制环境与强化学习控制环境后获取高速公路各路段的状态信息,使用TD 3算法对模型进行训练迭代,训练中将当前状态与控制动作在交通仿真环境中执行得到下一状态并根据综合奖励函数计算奖励得到真实经验,同时使用METANET模型在相同状态与动作下进行虚拟推演得到虚拟下一状态并根据引入物理损失项的综合奖励函数计算奖励得到虚拟经验,最后混合采样真实经验与虚拟经验更新模型参数;训练迭代结束后将稳定的控制策略应用于高速公路匝道仿真环境中。本发明提高了强化学习控制策略的稳定性和可解释性,增强了高速公路交通系统的运行效率和稳定性。
Need to check novelty before this filing date? Find Prior Art