The present application relates to a kind of layered
reinforcement learning driven car-road cooperation automatic driving longitudinal following control method and
system, belong to the field of automatic driving
vehicle control.For the problem that the generalization ability of existing following control method is insufficient, strategy update is easy to lose stability, easy to fall into
local optimum, adopt layered architecture, upper layer adapts intersection, high-speed straight road and ramp multiple typical traffic scenes, based on real-time car-road cooperation communication information
dynamic planning follow target, after normalizing
processing to multi-source heterogeneous traffic characteristics, generate environment
state vector;Lower layer uses proximal policy optimization
algorithm, outputs continuous accelerator or
brake command through actor-critic
double network, and updates network parameters by combining strategy entropy with
pruning loss with Jacobian correction.The method improves the cross-scene generalization ability and
perception robustness, effectively guarantees the stability and monotonicity of strategy update, avoids
instruction step mutation, breaks the exploration
deadlock, and balances
driving safety,
traffic efficiency and driving comfort.