A trajectory generation method based on deep
reinforcement learning for training game coefficients. The advantages of this invention are as follows: By organically combining deep
reinforcement learning and game theory
trajectory planning, this invention achieves automated and intelligent optimization of game coefficients. It automatically learns the
sensitivity coefficient α in the game objective function using algorithms such as DQN, completely replacing traditional manual parameter tuning methods and significantly reducing development costs and time. A reward function centered on the vehicle's forward distance guides the vehicle to pass quickly, while introducing penalty terms such as collision and acceleration to enhance safety and comfort. The generated trajectory outperforms manually tuned results in terms of
traffic efficiency, safety, and smoothness. Simultaneously, it preserves the
interpretability of the game
theory model; the optimized α value has a clear physical meaning and can dynamically adjust the driving style according to the real-time environment, improving the
system's scene adaptability and generalization robustness.