The application discloses an automatic driving model confrontation training method based on
double time scales and Nash equilibrium, and belongs to the technical field of automatic driving. The method comprises the following steps: establishing a strategy
pool and setting an asymmetric learning rate; performing closed-loop
simulation and experience collection; in the
inner loop, the number of update steps of a test network is greater than that of a host vehicle network; constructing a metagame payoff matrix and solving a mixed strategy Nash equilibrium; taking the Nash equilibrium
mixed model as a teacher model to perform knowledge
distillation to a single network; calculating three health indicators of
system availability, win rate climbing rate and strategy convergence stability and adaptively adjusting hyperparameters. The application eliminates catastrophic forgetting through a
double time scale asymmetric update mechanism, avoids mode collapse through Nash equilibrium solving, realizes vehicle-mounted low-power deployment through knowledge
distillation, realizes automatic adaptive training through
health indicator monitoring, and significantly improves the stability and robustness of automatic driving model confrontation training.