This invention discloses a lateral balance control method for unmanned bicycles based on adversarial
imitation learning. The method uses the lateral tilt angle, handlebar angle, and
angular velocity of the unmanned bicycle as the
system state, and the handlebar
control torque as the continuous control input. It combines expert demonstration data of the unmanned bicycle under stable lateral balance control conditions, collected or loaded, to initialize a policy network and a
discriminant network, constructing an adversarial
imitation learning framework. During training, the policy network interacts with the dynamic environment of the unmanned bicycle to generate state-action samples. These samples, along with the expert demonstration data, are simultaneously input into the
discriminant network. An
imitation reward
signal is constructed based on the output of the
discriminant network, and a policy
gradient method with
pruning constraints is used to alternately update the policy network and the discriminant network. This allows the control behavior generated by the policy network to gradually approximate the expert demonstration behavior in a distributional sense, thereby obtaining a stable lateral balance control strategy for the unmanned bicycle. By introducing an adversarial
imitation learning mechanism, this invention achieves effective learning of the control strategy without explicitly designing complex artificial reward functions. Furthermore, under disturbance conditions, the method can still maintain the lateral balance of the unmanned bicycle, demonstrating
good control stability and anti-interference ability, and has good
engineering feasibility and application promotion value.