Unmanned bicycle balance control based on adversarial imitation learning
By alternately optimizing the policy network and discriminant network under the adversarial imitation learning framework, the dependence of the lateral balance control of unmanned bicycles on the precise dynamic model is solved, realizing the autonomous and stable control of unmanned bicycles in complex environments and improving the anti-interference ability and control stability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUILIN UNIV OF ELECTRONIC TECH
- Filing Date
- 2026-04-10
- Publication Date
- 2026-07-03
AI Technical Summary
The lateral balance control method for unmanned bicycles is highly dependent on accurate dynamic models and lacks robustness and stable control capabilities under complex environments and uncertain disturbances.
By employing an adversarial imitation learning approach, an adversarial imitation learning framework is constructed using a policy network and a discriminant network. Expert demonstration data of unmanned bicycles under stable lateral balance control is used to achieve autonomous lateral balance control of unmanned bicycles, reducing the dependence on precise dynamic models and improving stability and anti-interference ability in complex environments.
It improves the anti-interference ability and stable control capability of unmanned bicycles under complex environments and uncertain disturbance conditions, enhances the learning efficiency and stability of control strategies, and achieves continuous, smooth control output that meets physical constraints.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This invention relates to the field of unmanned bicycle balance control, specifically to a method for lateral balance control of unmanned bicycles based on adversarial imitation learning. Background Technology
[0002] Unmanned bicycles possess advantages such as simple structure, high maneuverability, and low energy consumption, and have broad development prospects in applications such as intelligent logistics, short-distance inspection, and urban micro-transportation. However, due to the typical underactuated and unstable system characteristics of unmanned bicycles, the lateral tilt angle balance control problem remains one of the core technical challenges that urgently needs to be solved in the research and engineering application of unmanned bicycles.
[0003] Existing research on lateral balance control for autonomous bicycles mainly focuses on model-based control methods, such as linear quadratic regulators (LQR), active disturbance rejection control (ADRC), partial feedback linearization control, sliding mode control, and model predictive control (MPC). These methods typically rely on relatively accurate dynamic models to design control laws. However, autonomous bicycle systems are inevitably affected by parameter uncertainties, structural coupling, external disturbances, and road surface changes during actual operation. Furthermore, the modeling process often involves assumptions such as pure wheel rolling, flat ground, and ideal sensor measurements, leading to significant discrepancies between the theoretical model and the real system. When environmental conditions change or model parameters become mismatched, traditional model-based control methods often struggle to maintain stable control performance and their robustness is limited in the presence of these issues.
[0004] Reinforcement learning methods, which do not rely on precise mathematical models and can autonomously learn control strategies through continuous interaction with the environment, have been widely used in the field of intelligent control. These methods typically construct state-action-reward interaction mechanisms, allowing the agent to continuously correct its control behavior during trial and error, and optimize the strategy based on cumulative rewards, thereby reducing the impact of model uncertainty on control performance to some extent. However, existing reinforcement learning methods generally suffer from low sample utilization efficiency, slow convergence speed, and instability during training, especially in continuous state-action space tasks such as lateral balance control of unmanned bicycles, where these problems are more pronounced, thus limiting their further application in practical engineering scenarios.
[0005] Imitation learning is a method that directly learns control strategies using expert demonstration data. It establishes a mapping relationship from system state to control input through existing state-action samples, enabling the learned strategy to acquire relatively reasonable control behavior in the early stages of training. This reduces the ineffective exploration caused by the heavy reliance on random trial and error in the early training phase of pure reinforcement learning. Compared to pure reinforcement learning, imitation learning can utilize expert prior knowledge to narrow the policy search range and improve learning efficiency in the early stages of training. However, existing imitation learning methods still suffer from insufficient optimization of long-term control performance when facing complex continuous control tasks. When changes in state distribution, external disturbances, or environmental uncertainties occur during system operation, its control effect and closed-loop stability may still be affected. Therefore, imitation learning alone cannot fully meet the requirements of lateral balance control stability and robustness for unmanned bicycles under complex operating conditions.
[0006] Generative Adversarial Imitation Learning (GAIL) is a method that incorporates reinforcement learning concepts into the imitation learning framework. This method constructs an adversarial training mechanism between a policy network and a discriminator network, enabling the learned policy to approximate expert behavior distributions while continuously optimizing policy parameters based on environmental interactions. Compared to traditional reinforcement learning methods, this approach can improve training efficiency by leveraging expert demonstration data and, to some extent, reduce the difficulty of reward design; compared to general imitation learning methods, it can balance expert behavior approximation with long-term control performance optimization. However, existing GAIL methods still have several shortcomings in the continuous control task of lateral balance control for unmanned bicycles. These shortcomings include instability in the policy update process within a continuous action space, a lack of effective coordination between policy output and the physical constraints of the actuator, and susceptibility to factors such as changes in reward scale, policy distribution shifts, and external disturbances during training. Therefore, this paper proposes an GAIL control method to improve policy training stability and control output executability, which is of great significance for the lateral balance control of unmanned bicycles. Summary of the Invention
[0007] To address the shortcomings of existing lateral balance control methods for unmanned bicycles, such as high dependence on precise dynamic models and insufficient robustness and stability under complex environments and uncertain disturbances, this invention proposes a lateral balance control method for unmanned bicycles based on adversarial imitation learning. This method incorporates expert demonstration data of unmanned bicycles under stable lateral balance control conditions and constructs an adversarial imitation learning framework composed of a policy network and a discriminant network. Without requiring the manual design of complex physically meaningful reward functions, the method guides the control policy to gradually learn the stable lateral balance behavior of the unmanned bicycle through interaction with the environment, thereby achieving autonomous lateral balance control of the unmanned bicycle in complex environments.
[0008] The prototype of the unmanned bicycle and its simplified structural diagram are attached. Figure 1 Appendix Figure 1 In the diagram, B, H, F, and R represent the frame, front fork, front wheel, and rear wheel, respectively; M1, M2, M3, and M4 are the center of mass positions of the corresponding components.
[0009] To facilitate the explanation of the lateral balance control problem of unmanned bicycles, the relevant motion variables are described as follows:
[0010] set up The forward speed of the unmanned bicycle;
[0011] set up and These are the roll angle and angular velocity of the vehicle body, respectively.
[0012] set up and These are the handlebar angle and its angular velocity, respectively.
[0013] set up This is the control input torque acting on the handlebar steering mechanism.
[0014] Assuming an autonomous bicycle is traveling on a flat surface, its lateral balance control problem can be described as follows: by acquiring the current state information of the autonomous bicycle and calculating the handlebar control torque based on the state information. The control torque is applied to the handlebar steering mechanism, causing the handlebars to rotate to the corresponding angle, and the roll angle of the vehicle body is affected through vehicle dynamics coupling. Adjustments are made; during the control process, the roll angle of the vehicle is kept near the equilibrium point through continuous feedback and updates of the system status, thereby achieving lateral balance control of the unmanned bicycle.
[0015] Therefore, a parameterized control strategy is selected. The handlebar control torque is characterized, among which The parameters are the policy network parameters; the system state vector of the unmanned bicycle is selected as follows:
[0016]
[0017] This is used to characterize the current lateral posture and motion state of the unmanned bicycle.
[0018] Based on the above definitions of state and control input, the overall control algorithm flow of the unmanned bicycle lateral balance control method based on adversarial imitation learning proposed in this invention is as follows:
[0019] ST1: Collect or load expert demonstration data of unmanned bicycles under stable lateral balance control conditions. The expert demonstration data includes the state-action sequence of the unmanned bicycle during operation, which is used to characterize the distribution of expert lateral balance control behavior.
[0020] ST2: Initialize the policy network and the discriminant network, wherein the policy network is used to output the handlebar control torque according to the current state, and the discriminant network is used to distinguish whether the input state-action pair comes from expert demonstration data.
[0021] ST3: Repeat the following adversarial imitation learning training process:
[0022] ST3-1: Obtain the initial state of the unmanned bicycle;
[0023] ST3-2: Input the current state into the policy network, and generate continuous handlebar control torque by sampling or averaging according to the parameterized probability distribution output by the policy network.
[0024] ST3-3: Apply the control torque to the unmanned bicycle dynamics simulation environment, update the system state, and store the generated state-action pairs as policy samples;
[0025] ST3-4: Input the strategy sample and the expert demonstration sample into the discrimination network at the same time. The discrimination network outputs the probability value of the corresponding state-action pair belonging to the expert demonstration data.
[0026] ST3-5: Construct an adversarial imitation learning discrimination target based on the output of the discriminant network, and optimize the parameters of the discriminant network;
[0027] ST3-6: Construct a simulated reward signal based on the output of the discriminant network to measure the difference between the control behavior generated by the current policy and the control behavior demonstrated by the expert.
[0028] ST3-7: Based on the imitation reward signal, perform policy evaluation on the policy network to obtain the cumulative reward of the policy in the current round;
[0029] ST3-8: Calculate the cumulative return based on the policy sampling trajectory, and construct the advantage function estimate accordingly to evaluate the current policy update direction;
[0030] ST3-9: Based on the aforementioned advantage function estimation, the policy network parameters are updated using a policy gradient optimization method with pruning constraints, so that the state-action distribution generated by the policy gradually approximates the expert demonstration behavior distribution.
[0031] ST3-10: Alternately execute the optimization process of the policy network and the discriminant network. When the state-action distribution generated by the policy network is statistically consistent with the expert demonstration distribution, the policy network is considered to have converged, and the current training round ends.
[0032] ST4: Repeat step ST3 until the preset number of training rounds or the overall policy convergence condition is met.
[0033] ST5: Solidify the trained policy network and use it as the lateral balance controller for the unmanned bicycle to achieve autonomous lateral balance control; the control method is used to drive the unmanned bicycle actuator to output handlebar control torque to achieve lateral balance control.
[0034] To construct a simulation environment for the lateral balance control of an unmanned bicycle, and for generating expert demonstration data and training adversarial imitation learning algorithms, this invention uses a linearized dynamics model of the unmanned bicycle to describe its lateral motion characteristics. Its technical elements are as follows:
[0035] Let the system state vector be... Control input For an unmanned bicycle system, the nonlinear mechanical model of its lateral tilt angle can be expressed by the following affine equations:
[0036]
[0037] in: The lateral camber angle of the frame ; For the frame camber angle relative to time The first derivative; Let be an unknown nonlinear function of the system, where ; The handlebar control torque input to the system ; This refers to unknown disturbances in the system (including perturbations of internal structural parameters and external disturbances).
[0038] Step 1: Rewrite the classic LPV model in the field of unmanned bicycle research given in the reference "Linearized dynamics equations for the balance and steer of a bicycle: a benchmark and review" into state-space equations.
[0039] Step 2: Discretize the continuous-time state-space model described above to construct a discrete-time simulation model:
[0040]
[0041] in It is a discretized matrix.
[0042] Step 3: Embed the discrete dynamics model into the adversarial imitation learning training environment as the physical basis model for the interaction between the policy network and the environment, and use it to generate state transition data.
[0043] Using the above modeling method, this invention constructs a lateral balance simulation environment for unmanned bicycles suitable for adversarial imitation learning training.
[0044] Furthermore, the LPV mechanical model in step 1 is as follows:
[0045]
[0046] in: , and These are the frame lateral tilt angle and the handlebar swerve angle, respectively. Let be a symmetric invertible matrix representing the vehicle's mass and inertial properties; Here is the equivalent damping matrix of the system related to forward acceleration, steering, etc., where The horizontal linear velocity of the vehicle's center of mass; The equivalent stiffness matrix of the system related to gravitational acceleration, gyroscopic effect, and centrifugal effect. It is the acceleration due to gravity; , which is the control torque vector of the system.
[0047] Take state variables Then the LPV mechanical model can be rewritten as
[0048]
[0049] in: and These are the system's state transition matrix and control input matrix, respectively. To control the input, namely the handlebar torque.
[0050] In ST3-4, the technical elements of constructing the discriminant network are as follows:
[0051] In the lateral balance control method for unmanned bicycles based on adversarial imitation learning, a discriminant network is constructed in steps ST3-4 to distinguish between the state-action pairs generated by the policy network and the expert demonstration data, thereby providing imitation learning signals for the policy network.
[0052] The discriminant network uses the system state vector of the unmanned bicycle. Corresponding control actions As input, its output is the probability value of the state-action pair derived from expert demonstration data, and the mapping relationship is expressed as follows:
[0053]
[0054] in, To determine the network parameters, For the state dimension, For the action dimension.
[0055] In ST3-5, the optimization of the discriminant network involves the following technical elements:
[0056] In step ST3-5, the state-action samples generated by the policy network and the expert demonstration samples are simultaneously input into the discriminant network, and the discriminant network is trained using an adversarial learning mechanism.
[0057] The objective function for the discriminant network is defined as:
[0058]
[0059] in, Indicates the distribution of expert demonstration strategies. This represents the policy distribution generated by the current policy network.
[0060] By minimizing the above discriminant loss function, the discriminant network parameters are... The network is updated to gradually improve its ability to distinguish between expert demonstration behavior and policy generation behavior.
[0061] In ST3-6, the technical elements of constructing the imitation reward based on the discriminative output are as follows:
[0062] In steps ST3-6, to avoid manually designing reward functions that rely on dynamic models, a simulated reward signal is constructed based on the output of the discriminant network to measure the similarity between the current policy behavior and the expert demonstration behavior.
[0063] The imitation reward function is defined as follows:
[0064]
[0065] To improve numerical stability and prevent excessively large reward values from affecting the training process, the simulated reward signal is pruned, and its form is as follows:
[0066]
[0067] in, This is the preset reward cap.
[0068] Furthermore, the imitation reward function adopts the following numerically stable form:
[0069]
[0070] in, To determine the network's output values (logits) for state-action pairs, the above form is... Numerical stability is achieved to avoid numerical instability issues during training.
[0071] In ST3-7, the technical elements of the policy network structure and action generation are as follows:
[0072] In step ST3-7, a policy network is constructed to output the continuous handlebar control torque of the unmanned bicycle.
[0073] The policy network models continuous actions using a parameterized probability distribution, and its policy function is expressed as follows:
[0074]
[0075] in, The average action output by the policy network. It is a diagonal covariance matrix.
[0076] To satisfy physical execution constraints, a nonlinear mapping is applied to the output action of the policy network, and its final control action is expressed as:
[0077]
[0078] in, This represents the upper limit of the handlebar control torque amplitude.
[0079] In ST3-8, the technical elements of cumulative return and advantage function estimation are as follows:
[0080] In steps ST3-8, instead of constructing an explicit state-value function network separately, the cumulative reward at each time step is calculated based on the trajectory sampled during the policy-environment interaction process.
[0081]
[0082] in, As a discount factor, This is a simulated reward signal constructed based on a discriminant network.
[0083] Based on this, the cumulative returns are standardized to construct an estimate of the advantage function:
[0084]
[0085] in, To prevent small positive numbers with unstable values.
[0086] In ST3-9, the technical elements of policy network parameter updates are as follows:
[0087] In step ST3-9, based on the aforementioned advantage function estimation, the policy network parameters are updated using a policy gradient optimization method with pruning constraints. The objective function is defined as:
[0088]
[0089] in, This represents the probability ratio between the old and new strategies:
[0090]
[0091] This is the cutting factor.
[0092] The objective function is optimized using stochastic gradient descent, and the policy network parameters are updated. This limits the magnitude of policy updates and improves the stability of the training process.
[0093] In ST3-10, the technical elements of adversarial optimization and policy convergence determination are as follows:
[0094] In step ST3-10, the optimization process of the discriminant network and the policy network is executed alternately, and the distribution difference between the policy network's generated behavior and the expert's demonstrated behavior is continuously reduced through the adversarial learning mechanism.
[0095] When the policy network can stably maintain the lateral tilt angle and handlebar angle of the unmanned bicycle within the preset threshold range in consecutive training rounds, and the state-action distribution it generates is statistically consistent with the expert demonstration distribution, the policy network is considered to have converged, and the adversarial imitation learning training process ends.
[0096] The beneficial effects of this invention are:
[0097] 1. This invention proposes a balance control method for unmanned bicycles based on adversarial imitation learning. By introducing expert demonstration data and constructing an adversarial learning framework that alternately optimizes the policy network and the discriminant network, the control strategy gradually approximates the expert balance control behavior without the need for explicit design of a reward function. This reduces the dependence on precise dynamic models and improves the anti-interference ability and stable control capability of unmanned bicycles under complex environments and uncertain disturbance conditions.
[0098] 2. This invention is based on an adversarial imitation learning framework. It utilizes a discriminative network to construct an imitation reward signal from the discriminative output of the state-action pairs of an unmanned bicycle. It also combines a policy gradient optimization method with pruning constraints to update the policy network. This enables the learned control policy to continuously optimize the handlebar control torque output around the lateral balance task of the unmanned bicycle. While ensuring the stability of policy updates, it improves the control accuracy, training convergence speed, and control stability during the lateral tilt angle adjustment process. This effectively improves the problems of low sample utilization efficiency and unstable training in the continuous lateral balance control process of unmanned bicycles.
[0099] 3. This invention uses the lateral tilt angle, handlebar angle, and angular velocity of the unmanned bicycle as system state inputs, and the handlebar control torque as continuous control outputs, constructing a unified state-action mapping relationship to achieve data-driven lateral balance control strategy learning for the unmanned bicycle. This enables the controller to output continuous, smooth control commands that satisfy physical constraints under different operating states. It has the advantages of a clear control structure, simple implementation, and ease of engineering deployment and widespread application. Attached Figure Description
[0100] Figure 1 Unmanned bicycle prototype and its structural diagram
[0101] Figure 2 System control block diagram of an unmanned bicycle balance control method based on adversarial imitation learning
[0102] Figure 3 Roll angle curve of unmanned bicycle frame based on adversarial imitation learning
[0103] Figure 4 Handlebar angle curve of an unmanned bicycle based on adversarial imitation learning
[0104] Figure 5 Roll velocity curve of unmanned bicycle frame based on adversarial imitation learning
[0105] Figure 6 Curve of angular velocity of handlebars of an unmanned bicycle based on adversarial imitation learning
[0106] Figure 7Torque curve of unmanned bicycle handlebars based on adversarial imitation learning Detailed Implementation
[0107] To make the above-mentioned objectives, technical solutions, and beneficial effects of the present invention clearer, the present invention will be further described below in conjunction with specific embodiments. It should be understood that the embodiments described below are only for illustrating the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention. Any equivalent substitutions or modifications made by those skilled in the art without departing from the technical concept of the present invention should be included within the scope of protection of the present invention.
[0108] Example:
[0109] This embodiment addresses the challenges of lateral balance control in autonomous bicycles, which are susceptible to model uncertainties and external disturbances during operation. It proposes a balance control method based on adversarial imitation learning. This method aims to achieve lateral tilt angle stability in the autonomous bicycle. By introducing expert demonstration data and constructing an adversarial learning process that alternately optimizes a policy network and a discriminant network, autonomous balance control of the autonomous bicycle in complex environments is achieved. The specific steps are as follows:
[0110] Step 1: Based on the physical parameters of the unmanned bicycle and its components shown in Tables 1 and 2, establish a state-space model obtained by linearizing the classical bicycle dynamics model, which is used to construct a dynamic simulation environment for the lateral balance control of the unmanned bicycle.
[0111]
[0112] The formula is obtained from physical parameters under no-load conditions.
[0113]
[0114] in For vehicle speed, take the matrix. and From the first and third lines we can get:
[0115]
[0116]
[0117]
[0118]
[0119] Specifically, the simplified structural diagram of the research object is as follows: Figure 1 As shown, the physical parameter values corresponding to its LPV model are shown in Tables 1 and 2:
[0120] Table 1 Physical parameters of the unmanned bicycle
[0121]
[0122] Table 2 Physical parameters of unmanned bicycle components
[0123]
[0124] Step 2: State Space Design:
[0125] The state space of an autonomous bicycle is a set of states updated according to the LPV model:
[0126]
[0127] In the formula: The roll angle (roll velocity) of the vehicle body; The handlebar angle (handlebar angular velocity).
[0128] Step 3: Action space design and constraint implementation (corresponding environment and strategy output)
[0129] In this embodiment, the unmanned bicycle achieves lateral balance control by applying a control torque to the handlebar steering mechanism. Therefore, the handlebar control torque is selected as the continuous action output of the controller, and the continuous action space is defined as follows:
[0130]
[0131] in For handlebar control torque, This is the maximum permissible control torque. In this embodiment... This is a preset constant used to satisfy the physical constraints of the actuator.
[0132] To ensure that the policy output actions satisfy the action upper limit constraint, the policy network first outputs unconstrained actions. Then, the hyperbolic tangent function is used for amplitude mapping:
[0133]
[0134] Further trimming is performed on the execution side of the environment:
[0135]
[0136] The above design enables the controller to output continuous, smooth, and executable control commands under different operating conditions.
[0137] Step 4: Observation Normalization and Discriminative Reward Construction
[0138] To improve training stability and eliminate the influence of differences in the dimensions of various states, this embodiment calculates the state mean and standard deviation based on expert demonstration data and constructs an observation normalization operator. Let the original state be:
[0139]
[0140] The mean value was obtained based on expert demonstration data. with standard deviation And by setting a lower limit for the standard deviation to prevent numerical divergence, the normalized state is:
[0141]
[0142] in It is a small positive number. Both the policy network and the discriminant network use... As input state.
[0143] In the generative adversarial imitation learning framework, the discriminant network takes normalized states and actions as inputs and outputs logits:
[0144]
[0145] And from this, the discrimination probability is obtained:
[0146]
[0147] To avoid manually designing complex physical reward functions, this embodiment constructs a simulated reward signal based on the discriminator output. It adopts... Equivalent numerically stable form:
[0148]
[0149] The rewards were also trimmed to improve numerical stability.
[0150]
[0151] in This is the maximum reward.
[0152] Step 5: Generate the adversarial imitation learning training process
[0153] The training process in this embodiment is executed cyclically according to "each round of sampling - discriminator update - policy update", without using an experience replay pool, but instead using a rollout cache for each round of sampling. Each round of training includes the following sub-steps:
[0154] Step 5-1: Policy Sampling and Data Caching
[0155] Running the current policy network in the dynamic environment of an unmanned bicycle, the sampled length is... The trajectory data. For each time step Record normalized state ,action Termination mark And the log probability of the action under the old strategy. For future PPO-Clip updates:
[0156]
[0157] Step 5-2: Determine if a network update is detected
[0158] Sampling from expert demonstration dataset and normalized to Simultaneously, extract from the strategy sampling data The discriminator outputs logits:
[0159]
[0160] The discriminant network parameters are updated using a loss function in the form of binary cross-entropy logarithm (BCEWithLogits):
[0161]
[0162] Step 5-3: Constructing a policy objective based on discriminative rewards
[0163] For each time step, construct the imitation reward according to step 4 and calculate the cumulative reward:
[0164]
[0165] in Discount factor
[0166] Step 5-4: Advantage Estimation
[0167] This embodiment does not additionally train the state-value network, but instead constructs an advantage estimate by standardizing the reward:
[0168]
[0169] And can be used for Amplitude clipping is performed to enhance stability:
[0170]
[0171] Step 6: Policy Update and Convergence Determination
[0172] In this embodiment, the policy network is updated using Proximal Policy Optimization with Pruning Constraints (PPO-Clip). Let the log probability of the new policy be:
[0173]
[0174] Define the probability ratio:
[0175]
[0176] The objective function of PPO-Clip is:
[0177]
[0178] in Cutting factor
[0179] To improve exploration capabilities and prevent premature convergence of the strategy, an entropy regularization term is added to this embodiment:
[0180]
[0181] in This is the entropy coefficient.
[0182] During training, the trajectory data obtained by the interaction between the policy network and the unmanned bicycle dynamics simulation environment is used as the basis, and the simulated reward signal constructed by the output of the discriminant network is combined to iteratively optimize the parameters of the policy network.
[0183] In each training round, the policy network interacts with the simulation environment under the current parameters, collecting state-action trajectories for a fixed number of steps and recording the corresponding action log probabilities to construct the probability ratio between the old and new policies. Subsequently, based on the imitation reward signal obtained from the discriminator network output, the cumulative reward at each time step is calculated, and the cumulative reward is standardized to construct an advantage function estimate. On this basis, the policy network parameters are updated through a policy gradient objective function with pruning constraints, thereby limiting the magnitude of a single update and improving the numerical stability and convergence reliability of the training process.
[0184] During the strategy update process, a manually set fixed reward threshold is not used as the convergence criterion. Instead, the overall performance of the strategy is evaluated. Specifically, if the strategy network can keep the lateral tilt angle and handlebar angle of the unmanned bicycle within a preset safe range under given initial perturbation conditions in consecutive evaluation rounds, and does not trigger the instability termination condition during the control process, it indicates that the current strategy has a stable lateral balance control capability.
[0185] When the above stability criteria are met, or the number of training rounds reaches the preset upper limit, the policy network is determined to have converged, the adversarial imitation learning training process ends, and the trained policy network parameters are fixed and saved for subsequent simulation verification or actual control deployment.
[0186] To further verify the control performance of the proposed adversarial imitation learning-based lateral balance control method for unmanned bicycles, the simulation evaluation process used the dynamic environment of an unmanned bicycle as the test object. Under a given initial roll angle disturbance, Gaussian noise with a mean of 0 and a standard deviation of 0.005 was superimposed on the state vector after each simulation step state update to characterize uncertainties such as modeling error, parameter perturbation, external disturbances, and measurement noise. Under these conditions, the trained policy network was simulated, and the results are shown in the appendix. Figure 3 To be continued Figure 7 As shown:
[0187] See attached document Figure 3 Under the influence of environmental noise and disturbance, the roll angle of the unmanned bicycle frame It can still converge to the vicinity of the equilibrium point in a short time, and maintain fluctuations within a small range during subsequent operation without divergence or continuous deviation, indicating that the control method still has good lateral balance maintenance capability under the condition of uncertain disturbance.
[0188] See attached document Figure 4 Handlebar corner Throughout the control process, the changes remain within a limited range. Although there are some fluctuations due to disturbances, the overall changes are stable and without violent oscillations, indicating that the strategy network can generate control outputs in real time according to changes in system state, thereby effectively suppressing disturbances.
[0189] See attached document Figure 5 With appendix Figure 6 Chassis roll rate With handlebar angular velocity All values fluctuated around zero, without any continuous accumulation or divergence, indicating that the system has good damping characteristics and stability during dynamic response, and can effectively avoid energy amplification caused by disturbances.
[0190] See attached document Figure 7 Control the input handlebar torque Throughout the entire operation, the system remained within the preset constraints without any control saturation or abrupt changes, indicating that the control strategy can achieve continuous and smooth control while satisfying physical execution constraints, demonstrating good engineering feasibility.
[0191] Further analysis reveals that this invention introduces an adversarial imitation learning mechanism, enabling the policy network to learn the state-action mapping relationship in the lateral balance control of an unmanned bicycle under expert demonstration constraints. The imitation reward signal constructed from the output of the discriminant network is then used to continuously optimize the control strategy. Compared to control methods that rely solely on supervised fitting, this method maintains good lateral balance performance and control input smoothness even under given initial disturbances and small-amplitude Gaussian noise interference, indicating that it possesses good disturbance suppression and lateral balance maintenance capabilities.
[0192] In summary, even under simulation conditions incorporating environmental noise and disturbances, the method of this invention can still achieve stable lateral balance control of unmanned bicycles, and exhibits good performance in terms of stability, anti-interference ability, and continuous and abrupt control input, verifying the effectiveness and engineering application value of the method of this invention in complex environments.
[0193] It should be noted that the above convergence determination condition is only one implementation method in this embodiment. In actual applications, the stability determination method and training termination condition can be adjusted accordingly based on the structural parameters, operating conditions and control performance requirements of the unmanned bicycle, and should not be construed as limiting the scope of protection of this invention.
Claims
1. An unmanned bicycle lateral balance control method based on adversarial imitation learning, characterized in that, Includes the following steps: ST1: Collect or load expert demonstration data of unmanned bicycles under stable lateral balance control conditions. The expert demonstration data includes the state-action sequence of the unmanned bicycle during operation, which is used to characterize the distribution of expert lateral balance control behavior. ST2: Initialize the policy network and the discriminant network, wherein the policy network is used to output the handlebar control torque according to the current state, and the discriminant network is used to distinguish whether the input state-action pair comes from expert demonstration data. ST3: Repeat the following adversarial imitation learning training process: ST3-1: Obtain the initial state of the unmanned bicycle; ST3-2: Input the current state into the policy network, and generate continuous handlebar control torque by sampling or averaging according to the parameterized probability distribution output by the policy network. ST3-3: Apply the control torque to the unmanned bicycle dynamics simulation environment, update the system state, and store the generated state-action pairs as policy samples; ST3-4: Input the strategy sample and the expert demonstration sample into the discrimination network at the same time. The discrimination network outputs the probability value of the corresponding state-action pair belonging to the expert demonstration data. ST3-5: Construct an adversarial imitation learning discrimination target based on the output of the discriminant network, and optimize the parameters of the discriminant network; ST3-6: Construct a simulated reward signal based on the output of the discriminant network to measure the difference between the control behavior generated by the current policy and the control behavior demonstrated by the expert. ST3-7: Based on the imitation reward signal, perform policy evaluation on the policy network to obtain the cumulative reward of the policy in the current round; ST3-8: Calculate the cumulative return based on the policy sampling trajectory, and construct the advantage function estimate accordingly to evaluate the current policy update direction; ST3-9: Based on the aforementioned advantage function estimation, the policy network parameters are updated using a policy gradient optimization method with pruning constraints, so that the state-action distribution generated by the policy gradually approximates the expert demonstration behavior distribution. ST3-10: Alternately execute the optimization process of the policy network and the discriminant network. When the state-action distribution generated by the policy network is statistically consistent with the expert demonstration distribution, the policy network is considered to have converged, and the current training round ends. ST4: Repeat step ST3 until the preset number of training rounds or the overall policy convergence condition is met. ST5: Solidify the trained policy network and use it as the lateral balance controller for the unmanned bicycle to achieve autonomous lateral balance control; the control method is used to drive the unmanned bicycle actuator to output handlebar control torque to achieve lateral balance control.
2. The unmanned bicycle balance control method based on adversarial imitation learning according to claim 1, characterized in that: The state of the unmanned bicycle includes a vehicle body roll angle , a handlebar turning angle , a vehicle body roll angle velocity , and a handlebar turning angle velocity .
3. The unmanned bicycle balance control method based on adversarial imitation learning according to claim 1, characterized in that: The control action output by the strategy network is a continuous control torque acting on the steering mechanism of the unmanned bicycle handlebars, and the control torque satisfies a preset amplitude constraint.
4. The lateral balance control method for unmanned bicycles based on adversarial imitation learning according to claim 1, characterized in that: The discriminative network takes state-action pairs as input, outputs the probability value of the state-action pair belonging to the expert demonstration data, and constructs an adversarial imitation learning reward signal based on the probability value to measure the similarity between the control behavior generated by the current policy and the control behavior demonstrated by the expert.
5. The lateral balance control method for unmanned bicycles based on adversarial imitation learning according to claim 1, characterized in that: The policy network is updated using a proximal policy optimization method with pruning constraints. By limiting the update magnitude between the old and new policies, the stability and convergence speed of the training process are improved.