Reinforcement Learning Tracking Control Method for Event-Triggered Mobile Robots
Through the event-triggered reinforcement learning method, the problem of mobile robot trajectory tracking control in a dynamic environment is solved, and optimal trajectory tracking under external perturbation is achieved, reducing the computing burden and improving the control accuracy.
Patent Information
- Application Number
- CN202510486651.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-04-18
AI Technical Summary
Existing control algorithms are difficult to effectively solve the problem of trajectory tracking and control of mobile robots in dynamic environments, especially in the four-McNum wheel system, which is difficult to balance the requirements of control accuracy, calculation burden and real-time response.
The reinforcement learning tracking control method based on event trigger is adopted. By establishing a wheel sliding model and an error augmentation model, a value function and event triggering control strategy are designed, and a neural network is used to solve the approximate optimal control strategy to realize the optimal trajectory tracking control of mobile robots under external perturbations.
It effectively reduces the computing burden, realizes the optimal trajectory tracking control of mobile robots under external disturbances, ensures trajectory tracking accuracy, and reduces the complexity and energy consumption of the system.
Smart Images

Figure CN120010274B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of robot control, and particularly relates to a reinforcement learning tracking control method for a mobile robot based on event triggering. Background Art
[0002] With the rapid development of robot technology, especially in the field of mobile robots, the problem of trajectory tracking control has become a crucial research direction. Trajectory tracking means that the robot accurately follows a given path and minimizes the trajectory error as much as possible, which is of great significance for the robot's navigation, obstacle avoidance, and task execution in a dynamic environment.
[0003] Existing control technologies such as PID control, model predictive control, and adaptive control can effectively improve the control accuracy in some cases, but still face problems such as high computational complexity, dependence on the system model, and high real-time requirements. Especially when dealing with a system like a four-Mecanum-wheel mobile robot with high flexibility and complex dynamic characteristics, it is difficult for existing methods to balance the requirements of control accuracy, computational burden, and real-time response.
[0004] Event-triggered control is a dynamic event-driven control strategy. By setting trigger conditions, control updates are only performed when the system state changes exceed a certain preset threshold, thereby effectively reducing the computational burden and communication cost, and adapting to application scenarios with high real-time requirements and limited energy.
[0005] Therefore, a hybrid control strategy combining event-triggered control and reinforcement learning can effectively overcome the limitations of traditional methods. By reducing unnecessary computational and communication overheads, it can simultaneously achieve efficient optimization of control strategies, thereby improving the trajectory tracking performance while reducing the system complexity and energy consumption. Summary of the Invention
[0006] Based on the above problems, the present invention proposes a reinforcement learning tracking control method for a mobile robot based on event triggering, which solves the problem of difficult solution of the Hamilton-Jacobi-Bellman (HJB) equation by traditional control algorithms, effectively reduces the computational burden, and realizes the optimal trajectory tracking control of a mobile robot under external disturbances. The technical solution is as follows:
[0007] A reinforcement learning tracking control method for a mobile robot based on event triggering includes the following steps:
[0008] S1. Establish a wheel slip model of the mobile robot, and based on the wheel slip model, establish a forward kinematic model, an inverse kinematic model, and a dynamic model of the mobile robot under the influence of slip. Substitute the forward kinematic model and the inverse kinematic model under the influence of slip into its dynamic model to obtain a trajectory control model of the mobile robot;
[0009] S2. Set the desired trajectory of the mobile robot, introduce the tracking error and calculate the tracking error dynamics simultaneously, and combine the error tracking dynamics with the desired trajectory to obtain the error augmented model of the mobile robot;
[0010] S3. Design the value function, obtain the optimal control strategy to be solved, and add the event-triggered control scheme;
[0011] S4. Use neural network to solve the approximate optimal control strategy, and design the optimal trajectory tracking control algorithm for the mobile robot based on event-triggering and neural network;
[0012] S5. Analyze the uniformly ultimately bounded stability of the mobile robot, and complete the trajectory tracking control of the mobile robot based on event-triggering and reinforcement learning.
[0013] Preferably, substitute the forward kinematic model and the inverse kinematic model under the influence of the mobile robot's slip into its dynamic model to obtain the mobile robot trajectory control model, and the formula is as follows:
[0014] ;
[0015] ;
[0016] ;
[0017] ;
[0018] ;
[0019] ;
[0020] The mobile robot trajectory control model is:
[0021] ;
[0022] represents the position and azimuth angle of the mobile robot, represents the derivative of the three-dimensional column vector of, is the wheel radius of the mobile robot, is the slip factor, , is the equivalent moment of inertia of the wheel, is the viscous friction coefficient between the wheel and the ground, is the control input of the mobile robot trajectory control model, and are the state matrices with the azimuth angle as the parameter; is the three-dimensional column vector The derivative of
[0023] Preferably, the expected trajectory formula of the mobile robot in step S2 is:
[0024] ;
[0025] Where is the expected trajectory, is the expected position of the mobile robot, is the expected velocity of the mobile robot, and the superscript represents the transpose of the matrix;
[0026] The trajectory tracking error The calculation formula is:
[0027] ;
[0028] Where is the trajectory tracking error, is the current trajectory;
[0029] The dynamics of the tracking error The expression formula is:
[0030] ;
[0031] Define the error augmented model state as Then the expression formula of the error augmented model is:
[0032] ;
[0033] ;
[0034] ;
[0035] ;
[0036] ;
[0037] ;
[0038] ;
[0039] ;
[0040] ;
[0041] is the external disturbance, is the expected trajectory The derivative of is the control input of the error augmented model, is the steady-state control input.
[0042] Preferably, in step S3:
[0043] S31. Design the value function to obtain the Hamiltonian function;
[0044] S32. Obtain the HJB equation to be solved and the optimal control strategy according to the Hamiltonian function;
[0045] S33. Design the event-triggered control scheme to obtain the event-triggered control strategy.
[0046] Preferably, the expression formula of the value function in step S31 is:
[0047] ;
[0048] where represents the value function, is the discount factor, is the attenuation exponent, is a constant, s-t is the time interval, is the control input of the error augmented model, , is the state of the error augmented model, is the disturbance factor, ;
[0049] and are both constant matrices, and the formulas are as follows,
[0050] ;
[0051] ;
[0052] ;
[0053] is the identity matrix;
[0054] Derive the value function with respect to the error augmented model to obtain the approximate Hamiltonian function expression formula as:
[0055] ;
[0056] where represents the Hamiltonian function, represents the derivative of the value function with respect to the state of the error augmented model.
[0057] Preferably, the HJB equation in step S32 is:
[0058] ;
[0059] Among them, represents the optimal Hamiltonian function, represents the optimal control strategy, represents the optimal value function;
[0060] Optimal control strategy is:
[0061] ;
[0062] Among them, is the optimal control strategy, is the control strategy of the admissible set, represents the derivative of the optimal value function with respect to the state of the error augmented model, is the state of the error augmented model, represents a four-dimensional positive definite matrix.
[0063] Preferably, the event-triggered control scheme described in step S33 is:
[0064] ;
[0065] Among them represents the triggering error, represents the state of the error augmented model at the current moment, represents the th state of the error augmented model at the triggering moment, and are positive definite symmetric matrices;
[0066] The event-triggered control strategy is:
[0067] ;
[0068] represents the control input at the moment of represents the control input at the moment of
[0069] Preferably, step S4 includes the following sub-steps:
[0070] S41. Approximate the optimal value function using a single-layer neural network, and calculate the optimal control strategy according to the analytical relationship of the HJB equation;
[0071] S42. Define the Bellman error of the HJB function, and design the update strategy of the weights of the neural network according to the gradient descent method.
[0072] Preferably, the following formula is included in step S41:
[0073] ;
[0074] Among them, represents the optimal value function, is the transpose of the weights of the neural network, is the activation function of the neural network, is the approximation error of the neural network;
[0075] The expression formula of the optimal control strategy is:
[0076] ;
[0077] Among them, represents the optimal control strategy, is a four-dimensional constant matrix, is the transpose of the system state matrix ;
[0078] According to the Weierstrass high-order approximation theory, approximate the optimal value function and the optimal policy function to obtain the estimated values of the optimal value function and the optimal policy function as follows:
[0079] ;
[0080] Among them, and respectively represent the estimated values of the weights of the neural network.
[0081] Preferably, step S42 includes the following formula:
[0082] Substitute the estimated values of the optimal value function and the optimal policy function into the Hamiltonian equation to obtain the approximate Hamiltonian equation as:
[0083] ;
[0084] Among them, is the approximate Hamiltonian function;
[0085] Define the calculation formula of the Bellman error of the Hamiltonian function as:
[0086] ;
[0087] Among them, is the approximate Hamiltonian function, is the optimal Hamiltonian function;
[0088] Bellman error The expression formula is:
[0089] ;
[0090] The update strategy of the weight is:
[0091] ;
[0092] Among them, represents the derivative of the neural network weight , represents the derivative of the neural network weight , and represents the parameter matrix, , , and are positive constants, is the constant forgetting factor.
[0093] Preferably, analyzing the uniform ultimate boundedness of the mobile robot system in step S5 includes the following formula:
[0094] Select the Lyapunov function of the entire mobile robot trajectory tracking control system as:
[0095] ;
[0096] Among them, is the optimal value function, is the inverse matrix of the dynamic gain matrix , and are positive constants, and are the weight estimation errors, , , when , its derivative is obtained as:
[0097] ;
[0098] Among them, represents the derivative of the Lyapunov function , is the derivative of the inverse matrix of the dynamic gain matrix , represents the derivative of the optimal value function, and They are the derivatives of the weight estimation errors of the Critic network and the Actor respectively. According to the weight update strategy, the weight estimation error of the neural network and are differentiated to obtain:
[0099] ;
[0100] Among them, represents the normalization parameter matrix, , represents the parameter matrix, , is a scalar related to ;
[0101] and are transition matrices, and the formulas are as follows:
[0102] ;
[0103] ;
[0104] ;
[0105] Substitute the above formula into the derivative of the Lyapunov function to obtain:
[0106] ;
[0107] ;
[0108] ;
[0109] , ;
[0110] ;
[0111] When the following conditions are met, the derivative of the Lyapunov function is:
[0112] ;
[0113] When , if the above formula is satisfied, then there is:
[0114] ;
[0115] Among them is the optimal value function at the moment of , , be the optimal value function at a certain moment, the weight estimation of the Critic network error, the weight estimation of the Actor network error and the dynamic gain matrix the inverse matrix of; = ; , , , , , are all positive constants, respectively satisfying , , , , , ; , , , are all positive definite symmetric matrices of appropriate dimensions, and are respectively the minimum eigenvalue and the maximum eigenvalue of the matrix, , are all given constants.
[0116] Compared with the prior art, the beneficial effects of this application are as follows:
[0117] For the trajectory tracking control problem of a four-Mecanum wheel mobile robot, this application uses the Actor-Critic synchronous learning algorithm to solve the optimal control strategy, uses the interaction data for policy iteration update, adjusts the controller output according to the disturbance in real time, and at the same time introduces an event-triggering mechanism, and the controller only updates the control signal when the triggering threshold is met; based on the optimal control theory, the uniform ultimate bounded stability of the four-Mecanum wheel mobile robot system is analyzed. This application solves the trajectory tracking problem of the mobile robot in the presence of external disturbances such as sliding, and ensures the trajectory tracking accuracy while effectively reducing the computational burden. BRIEF DESCRIPTION OF THE DRAWINGS
[0118] Figure 1 is a schematic diagram of a four-Mecanum wheel omnidirectional mobile robot and a coordinate system;
[0119] Figure 2 is a trajectory tracking error graph using the control method of the present invention;
[0120] Figure 3 is a trajectory tracking error graph using the time-triggered reinforcement learning control method;
[0121] Figure 4 The expected and actual movement trajectories of a mobile robot in the direction when using the control method of the present invention;
[0122] Figure 5 The expected and actual movement trajectories of a mobile robot in the direction when using the control method of the present invention;
[0123] Figure 6 The expected and actual direction angles of a mobile robot during the movement process when using the control method of the present invention;
[0124] Figure 7 The convergence curve graph of the weights of the Actor neural network when using the control method of the present invention ;
[0125] Figure 8 The convergence curve graph of the weights of the Critic neural network when using the control method of the present invention ;
[0126] Figure 9 The triggering moment and triggering time interval graph when using the control method of the present invention within 120 - 125 seconds;
[0127] Figure 10 The triggering moment and triggering time interval graph when using the time - triggered reinforcement learning control method within 120 - 125 seconds;
[0128] Figure 11 The triggering times graph when using the control method of the present invention and the time - triggered reinforcement learning control method. Detailed implementation mode
[0129] The technical solution of the present application will be described in detail below through specific embodiments and the accompanying drawings. It should be understood that the specific features in the embodiments of the present application are detailed descriptions of the technical solution of the present application, rather than limitations on the technical solution of the present application, and the specific technical features can be combined with each other.
[0130] A reinforcement learning tracking control method for an event - triggered mobile robot includes the following steps:
[0131] S1. Establish a wheel slip model of the mobile robot, establish a forward kinematic model, an inverse kinematic model, and a dynamic model of the mobile robot under the influence of slip according to the wheel slip model, and substitute the forward kinematic model and the inverse kinematic model under the influence of slip into its dynamic model to obtain a trajectory control model of the mobile robot;
[0132] By controlling the speeds of the four wheels of the mobile robot, omnidirectional movement of the robot is achieved. The wheel slip model is as follows:
[0133] ;
[0134] The forward kinematic model, inverse kinematic model, and dynamic model under sliding are as follows:
[0135] ;
[0136] Among them, is the angular velocity of rotation of the th wheel, is the linear velocity of the th wheel, , represents the position and azimuth angle of the mobile robot, represents the derivative of the three-dimensional column vector , is the wheel radius of the mobile robot, is the sliding factor, , represents the angular velocities of rotation of the four wheels, is the derivative of the four-dimensional column vector , is the equivalent moment of inertia of the wheel, is the viscous friction coefficient between the wheel and the ground, is the control input of the trajectory control model of the mobile robot, and are state matrices with the azimuth angle as a parameter;
[0137] ;
[0138] ;
[0139] is half of the width of the mobile robot, is half of the length of the mobile robot.
[0140] Substituting the forward and inverse kinematic models of the mobile robot into its dynamic model, we get:
[0141] ;
[0142] ;
[0143] ;
[0144] ;
[0145] ;
[0146] ;
[0147] The trajectory control model of the mobile robot is:
[0148] ;
[0149] where, is the linear velocity of the mobile robot's movement, is the angular velocity of self-rotation, is the derivative of the three-dimensional column vector .
[0150] S2. Set the desired trajectory of the mobile robot, introduce the tracking error and calculate the tracking error dynamics at the same time, and combine the error tracking dynamics with the desired trajectory to obtain the error augmented model of the mobile robot;
[0151] The formula for the desired trajectory of the mobile robot is:
[0152] ;
[0153] is the desired trajectory, is the desired position of the mobile robot, is the desired velocity of the mobile robot, and the superscript represents the transpose of the matrix;
[0154] The trajectory tracking error The calculation formula is:
[0155] ;
[0156] where, is the trajectory tracking error, is the current trajectory;
[0157] The dynamics of the said tracking error The expression formula is:
[0158] .
[0159] Define the state of the error tracking model , then the expression formula of the said error augmented model is:
[0160] ;
[0161] ;
[0162] ;
[0163] ;
[0164] ;
[0165] ;
[0166] ;
[0167] is the external disturbance, is the desired trajectory derivative of, is the control input of the error augmented model, is the steady-state control input.
[0168] S3. Design the value function, obtain the optimal control strategy, and add the event-triggered control scheme;
[0169] S31. Design the value function to obtain the Hamiltonian function;
[0170] ;
[0171] where, represents the value function, is the discount factor, is the decay exponent, is a constant, s-t is the time interval, is the control input of the error augmented model, , is the state of the error augmented model, is the disturbance factor, ;
[0172] and are both constant matrices, and the formulas are as follows:
[0173] ;
[0174] ;
[0175] ;
[0176] is the identity matrix;
[0177] Derive the value function with respect to the error augmented model to obtain the expression formula of the approximate Hamiltonian function as:
[0178] ;
[0179] where, represents the Hamiltonian function, represents the derivative of the value function with respect to the state of the error augmented model.
[0180] S32. Obtain the HJB equation to be solved and the optimal control strategy according to the Hamiltonian function;
[0181] The HJB equation is:
[0182] ;
[0183] where, represents the optimal Hamiltonian function, represents the optimal control strategy, represents the optimal value function;
[0184] The optimal control strategy is:
[0185] ;
[0186] where, is the optimal control strategy, is the control strategy of the admissible set, represents the derivative of the optimal value function with respect to the state of the error augmented model, is the state of the error augmented model, represents a four-dimensional positive definite matrix.
[0187] S33. Design an event-triggered control scheme to obtain the event-triggered control strategy;
[0188] The event-triggered condition is:
[0189] ;
[0190] where represents the triggering error, represents the state of the error augmented model at the current time, represents the th state of the error augmented model at the triggering time, and are positive definite symmetric matrices;
[0191] The event-triggered control strategy is:
[0192] ;
[0193] represents the control input at time represents the control input at time
[0194] S4. Solve the approximate optimal control law and design an optimal trajectory tracking control algorithm for a mobile robot based on event-triggering and neural networks.
[0195] S41. Approximate the optimal value function using a single-layer neural network, and calculate the optimal control strategy according to the analytical relationship of the equation;
[0196] ;
[0197] where, represents the optimal value function, is the transpose of the weights of the neural network, is the activation function of the neural network, is the approximation error of the neural network;
[0198] The expression formula of the optimal control strategy is:
[0199] ;
[0200] where, represents the optimal control strategy, is a four-dimensional constant matrix, is the transpose of the error-augmented model state matrix ;
[0201] According to the Weierstrass high-order approximation theory, approximate the optimal value function and the optimal policy function to obtain the estimated values of the optimal value function and the optimal policy function as follows:
[0202] ;
[0203] where, and respectively represent the estimated values of the weights of the neural network.
[0204] S42. Define the Bellman error of the function, and design the weight update law of the neural network according to the gradient descent method.
[0205] Substitute the estimated values of the optimal value function and the optimal policy function into the Hamiltonian equation to obtain the approximate Hamiltonian equation as:
[0206] ;
[0207] where, is the approximate Hamiltonian function.
[0208] The calculation formula for the Bellman error of the Hamiltonian function is defined as:
[0209] ;
[0210] where, is the approximate Hamiltonian function, is the optimal Hamiltonian function;
[0211] The Bellman error has the following expression formula:
[0212] ;
[0213] The update strategy for the weights is:
[0214] ;
[0215] where, represents the derivative of the neural network weight , represents the derivative of the neural network weight , and represent parameter matrices, , , and are positive constants, is the constant forgetting factor.
[0216] S5. Analyze the uniformly ultimately bounded stability of the mobile robot and complete the trajectory tracking control of the mobile robot based on event-triggering and reinforcement learning.
[0217] Select the Lyapunov function of the entire mobile robot trajectory tracking control system as:
[0218] ;
[0219] where, is the optimal value function, is the inverse matrix of the dynamic gain matrix , and are positive constants, and are the weight estimation errors, , , when , its derivative is obtained as:
[0220] ;
[0221] where, Denote the Lyapunov function 's derivative, is the dynamic gain matrix 's inverse matrix's derivative, Denote the derivative of the optimal value function, and are the derivatives of the weight estimation errors of the Critic network and the Actor respectively. Update the optimal control strategy according to the weights, and take the derivatives of the weight estimation errors and of the neural network to obtain:
[0222] ;
[0223] where, denotes the normalized parameter matrix, , denotes the parameter matrix, , is a scalar related to , and are transition matrices, , , , and are positive constants.
[0224] Substitute the above formula into the derivative of the Lyapunov function to obtain:
[0225] ;
[0226] ;
[0227] ;
[0228] , ;
[0229] ;
[0230] When the following conditions are satisfied, the derivative of the Lyapunov function is:
[0231] ;
[0232] When , if the above formula is satisfied:
[0233] ;
[0234] where is the optimal value function at time , , is the optimal value function at time , the weight estimation of the Critic network error of and the weight estimation error of the Actor network and the inverse matrix of the dynamic gain matrix = ; , , , , , are all positive constants, satisfying respectively , , , , , ; , , , are all positive definite symmetric matrices of appropriate dimensions, and are respectively the minimum eigenvalue and the maximum eigenvalue of the matrix, are all constants.
[0235] In summary, according to the standard Lyapunov stability theorem, the state of the error tracking model and the estimation errors of the weights and are all uniformly ultimately bounded.
[0236] Given the desired trajectory for the mobile robot to track, and the initial position .
[0237] The simulation parameters are selected as shown in Table 1:
[0238] Table 1 Simulation Parameters
[0239] .
[0240] Figure 2 and Figure 3 are respectively the comparison diagrams of the tracking errors of the control method given by the present invention and the time-triggered reinforcement learning control method. The blue line e1 is the tracking error in the x direction, the red line e2 is the tracking error in the y direction, and the black line e3 is the tracking error of the direction angle of Figure 4 , Figure 5 ,Figure 6 The tracking curves of the robot in the x-direction, y-direction, and orientation angle respectively. The red line is the desired trajectory, and the blue line is the response result corresponding to the control method proposed by the present invention.
[0241] Figure 7 , Figure 8 are respectively the curves of the weights of the Actor neural network and the Critic neural network proposed by the present invention changing with time. The two figures reflect the training process of the neural network. As the training iterates, the weight values converge, and the convergence values of the weights of the two neural networks are equal.
[0242] Figure 9 and Figure 10 are respectively the trigger time interval diagrams of the control method proposed by the present invention and the time-triggered reinforcement learning control method. The x-coordinate is the trigger moment within 120 - 125 seconds, and the y-coordinate is the time interval between the current trigger and the previous trigger. Figure 11 is the trigger count diagram of the control method proposed by the present invention and the time-triggered reinforcement learning control method. Within 0 - 250 seconds, the time-triggered reinforcement learning control method was triggered 141,979 times, while the control method proposed by the present invention was triggered 81,646 times, and the trigger count decreased by 42.49%. Combining Figures 2 - 11 it can be seen that the control method proposed by the present invention can effectively reduce the computational burden while ensuring the tracking speed and tracking accuracy.
[0243] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present application, several improvements and deformations can still be made, and these improvements and deformations should also be regarded as the protection scope of the present application.
Claims
1. A reinforcement learning tracking control method for a mobile robot based on event triggering, characterized in that: The following steps are involved: S1. Establish a wheel sliding model of the mobile robot, establish a forward kinematics model, an inverse kinematics model and a dynamic model of the mobile robot under the influence of sliding according to the wheel sliding model, substitute the forward kinematics model and the inverse kinematics model under the influence of sliding into its dynamic model to obtain a trajectory control model of the mobile robot; S2. Set the desired trajectory of the mobile robot, introduce the tracking error and calculate the tracking error dynamics, combine the tracking error dynamics with the desired trajectory, and obtain the error augmented model of the mobile robot; S3. Design the value function, obtain the optimal control strategy to be solved, and add the event-triggered control scheme; S31. Design a value function and obtain a Hamiltonian function; The value function expression formula is: Where V(z(t)) represents the value function, γ is the discount factor, and c -γ(s-t) is the decay exponent, c is a constant, st is the time interval, u(z) is the control input of the error augmentation model, u(z) = τ-τ r , z is the state of the error augmentation model, λ M (z) is the disturbance factor, λ M (z) = || d(q, v q )||·||g(x q )||; and R are both constant matrices, and the formula is as follows: Q = 6I; R = 4I; I is the identity matrix; Derivative the value function with respect to the error augmentation model to obtain the approximate Hamiltonian function, which is expressed as: Where H(·) represents the Hamiltonian function, represents the derivative of the value function with respect to the state of the error augmented model; S32. According to the Hamiltonian function, the HJB equation and the optimal control strategy to be solved are obtained; S33. Design an event-triggered control scheme to obtain an event-triggered control strategy; S4. Use neural networks to solve approximate optimal control strategies and design an optimal trajectory tracking control algorithm for mobile robots based on event triggering and neural networks; S5. Analyze the consistent ultimate bounded stability of the mobile robot and complete the event-triggered reinforcement learning trajectory tracking control of the mobile robot.
2. The event-triggered mobile robot reinforcement learning tracking control method according to claim 1 is characterized in that: Substituting the forward kinematics model and inverse kinematics model of the mobile robot under the influence of sliding into its dynamic model, the trajectory control model of the mobile robot is obtained, and the formula is as follows: The expression formula of the mobile robot trajectory control model is: q=(xy θ) T represents the position and azimuth of the mobile robot, represents the derivative of the three-dimensional column vector q, r is the wheel radius of the mobile robot, ρ is the sliding factor, 0<ρ<1, I0 is the equivalent moment of inertia of the wheel, η0 is the viscous friction coefficient between the wheel and the ground, τ=[τ1 τ2 τ3 τ4] T is the control input of the trajectory control model of the mobile robot, H(θ) and P(θ) are the state matrices with the azimuth angle θ as the parameter; is a three-dimensional column vector v q The derivative of .
3. The event-triggered mobile robot reinforcement learning tracking control method according to claim 1, characterized in that: The formula for the expected trajectory of the mobile robot in step S2 is: Among them, x r is the expected trajectory, q r is the desired position of the mobile robot, v qr is the expected velocity of the mobile robot, and the superscript T represents the transpose of the matrix; The calculation formula of trajectory tracking error e is: e=x q -x r ; Where e is the trajectory tracking error, x q is the actual trajectory; Tracking Error Dynamics The expression formula is: Define the state of the error augmentation model as Then the expression formula of the error augmentation model is: ξ=G(z)d(q,v q ); ξ is the external disturbance, is the expected trajectory x r The derivative of , u(z) is the control input of the error augmentation model, τ r is the steady-state control input.
4. The event-triggered mobile robot reinforcement learning tracking control method according to claim 1, characterized in that: The HJB equation in step S32 is: Among them, H * (·) represents the optimal Hamiltonian function, u * (z) represents the optimal control strategy, V * (z) represents the optimal value function; Optimal control strategy u * (z) is: Among them, u * (z) is the optimal control strategy, π(Ω) is the admissible set of control strategies u(z), represents the derivative of the optimal value function with respect to the state of the error augmented model, z is the state of the error augmented model, and R represents a four-dimensional positive definite matrix; The event triggering control scheme in step S33 is: where d j (t) = z(t)-z(t j ) represents the trigger error, z(t) represents the state of the error augmentation model at the current moment, z(t j ) represents the state of the error augmentation model at the jth triggering moment, Ω1 and Ω2 are positive definite symmetric matrices; The event trigger control strategy is: u(z(t))=u(z(t j )),t∈t j ,t j+1 ); u(z(t)) represents the control input at time t, u(z(t j )) means t j Control input at any time.
5. The event-triggered mobile robot reinforcement learning tracking control method according to claim 1, characterized in that: Step S4 includes the following sub-steps: S41. Use a single-layer neural network to approximate the optimal value function and calculate the optimal control strategy based on the analytical relationship of the HJB equation; S42. Define the Bellman error of the HJB function and design the update strategy of the neural network weights based on the gradient descent method.
6. The method for tracking and controlling a mobile robot based on event triggering according to claim 5, characterized in that: Step S41 includes the following formula: V * (from)=In T φ(z)+ε v (With); Among them, V * (z) represents the optimal value function, W T is the transpose of the weights of the neural network, φ(z) is the activation function of the neural network, ε v (z) is the approximation error of the neural network; The expression formula of the optimal control strategy is: Among them, u * (z) represents the optimal control strategy, R is a four-dimensional constant matrix, G T (z) is the transpose of the system state matrix G(z); According to Weierstrass high-order approximation theory, the optimal value function V * (z) and the optimal policy function u * (z) is approximated to obtain the optimal value function V * (z) and the optimal policy function u * The estimated values of (z) are: in, and They represent the estimated weights of the neural network respectively.
7. The method for tracking and controlling a mobile robot based on event triggering according to claim 5, characterized in that: Step S42 includes the following formula: The optimal value function V * (z) and the optimal policy function u * Substituting the estimated value of (z) into the Hamiltonian equation, we obtain the approximate Hamiltonian equation for: in, is the approximate Hamiltonian function; The calculation formula for the Bellman error that defines the Hamiltonian function is: in, is the approximate Hamiltonian function, H * is the optimal Hamiltonian function; Bellman Error The expression formula is: The weight update strategy is: in, Represents the neural network weights The derivative of Represents the neural network weights The derivatives of σ(t) and m σ represents the parameter matrix, η c and v are positive constants, and β∈(0,1) is a constant forgetting factor.
Citation Information
Patent Citations
Robot trajectory tracking optimal control method based on an event trigger mechanism
CN113093548A