Reinforced learning tracking control method of mobile robot based on event triggering
By adopting event-triggered reinforcement learning tracking control method in the four-mcnum wheel mobile robot, the problem of difficulty in balancing control accuracy and calculation burden in the prior art is solved, and the optimal trajectory tracking control under external disturbance conditions is achieved.
Patent Information
- Application Number
- CN202510486651.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-18
AI Technical Summary
Existing control technology is difficult to balance control accuracy, calculation burden and real-time response requirements in four-McNum wheel mobile robots, especially in the face of external disturbances.
A reinforcement learning tracking control method based on event trigger is adopted, and a value function and optimal control strategy are designed by establishing a wheel sliding model and an error augmentation model, and a neural network is used to solve the approximate optimal control strategy, and an event trigger control mechanism is introduced.
It effectively reduces the computing burden, realizes the optimal trajectory tracking control of mobile robots under external disturbances, improves trajectory tracking performance, and ensures the stability of the system.
Smart Images

Figure CN120010274A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of robot control, and specifically relates to a reinforcement learning tracking control method of a mobile robot based on event triggering. Background Art
[0002] With the rapid development of robotics technology, especially in the field of mobile robots, trajectory tracking control has become a crucial research direction. Trajectory tracking refers to the robot accurately following a given path and minimizing trajectory errors, which is of great significance for the robot's navigation, obstacle avoidance and task execution in a dynamic environment.
[0003] Although existing control technologies such as PID control, model predictive control, and adaptive control can effectively improve control accuracy in some cases, they still face problems such as computational complexity, dependence on system models, and high real-time requirements. Especially when facing a system with high flexibility and complex dynamic characteristics such as a four-Mecanum wheel mobile robot, existing methods find it difficult to balance the requirements of control accuracy, computational burden, and real-time response.
[0004] Event-triggered control is a dynamic event-driven control strategy that sets trigger conditions and only performs control updates when the system state changes exceed a preset threshold, thereby effectively reducing the computational burden and communication costs, and adapting to application scenarios with high real-time requirements and limited energy.
[0005] Therefore, the hybrid control strategy that combines event-triggered control and reinforcement learning can effectively overcome the limitations of traditional methods by reducing unnecessary computational and communication overheads while achieving efficient control strategy optimization, thereby improving trajectory tracking performance while reducing system complexity and energy consumption. Summary of the invention
[0006] Based on the above problems, the present invention proposes a reinforcement learning tracking control method for a mobile robot based on event triggering, which solves the problem that the traditional control algorithm is difficult to solve the Hamilton-Jacobi-Bellman (HJB) equation, effectively reduces the computational burden, and realizes the optimal trajectory tracking control of the mobile robot under external disturbance. Its technical solution is: A reinforcement learning tracking control method for a mobile robot based on event triggering comprises the following steps: S1. Establish a wheel sliding model of the mobile robot, establish a forward kinematics model, an inverse kinematics model and a dynamics model of the mobile robot under the influence of sliding according to the wheel sliding model, substitute the forward kinematics model and the inverse kinematics model under the influence of sliding into its dynamics model to obtain a trajectory control model of the mobile robot; S2. Set the desired trajectory of the mobile robot, introduce the tracking error and calculate the tracking error dynamics, combine the error tracking dynamics with the desired trajectory, and obtain the error augmented model of the mobile robot; S3. Design the value function, obtain the optimal control strategy to be solved, and add the event-triggered control scheme; S4. Use neural networks to solve approximate optimal control strategies and design an optimal trajectory tracking control algorithm for mobile robots based on event triggering and neural networks; S5. Analyze the consistent ultimate bounded stability of the mobile robot and complete the trajectory tracking control of the mobile robot based on event triggering and reinforcement learning.
[0007] Preferably, the forward kinematics model and the inverse kinematics model of the mobile robot under the influence of sliding are substituted into its dynamic model to obtain the trajectory control model of the mobile robot, and the formula is as follows: ; ; ; ; ; ; The trajectory control model of the mobile robot is: ; represents the position and azimuth of the mobile robot, Represents a three-dimensional column vector The derivative of is the wheel radius of the mobile robot, is the sliding factor, , is the equivalent moment of inertia of the wheel, is the viscous friction coefficient between the wheel and the ground, is the control input of the mobile robot trajectory control model, and Azimuth is the state matrix of the parameters; is a three-dimensional column vector The derivative of .
[0008] Preferably, the formula for the expected trajectory of the mobile robot in step S2 is: ; in, is the expected trajectory, is the desired position of the mobile robot, is the expected speed of the mobile robot, Represents the transpose of a matrix; Trajectory tracking error The calculation formula is: ; in, is the trajectory tracking error, is the current trajectory; Tracking Error Dynamics The expression formula is: ; Define the error augmentation model state as , then the expression formula of the error augmentation model is: ; ; ; ; ; ; ; ; ; is the external disturbance, The expected trajectory The derivative of is the control input of the error augmentation model, is the steady-state control input.
[0009] Preferably, in step S3: S31. Design a value function and obtain a Hamiltonian function; S32. According to the Hamiltonian function, the HJB equation to be solved and the optimal control strategy are obtained; S33. Design an event-triggered control scheme and obtain an event-triggered control strategy.
[0010] Preferably, the value function in step S31 is expressed as: ; in, represents the value function, is the discount factor, is the decay exponent, is a constant, st is the time interval, is the control input of the error augmentation model, , is the state of the error augmentation model, is the disturbance factor, ; and are all constant matrices, and the formula is as follows: ; ; ; is the identity matrix; Derivative the value function with respect to the error augmentation model, and the approximate Hamiltonian function expression formula is obtained as follows: ; in, represents the Hamiltonian function, Represents the derivative of the value function with respect to the state of the error augmented model.
[0011] Preferably, the HJB equation in step S32 is: ; in, represents the optimal Hamiltonian function, represents the optimal control strategy, represents the optimal value function; Optimal control strategy for: ; in, is the optimal control strategy, For control strategy The admissible set of represents the derivative of the optimal value function with respect to the state of the error augmented model, is the state of the error augmentation model, represents a 4D positive definite matrix.
[0012] Preferably, the event triggering control scheme in step S33 is: ; in Indicates the trigger error, represents the state of the error augmentation model at the current moment, Indicates The state of the error augmented model at the time of the next trigger, and is a positive definite symmetric matrix; The event trigger control strategy is: ; express The control input at the moment, express Control input at any time.
[0013] Preferably, step S4 comprises the following sub-steps: S41. Use a single-layer neural network to approximate the optimal value function and calculate the optimal control strategy based on the analytical relationship of the HJB equation; S42 defines the Bellman error of the HJB function and designs the update strategy of the neural network weights based on the gradient descent method.
[0014] Preferably, step S41 includes the following formula: ; in, represents the optimal value function, is the transpose of the neural network weights, is the activation function of the neural network, is the approximation error of the neural network; The expression formula of the optimal control strategy is: ; in, represents the optimal control strategy, is a four-dimensional constant matrix, is the system state matrix The transpose of According to Weierstrass's high-order approximation theory, the optimal value function and the optimal policy function Approximate and obtain the optimal value function and the optimal policy function The estimated values are: ; in, and They represent the estimated weights of the neural network respectively.
[0015] Preferably, step S42 includes the following formula: The optimal value function and the optimal policy function Substituting the estimated value of into the Hamiltonian equation, we get the approximate Hamiltonian equation for: ; in, is the approximate Hamiltonian function; The calculation formula for the Bellman error that defines the Hamiltonian function is: ; in, is the approximate Hamiltonian function, is the optimal Hamiltonian function; Bellman Error The expression formula is: ; The weight update strategy is: ; in, Represents the neural network weights The derivative of Represents the neural network weights The derivative of and represents the parameter matrix, , , and is a positive constant, is the constant forgetting factor.
[0016] Preferably, the analysis of the consistent final boundedness of the mobile robot system in step S5 includes the following formula: Select the Lyapunov function of the entire mobile robot trajectory tracking control system for: ; in, is the optimal value function, is the dynamic gain matrix The inverse matrix of and is a positive constant, and is the weight estimation error, , ,when When , we can derive it to get: ; in, represents the Lyapunov function The derivative of is the dynamic gain matrix The derivative of the inverse matrix of represents the derivative of the optimal value function, and They are the derivatives of the weight estimation errors of the Critic network and the Actor respectively. According to the weight update strategy, the weight estimation error of the neural network is and The derivative is: ; in, represents the normalized parameter matrix, , represents the parameter matrix, , yes A scalar quantity related to ; and is the transition matrix, and the formula is as follows: ; ; ; Substitute the above formula into the derivative of the Lyapunov function In the result: ; ; ; , ; ; The derivative of the Lyapunov function is : ; when If the above equation is satisfied, then: ; in for The optimal value function at the moment, , , for The optimal value function at the moment and the weight estimation of the critic network The error of the Actor network weight estimation The error and dynamic gain matrices The inverse matrix of = ; , , , , , are all positive constants, satisfying , , , , , ; , , , are all positive definite symmetric matrices of suitable dimensions, and are the minimum and maximum eigenvalues of the matrix, respectively. , are all given constants.
[0017] Compared with the prior art, the present invention has the following beneficial effects: Aiming at the trajectory tracking control problem of a four-Mecanum-wheeled mobile robot, this application uses the Actor-Critic synchronous learning algorithm to solve the optimal control strategy, uses interactive data to iteratively update the strategy, adjusts the controller output in real time according to the disturbance, and introduces an event trigger mechanism. The controller only updates the control signal when the trigger threshold is met; based on the optimal control theory, the consistent ultimate bounded stability of the four-Mecanum-wheeled mobile robot system is analyzed. This application solves the trajectory tracking problem of mobile robots in the presence of external disturbances such as sliding, and ensures the trajectory tracking accuracy while effectively reducing the computational burden. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 The schematic diagram and coordinate system of the four-Mecanum-wheel omnidirectional mobile robot; Figure 2 is a trajectory tracking error diagram using the control method of the present invention; Figure 3 Figure 2 is the trajectory tracking error graph using the time-triggered reinforcement learning control method; Figure 4 To use the control method of the present invention to move the robot Expected movement trajectory and actual movement trajectory diagram in direction; Figure 5 To use the control method of the present invention to move the robot Expected movement trajectory and actual movement trajectory diagram in direction; Figure 6 A diagram of the expected direction angle and the actual direction angle during the movement of the mobile robot using the control method of the present invention; Figure 7 Actor neural network weights for using the control method of the present invention The convergence curve diagram of Figure 8 To use the Critic neural network weights of the control method of the present invention The convergence curve diagram of Fig. 9 A diagram showing the triggering time and triggering time interval of the control method of the present invention within 120-125 seconds; Fig.10 A diagram showing the triggering moment and triggering time interval using a time-triggered reinforcement learning control method within 120-125 seconds; Fig.11 A diagram showing the number of triggers when using the control method of the present invention and a time-triggered reinforcement learning control method. DETAILED DESCRIPTION
[0019] The technical solution of the present application is described in detail below through specific embodiments and drawings. It should be understood that the embodiments of the present application and the specific features in the embodiments are detailed descriptions of the technical solution of the present application, rather than limitations on the technical solution of the present application, and the specific technical features may be combined with each other.
[0020] A reinforcement learning tracking control method for a mobile robot based on event triggering comprises the following steps: S1. Establish a wheel sliding model of the mobile robot, establish a forward kinematics model, an inverse kinematics model and a dynamic model of the mobile robot under the influence of sliding according to the wheel sliding model, substitute the forward kinematics model and the inverse kinematics model under the influence of sliding into its dynamic model to obtain a trajectory control model of the mobile robot; By controlling the speed of the four wheels of the mobile robot, the robot can move in all directions. The wheel sliding model is as follows: ; The forward kinematics model, inverse kinematics model and dynamics model under the influence of sliding are as follows: ; in, For the The angular velocity of the wheel rotation, For the The linear speed of the wheels, , represents the position and azimuth of the mobile robot, Represents a three-dimensional column vector The derivative of is the wheel radius of the mobile robot, is the sliding factor, , represents the rotation angular velocity of the four wheels, is a four-dimensional column vector The derivative of is the equivalent moment of inertia of the wheel, is the viscous friction coefficient between the wheel and the ground, is the control input of the mobile robot trajectory control model, and Azimuth is the state matrix of the parameters; ; ; is half the width of the mobile robot, is half the length of the mobile robot.
[0021] Substituting the forward and inverse kinematics model of the mobile robot into its dynamic model, we can obtain: ; ; ; ; ; ; The trajectory control model of the mobile robot is: ; in, is the linear velocity of the mobile robot, is the angular velocity of rotation, is a three-dimensional column vector The derivative of .
[0022] S2. Set the desired trajectory of the mobile robot, introduce the tracking error and calculate the tracking error dynamics, combine the error tracking dynamics with the desired trajectory, and obtain the error augmented model of the mobile robot; The formula for the expected trajectory of the mobile robot is: ; is the expected trajectory, is the desired position of the mobile robot, is the expected speed of the mobile robot, Represents the transpose of a matrix; Trajectory tracking error The calculation formula is: ; in, is the trajectory tracking error, is the current trajectory; The dynamics of the tracking error The expression formula is: .
[0023] Defines the state of the error tracking model , then the expression formula of the error augmentation model is: ; ; ; ; ; ; ; is the external disturbance, The expected trajectory The derivative of is the control input of the error augmentation model, is the steady-state control input.
[0024] S3. Design the value function, obtain the optimal control strategy, and add the event-triggered control scheme; S31. Design a value function and obtain a Hamiltonian function; ; in, represents the value function, is the discount factor, is the decay exponent, is a constant, st is the time interval, is the control input of the error augmentation model, , is the state of the error augmentation model, is the disturbance factor, ; and They are all constant matrices, and the formula is as follows: ; ; ; is the identity matrix; The value function is derived from the error augmentation model to obtain the approximate Hamiltonian function expression formula: ; in, represents the Hamiltonian function, Represents the derivative of the value function with respect to the state of the error augmented model.
[0025] S32. According to the Hamiltonian function, the HJB equation to be solved and the optimal control strategy are obtained; The HJB equation is: ; in, represents the optimal Hamiltonian function, represents the optimal control strategy, represents the optimal value function; Optimal control strategy for: ; in, is the optimal control strategy, For control strategy The admissible set of represents the derivative of the optimal value function with respect to the error augmented model state, is the state of the error augmentation model, represents a 4D positive definite matrix.
[0026] S33. Design an event-triggered control scheme to obtain an event-triggered control strategy; The event trigger conditions are: ; in Indicates the trigger error, represents the state of the error augmentation model at the current moment, Indicates The state of the error augmented model at the time of the next trigger, and is a positive definite symmetric matrix; The event trigger control strategy is: ; express The control input at the moment, express Control input at any time.
[0027] S4. Solve the approximate optimal control law and design the optimal trajectory tracking control algorithm of the mobile robot based on event triggering and neural network.
[0028] S41. Use a single-layer neural network to approximate the optimal value function and The analytical relationship of the equation is used to calculate the optimal control strategy; ; in, represents the optimal value function, is the transpose of the neural network weights, is the activation function of the neural network, is the approximation error of the neural network; The expression formula of the optimal control strategy is: ; in, represents the optimal control strategy, is a four-dimensional constant matrix, The error augmented model state matrix The transpose of According to Weierstrass's high-order approximation theory, the optimal value function and the optimal policy function Approximate and obtain the optimal value function and the optimal policy function The estimated values are: ; in, and They represent the estimated weights of the neural network respectively.
[0029] S42. Definition The Bellman error of the function is used to design the weight update law of the neural network based on the gradient descent method.
[0030] The optimal value function and the optimal policy function Substituting the estimated value of into the Hamiltonian equation, we get the approximate Hamiltonian equation for: ; in, is the approximate Hamiltonian function.
[0031] The calculation formula for the Bellman error that defines the Hamiltonian function is: ; in, is the approximate Hamiltonian function, is the optimal Hamiltonian function; Bellman Error The expression formula is: ; The weight update strategy is: ; in, Represents the neural network weights The derivative of Represents the neural network weights The derivative of and represents the parameter matrix, , , and is a positive constant, is the constant forgetting factor.
[0032] S5. Analyze the consistent ultimate bounded stability of the mobile robot and complete the trajectory tracking control of the mobile robot based on event triggering and reinforcement learning.
[0033] Select the Lyapunov function of the entire mobile robot trajectory tracking control system for: ; in, is the optimal value function, is the dynamic gain matrix The inverse matrix of and is a positive constant, and is the weight estimation error, , ,when When , we can derive it to get: ; in, represents the Lyapunov function The derivative of is the dynamic gain matrix The derivative of the inverse matrix of represents the derivative of the optimal value function, and They are the derivatives of the weight estimation errors of the Critic network and the Actor respectively. The optimal control strategy is updated according to the weights. and The derivative is: ; in, represents the normalized parameter matrix, , represents the parameter matrix, , yes A scalar quantity related to , and is the transition matrix, , , , and is a positive constant.
[0034] Substitute the above formula into the derivative of the Lyapunov function In the result: ; ; ; , ; ; The derivative of the Lyapunov function is : ; when If the above formula satisfies: ; in for The optimal value function at the moment, , , for The optimal value function at the moment and the weight estimation of the critic network The error of the Actor network weight estimation The error and dynamic gain matrices The inverse matrix of = ; , , , , , are all positive constants, satisfying , , , , , ; , , , are all positive definite symmetric matrices of suitable dimensions, and are the minimum and maximum eigenvalues of the matrix, respectively. are all constants.
[0035] In summary, according to the standard Lyapunov stability theorem, the error tracks the state of the model and the estimated error of the weights and are all uniformly ultimately bounded.
[0036] Given the desired trajectory that the mobile robot tracks , initial position .
[0037] The simulation parameters are shown in Table 1: Table 1 Simulation parameters .
[0038] Figure 2 and Figure 3 The following are the tracking error comparison diagrams of the control method provided by the present invention and the time-triggered reinforcement learning control method. The blue line e1 is the tracking error in the x direction, the red line e2 is the tracking error in the y direction, and the black line e3 is the direction angle. The tracking error, Figure 4 , Figure 5 , Figure 6 The robot is in the x direction, y direction and direction angle respectively The red line is the expected trajectory, and the blue line is the response result corresponding to the control method proposed in the present invention.
[0039] Figure 7 , Figure 8 They are respectively the weight change curves of the Actor neural network and the Critic neural network proposed in the present invention over time. The two figures reflect the training process of the neural network. With the iteration of training, the weight values converge, and the weight convergence values of the two neural networks are equal.
[0040] Fig. 9 and Fig.10 The trigger time interval diagrams of the control method proposed in the present invention and the time-triggered reinforcement learning control method are shown respectively, where the x coordinate is the trigger time within 120-125 seconds, and the y coordinate is the time interval between the current trigger and the last trigger. Fig.11 The triggering times of the control method proposed in the present invention and the time-triggered reinforcement learning control method are shown in Figure 2. Within 0-250 seconds, the time-triggered reinforcement learning control method was triggered 141,979 times, while the control method proposed in the present invention was triggered 81,646 times, and the triggering times were reduced by 42.49%. Figure 2-Figure 11 It can be seen that the control method proposed in the present invention can effectively reduce the calculation burden while ensuring the tracking speed and tracking accuracy.
[0041] The above is only a preferred implementation of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A reinforcement learning tracking control method for a mobile robot based on event triggering, characterized in that: The following steps are involved: S1. Establish a wheel sliding model of the mobile robot, establish a forward kinematics model, an inverse kinematics model and a dynamics model of the mobile robot under the influence of sliding according to the wheel sliding model, substitute the forward kinematics model and the inverse kinematics model under the influence of sliding into its dynamics model to obtain a trajectory control model of the mobile robot; S2. Set the desired trajectory of the mobile robot, introduce the tracking error and calculate the tracking error dynamics, combine the tracking error dynamics with the desired trajectory, and obtain the error augmented model of the mobile robot; S3. Design the value function, obtain the optimal control strategy to be solved, and add the event-triggered control scheme; S4. Use neural networks to solve approximate optimal control strategies and design an optimal trajectory tracking control algorithm for mobile robots based on event triggering and neural networks; S5. Analyze the consistent ultimate bounded stability of the mobile robot and complete the event-triggered reinforcement learning trajectory tracking control of the mobile robot.
2. The event-triggered mobile robot reinforcement learning tracking control method according to claim 1 is characterized in that: Substituting the forward kinematics model and inverse kinematics model of the mobile robot under the influence of sliding into its dynamic model, the trajectory control model of the mobile robot is obtained, and the formula is as follows: ; ; ; ; ; ; The expression formula of the mobile robot trajectory control model is: ; represents the position and azimuth of the mobile robot, Represents a three-dimensional column vector The derivative of is the wheel radius of the mobile robot, is the sliding factor, , is the equivalent moment of inertia of the wheel, is the viscous friction coefficient between the wheel and the ground, is the control input of the mobile robot trajectory control model, and Azimuth is the state matrix of the parameters; is a three-dimensional column vector The derivative of .
3. The event-triggered mobile robot reinforcement learning tracking control method according to claim 1, characterized in that: The formula for the expected trajectory of the mobile robot in step S2 is: ; in, is the expected trajectory, is the desired position of the mobile robot, is the expected speed of the mobile robot, Represents the transpose of a matrix; Trajectory tracking error The calculation formula is: ; in, is the trajectory tracking error, is the actual trajectory; Tracking Error Dynamics The expression formula is: ; Define the state of the error augmentation model as , then the expression formula of the error augmentation model is: ; ; ; ; ; ; ; is the external disturbance, The expected trajectory The derivative of is the control input of the error augmentation model, is the steady-state control input.
4. The event-triggered mobile robot reinforcement learning tracking control method according to claim 1, characterized in that: In step S3: S31. Design a value function and obtain a Hamiltonian function; S32. According to the Hamiltonian function, the HJB equation and the optimal control strategy to be solved are obtained; S33. Design an event-triggered control scheme and obtain an event-triggered control strategy.
5. The event-triggered mobile robot reinforcement learning tracking control method according to claim 4 is characterized in that: The value function expression formula in step S31 is: ; in, represents the value function, is the discount factor, is the decay exponent, is a constant, st is the time interval, is the control input of the error augmentation model, , is the state of the error augmentation model, is the disturbance factor, ; and They are all constant matrices, and the formula is as follows: ; ; ; is the identity matrix; Derivative the value function with respect to the error augmentation model to obtain the approximate Hamiltonian function, which is expressed as: ; in, represents the Hamiltonian function, Represents the derivative of the value function with respect to the state of the error augmented model.
6. The event-triggered mobile robot reinforcement learning tracking control method according to claim 4, characterized in that: The HJB equation in step S32 is: ; in, represents the optimal Hamiltonian function, represents the optimal control strategy, represents the optimal value function; Optimal control strategy for: ; in, is the optimal control strategy, For control strategy The admissible set of represents the derivative of the optimal value function with respect to the state of the error augmented model, is the state of the error augmentation model, represents a four-dimensional positive definite matrix; The event triggering control scheme in step S33 is: ; in Indicates the trigger error, represents the state of the error augmentation model at the current moment, Indicates The state of the error augmented model at the time of the next trigger, and is a positive definite symmetric matrix; The event trigger control strategy is: ; express The control input at the moment, express Control input at any time.
7. The event-triggered mobile robot reinforcement learning tracking control method according to claim 1, characterized in that: Step S4 includes the following sub-steps: S41. Use a single-layer neural network to approximate the optimal value function and calculate the optimal control strategy based on the analytical relationship of the HJB equation; S42. Define the Bellman error of the HJB function and design the update strategy of the neural network weights based on the gradient descent method.
8. The event-triggered mobile robot reinforcement learning tracking control method according to claim 7, characterized in that: Step S41 includes the following formula: ; in, represents the optimal value function, is the transpose of the neural network weights, is the activation function of the neural network, is the approximation error of the neural network; The expression formula of the optimal control strategy is: ; in, represents the optimal control strategy, is a four-dimensional constant matrix, is the system state matrix The transpose of According to Weierstrass's high-order approximation theory, the optimal value function and the optimal policy function Approximate and obtain the optimal value function and the optimal policy function The estimated values are: ; in, and They represent the estimated weights of the neural network respectively.
9. The method for tracking and controlling a mobile robot based on event triggering according to claim 7, characterized in that: Step S42 includes the following formula: The optimal value function and the optimal policy function Substituting the estimated value of into the Hamiltonian equation, we get the approximate Hamiltonian equation for: ; in, is the approximate Hamiltonian function; The calculation formula for the Bellman error that defines the Hamiltonian function is: ; in, is the approximate Hamiltonian function, is the optimal Hamiltonian function; Bellman Error The expression formula is: ; The weight update strategy is: ; in, Represents the neural network weights The derivative of Represents the neural network weights The derivative of and represents the parameter matrix, , , and is a positive constant, is the constant forgetting factor.
10. The event-triggered mobile robot reinforcement learning tracking control method according to claim 1, characterized in that: The consistent final boundedness of the mobile robot system analyzed in step S5 includes the following formula: Select the Lyapunov function of the entire mobile robot trajectory tracking control system for: ; in, is the optimal value function, is the dynamic gain matrix The inverse matrix of and is a positive constant, and is the weight estimation error, , ,when When , we can derive it to get: ; in, represents the Lyapunov function The derivative of is the dynamic gain matrix The derivative of the inverse matrix of represents the derivative of the optimal value function, and They are the derivatives of the weight estimation errors of the Critic network and the Actor respectively. According to the weight update strategy, the weight estimation error of the neural network is and The derivative is: ; in, represents the normalized parameter matrix, , represents the parameter matrix, , yes A scalar quantity related to ; and is the transition matrix, and the formula is as follows: ; ; ; Substitute the above formula into the derivative of the Lyapunov function In the result: ; ; ; , ; ; The derivative of the Lyapunov function is : ; when If the above equation is satisfied, then: ; in for The optimal value function at the moment, , , and for The optimal value function at the moment and the weight estimation of the critic network The error of the Actor network weight estimation The error and dynamic gain matrices The inverse matrix of = ; , , , , , are all positive constants, satisfying , , , , , ; , , , are all positive definite symmetric matrices of suitable dimensions, and are the minimum and maximum eigenvalues of the matrix, respectively. , are all given constants.
Citation Information
Patent Citations
Wheeled mobile robot event trigger tracking control method based on deterministic learning
CN112051734A
Robot trajectory tracking optimal control method based on an event trigger mechanism
CN113093548A
Decentralized tracking control method for mechanical arm based on event triggering-neural dynamic programming
CN113211446A
Unmanned aerial vehicle trajectory tracking method based on reinforcement learning and event triggering
CN116300991A
Omnidirectional trolley trajectory tracking optimal control method based on reinforcement learning
CN118838360A
Cited By
Networked mobile robot event triggering data driving control method and system
CN120722756A
Mechanical arm system robust tracking control method based on self-adaptive dynamic programming
CN121589821A
Mobile robot dynamic event trigger control method based on reinforcement learning
CN122308123A
Reinforcement learning based dynamic event-triggered control method for mobile robots
CN122308123B