Mechanical arm self-adaptive preset performance control method based on reinforcement learning
By adopting an adaptive preset performance control method based on reinforcement learning, the tracking control problem of the robotic arm under external disturbances and limited communication resources is solved. This method achieves high-precision tracking of the robotic arm system in both transient and steady states, avoids Zeno's phenomenon, and saves control resources.
Patent Information
- Application Number
- CN202511526853.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-24
- Publication Date
- 2026-01-20
AI Technical Summary
Under conditions of external disturbances and limited communication resources, the optimal/adaptive and event-triggered control of existing robotic arms cannot simultaneously guarantee the tracking accuracy of transient and steady states, and there is a risk of Zeno's phenomenon, which leads to waste of control resources and a decrease in accuracy.
An adaptive preset performance control method based on reinforcement learning is adopted. By constructing a dynamic model of a multi-link robotic arm, setting tracking error variables and performing equivalent error transformation, and combining backstepping method and command filter, Hamilton-Jacobi-Bellman equation and neural network are used for optimization control. A relative threshold event triggering strategy is designed to ensure that the tracking error of the robotic arm is within the preset range.
This effectively avoids Zeno's phenomenon, ensures the tracking accuracy of the robotic arm under external disturbances and limited communication resources, reduces control resource consumption, and improves the system's tracking performance and engineering usability.
Smart Images

Figure CN121361084A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of mechanical arm control, and particularly relates to a mechanical arm adaptive preset performance control method based on reinforcement learning. BACKGROUND
[0002] With the in-depth promotion of the industrial 4.0 era, industrial mechanical arms are more and more widely applied in the fields of intelligent manufacturing, logistics and warehousing, medical surgery and other automation fields due to their high precision, high flexibility and high automation. At the same time, with the development of control technology and the increasing of global energy crisis, how to realize the efficient use of control resources under the premise of guaranteeing control performance has attracted more and more attention and research. As an important branch of modern control theory, optimal control theory can minimize the consumption of control resources while achieving the expected control target of the system by constructing a performance index function containing control objectives and constraint conditions. Under this background, the research on the optimal control of mechanical arms has important academic value and practical significance for improving the control performance of industrial mechanical arms and realizing efficient control.
[0003] Preset performance control sets performance boundaries for the variables of the controlled system, then maps the constrained system into an equivalent unconstrained system for controller design, and at the same time, takes into account the transient performance and steady-state performance of the system. The preset performance control of the mechanical arm not only improves the tracking accuracy, but also simplifies the difficulty of controller design, which has important value for the expansion application of the mechanical arm in the high-end manufacturing field such as precision machining and micro-operation.
[0004] In actual operation, the robot arm often faces external disturbance and limited communication resources, so the controller design must consider how to reduce the signal transmission frequency between the actuator and the controller. Event-triggered control can effectively solve this problem, however, an important consideration of event-triggered control is to avoid the occurrence of Zeno phenomenon, that is, the triggering condition does not allow infinite triggering in a finite time. Studies have shown that in the absence of external disturbance, the event-triggered strategy will not appear Zeno phenomenon. When there is any small external disturbance in the system, avoiding Zeno phenomenon will be a challenging task. Therefore, how to avoid the Zeno phenomenon of the robot arm system with external disturbance has attracted the attention of many scholars. L. Xing et al. disclosed an event-triggered adaptive control for a class of uncertain nonlinear systems, in the controller design, it is more appropriate to consider the event-triggered threshold as a dynamically changing threshold [L. Xing, C. Wen, Z. Liu, H. Su, and J. Cai, “Event-triggered adaptive control for a class of uncertain nonlinear systems,” IEEE Transactions on Automatic Control, vol. 62, pp. 2071-2076, Apr. 2017.]. This means that when the amplitude of the control signal is large, the triggering threshold is also relatively large, so that the control signal remains constant for a long time. For the system state close to the equilibrium point, a smaller triggering threshold can be used. Therefore, the event-triggered relative threshold strategy of the robot arm system is also worthy of further research.
[0005] The existing optimal / adaptive and event-triggered control for robot arms generally has the problem that the preset performance error is difficult to guarantee in both transient and steady states under external disturbance and limited communication resources. In order to reduce communication load, a conservative triggering threshold is used, which leads to a decrease in precision. Triggering too frequently can easily induce Zeno risk. At the same time, backstepping design often requires high-order derivatives, which is complex to implement and has a heavy parameter tuning burden. It is difficult to balance stability and optimality. Based on the above shortcomings, the present application proposes a control method under the conditions of external disturbance and limited communication resources. On the premise of strictly avoiding high-frequency triggering / Zeno risk, the communication and control overhead are reduced, and the tracking error of the robot arm satisfies the preset performance constraint in both transient and steady states, thereby improving the tracking accuracy and engineering usability of the system. SUMMARY
[0006] In view of the above shortcomings in the prior art, the present application provides a robot arm adaptive preset performance control method based on reinforcement learning, which guarantees the tracking performance of the robot arm system and keeps the tracking error within the preset range under the conditions of external disturbance and limited communication resources.
[0007] To achieve the above object, the technical scheme adopted by the present application is as follows:
[0008] A mechanical arm adaptive preset performance control method based on reinforcement learning, comprising the following steps: A dynamics model of a multi-link mechanical arm with external disturbance is constructed; A tracking error variable of the multi-link mechanical arm is set, a preset performance control function is established to constrain the tracking error, and an equivalent error conversion is performed on the tracking error to obtain an unconstrained error variable; A compensation tracking error signal based on the backstepping method and the command filter is established according to the unconstrained error variable, and a performance index related to the compensation tracking error signal, a virtual control signal and an actual control signal is constructed to establish a Hamilton-Jacobi-Bellman equation; A reinforcement learning method based on a neural network is used to solve the Hamilton-Jacobi-Bellman equation, and a Lyapunov function related to the compensation tracking error signal and a neural network weight estimation error is constructed in combination with the Lyapunov stability theory to determine the virtual control signal, the actual control signal and the parameter adaptive law.
[0009] Further, the dynamics model of the multi-link mechanical arm with external disturbance is specifically as follows: ; wherein, is the joint angle of the mechanical arm, is the angular velocity of the mechanical arm, is the angular acceleration of the mechanical arm, M is a symmetric positive inertia matrix, C is a Coriolis force matrix, G is a gravity vector, is the control torque, is the external disturbance; the state variable is defined as , the system output is , and the control input is , and the dynamics model of the multi-link mechanical arm is represented as: ; wherein, , .
[0010] Further, the tracking error variable of the multi-link mechanical arm is specifically designed as: ; wherein, is the tracking error, is the system output, is the desired trajectory.
[0011] Further, the tracking error is converted into an equivalent error Setting constraints wherein, is the number of links of the robot arm, is a preset performance control function, are positive parameters designed respectively, is a lower bound proportional coefficient, is an upper bound proportional coefficient, t is a current time, and T is a transpose symbol.
[0012] Further, the tracking error is converted into an equivalent error ; wherein, is a conversion function, is a converted tracking error, ;
[0013] Further, the tracking error is converted into an equivalent error .
[0014] Further, a compensation tracking error signal based on a backstepping method and a command filter is established according to the unconstrained error variable, comprising:
[0015] Setting a coordinate transformation related to the tracking error is: ; wherein, is an unconstrained tracking error variable, is a velocity error variable, is an output of the command filter, and s is an unconstrained error variable;
[0016] Setting a compensation tracking error signal is: ; wherein, is a compensation tracking error signal, is a compensation signal.
[0017] Further, the output of the command filter is obtained according to the following formula: ; is an output of the command filter, ; wherein, is a filtered virtual control output of the i th link, is a command filter bandwidth of the i th link, is an auxiliary state variable of the filter corresponding to the i th link, is a filtered virtual control output initial value of the i th link, the initial value of the auxiliary state variable of the filter corresponding to the i-th connecting rod, the damping ratio of the filter corresponding to the i-th connecting rod, , the virtual control signal of the input filter.
[0018] Further, by constructing a performance index related to the compensated tracking error signal, the virtual control signal and the actual control signal, a Hamilton-Jacobi-Bellman equation is established, including:
[0019] By constructing a performance index related to the compensated tracking error signal, the virtual control signal and the actual control signal, it is: ; ; wherein, the optimal cumulative cost of the virtual control subsystem, the optimal cumulative cost of the actual control subsystem, the minimum value of the admissible control set, the virtual control signal, the actual control signal, the optimal virtual control signal, the optimal actual control signal, the instantaneous cost function, , the optimal instantaneous cost function, the admissible control set, the compact set containing the origin, the integral dummy variable; The Hamilton-Jacobi-Bellman equation is established as: ; ; wherein, the Hamilton-Jacobi-Bellman equation of the virtual control subsystem, the Hamilton-Jacobi-Bellman equation of the actual control subsystem, T is the transpose symbol, the derivative of the compensated tracking error signal.
[0020] Further, a neural network-based reinforcement learning method is used to solve the Hamilton-Jacobi-Bellman equation, and a Lyapunov function related to the compensated tracking error signal and the neural network weight estimation error is constructed based on Lyapunov stability theory, to determine the virtual control signal, the actual control signal and the parameter adaptive law, including: Step 1, calculate : The first continuous auxiliary function is constructed, and a radial basis function neural network is used to estimate the continuous auxiliary function, to obtain the partial derivative of the performance index with respect to and the optimal virtual control signal as: ; ; wherein is a preset performance transformation function, , is a design parameter, is a diagonal matrix, , is an ideal weight vector of the neural network, is a basis function vector of the neural network, is an estimation error of the neural network, is a derivative of the desired trajectory, is a derivative of the preset performance control function.
[0021] The optimal virtual control signal and the partial derivative of the performance index with respect to are taken as an actor and a critic of a reinforcement learning method, specifically as: ; ; wherein is an estimated value of the optimal virtual control signal, is an estimated value of , is a weight of the actor and critic neural network.
[0022] A condition for constructing a neural network parameter adaptive law is constructed by using a Bellman residual, and the neural network parameter adaptive law is set to be: ; ; wherein is a derivative of , are learning rates of the actor and critic neural networks respectively, is a derivative of .
[0023] In step 2, the is calculated as: A second continuous auxiliary function is constructed, and a radial basis function neural network is used to estimate the continuous auxiliary function, to obtain the partial derivative of the performance index with respect to and the optimal virtual control signal as: ; ; where, is a design parameter, f is a nonlinear dynamics equation in the system dynamics model, is the derivative of the command filter output , is the ideal weight vector of the neural network, is the basis function vector of the neural network, is the estimation error of the neural network.
[0024] The optimal virtual control signal and the partial derivative of the performance index with respect to are taken as the actor and critic of the reinforcement learning method, specifically: ; ; where, is the estimated value of the optimal virtual control signal, is the estimated value of , are the weights of the actor and critic neural networks, respectively.
[0025] The condition for constructing the neural network parameter adaptive law using the Bellman residual is constructed, and by constructing a positive definite function, the neural network parameter adaptive law is set as: ; where, is the derivative of , are the learning rates of the actor and critic neural networks, respectively, is the derivative of .
[0026] Further, the relative threshold event-triggered strategy is used to update the signals transmitted between the actuator and the controller, specifically: ; where, is the designed intermediate control signal, , , , are all design parameters, is the actual control signal of the i-th link, denotes the time of the i-th link k-th control input update, is the maximum lower bound of the set, is the event-triggered error, .
[0027] The present application has the following beneficial effects:
[0028] (1) The present application designs a mechanical arm adaptive preset performance control method based on reinforcement learning for a mechanical arm system, which can effectively solve the mechanical arm tracking control problem under the existence of external disturbance and limited communication resources.
[0029] (2) The present application guarantees the transient and steady-state tracking accuracy through preset performance constraints, and uses the Actor-Critic reinforcement learning strategy to solve the HJB equation to realize the performance function minimization, which minimizes the control input under the premise of guaranteeing the tracking accuracy to meet the constraints, thereby saving the control resources.
[0030] (3) The present application successfully avoids the "calculation explosion" problem caused by repeatedly solving the derivative of the virtual control input signal by introducing an instruction filter, and compensates for the error of the filter, thereby effectively improving the tracking performance of the system.
[0031] (4) The present application adopts an event-triggered strategy of a relative threshold, which effectively reduces the signal transmission frequency between the controller and the actuator while guaranteeing the tracking performance, thereby saving the communication resources. BRIEF DESCRIPTION OF DRAWINGS
[0032] Figure 1 It is a flowchart of a mechanical arm adaptive preset performance control method based on reinforcement learning; Figure 2 It is a state of a 2-link mechanical arm system and a trajectory diagram of a reference signal; Figure 3 It is a trajectory diagram of the tracking error of the 2-link mechanical arm system; Figure 4 It is a trajectory diagram of the adaptive parameter ; Figure 5 It is a schematic diagram of the event-triggered interval and the number of triggers of the event-triggered control signal and ; Figure 6 It is a schematic diagram of the event-triggered interval and the number of triggers of the event-triggered control signal and . DETAILED DESCRIPTION
[0033] The specific embodiments of the present application are described below to facilitate the understanding of the present application for those skilled in the art, but it should be clear that the present application is not limited to the scope of the specific embodiments, and for those skilled in the art, it is obvious that various changes are within the spirit and scope of the present application defined and determined by the appended claims, and all the inventions utilizing the concept of the present application are within the scope of protection.
[0034] The embodiment of the present application is directed to a mechanical arm system, which solves the tracking control problem under external disturbance and limited communication resources. First, a preset performance control function is established to constrain the tracking error, and the corresponding equivalent error transformation is performed to obtain an unconstrained error variable. Then, combining the Backstepping method, the command filter, the Actor-Critic reinforcement learning based on neural network and the Lyapunov stability theory, an adaptive event-triggered optimal controller based on reinforcement learning is developed. Under the condition of external disturbance and limited communication resources, the proposed control strategy not only ensures that the tracking error of the closed-loop system meets the preset performance, but also ensures that all signals in the system are bounded, in addition, there is a lower limit for the time interval of event triggering, which avoids Zeno behavior.
[0035] As shown in Figure 1 , the adaptive preset performance control method for a mechanical arm based on reinforcement learning provided by the embodiment of the present application includes the following steps S1 to S4:
[0036] S1, a dynamic model of a multi-link mechanical arm with external disturbance is constructed.
[0037] In the present embodiment, the step S1 of constructing the dynamic model of the multi-link mechanical arm with external disturbance is specifically: ; wherein, is the joint angle of the mechanical arm, is the angular velocity of the mechanical arm, is the angular acceleration of the mechanical arm, M is a symmetric positive inertia matrix, C is a Coriolis force matrix, G is a gravity vector, is the control torque, is the external disturbance.
[0038] The state variable is defined as , the system output is , and the control input is , then the dynamic model of the n-link mechanical arm can be arranged as: ; wherein, , .
[0039] In this embodiment, the dynamics model of the n-link robot arm with external disturbance satisfies the following conditions:
[0040] Condition 1, reference signal and are smooth and bounded.
[0041] Condition 2, unknown external disturbance is bounded, that is where is an unknown positive constant, denotes the 2-norm.
[0042] S2, set the tracking error variable of the multi-link robot arm, establish the preset performance control function to constrain the tracking error, and perform equivalent error conversion on the tracking error to obtain an unconstrained error variable.
[0043] In this embodiment, the tracking error variable of the multi-link robot arm is specifically designed as: ; where is the tracking error, is the system output, is the desired trajectory.
[0044] The tracking error is set to be constrained where is the number of robot arm links, is the preset performance control function, are the designed positive parameters, respectively, is the lower bound proportion coefficient, is the upper bound proportion coefficient, t is the current time, and T is the transpose symbol.
[0045] This embodiment converts the tracking error to: ; where is the converted error variable, is the conversion function, which is specifically: ; to obtain the unconstrained error variable .
[0046] S3, establish a compensation tracking error signal based on the backstepping method and the command filter according to the unconstrained error variable, and establish the Hamilton-Jacobi-Bellman (HJB) equation by constructing a performance index related to the compensation tracking error signal, the virtual control signal, and the actual control signal.
[0047] In this embodiment, the coordinate transformation related to tracking error is defined and tracked as: ; where, is the unconstrained tracking error variable, is the velocity error variable, is the command filter output, which has the form: ; where, is the filtered virtual control output of the ith link, is the command filter bandwidth of the ith link, , is the auxiliary state variable of the ith filter, is the initial value of the filtered virtual control output of the ith joint, is the initial value of the auxiliary state variable of the ith filter, is the damping ratio of the ith filter, , , is the virtual control signal input to the filter. The command filter used solves the problem of complexity explosion caused by repeated derivation of virtual signals in backstepping control, but the error caused by the filter will affect the control performance of the system. In the following, a compensation tracking error signal is introduced to eliminate the influence of the filter error.
[0048] The compensation tracking error signal is defined as: ; where, is the compensation signal, which is designed as: ; where, are all design parameters, is a pre-set performance transformation function, , .
[0049] The performance index is defined as: ; ; where, is the optimal cumulative cost of the virtual control subsystem, is the optimal cumulative cost of the actual control subsystem, is the minimum value of the admissible control set, is the virtual control signal, is the actual control signal, is the optimal virtual control signal, is the optimal actual control signal, is the admissible control set, is a compact set containing the origin, is an integral dummy variable, is the instantaneous cost function, , is the optimal instantaneous cost function defined as:
[0050] .
[0051] Further, the HJB equation is obtained as: ; ; where, is the Hamilton-Jacobi-Bellman equation of the virtual control subsystem, is the Hamilton-Jacobi-Bellman equation of the actual control subsystem, T is the transpose symbol, is the derivative of the compensation tracking error signal. By solving the HJB equation, the optimal control can be obtained by minimizing the performance index .
[0052] S4, the Hamilton-Jacobi-Bellman equation is solved by using the neural network-based reinforcement learning method, and a Lyapunov function related to the virtual control signal, the actual control signal and the parameter adaptive law is constructed by combining the Lyapunov stability theory and the estimation error of the neural network weight.
[0053] In this embodiment, in the first step, the , the optimal virtual control signal can be obtained as ; The first continuous auxiliary function is introduced as: ; where, is a diagonal matrix, The radial basis function neural network has the universal approximation property and can be used to estimate the continuous auxiliary function . Therefore, there exists a radial basis function neural network such that: ; where, denotes the ideal weight vector of the neural network, denotes the number of rules of the neural network, denotes the basis function vector of the neural network, and the Gaussian function is selected , are the center and width of the Gaussian function, respectively, estimation error of the neural network satisfies where is an unknown constant.
[0054] The partial derivative of the performance index with respect to and the optimal virtual control signal are: ; .
[0055] Since is unknown, the optimal virtual control signal is not available. This embodiment uses the Actor-Critic policy to obtain an approximate solution to the HJB equation, and selects the partial derivative of the performance index with respect to as the actor and critic, respectively. The actor is responsible for generating the control input of the robot arm, and the critic is responsible for evaluating the value of the robot arm action to provide feedback information for the actor to achieve input optimization. The generated actor and critic are: ; ; where is the estimated value of the optimal virtual control signal , and is the estimated value of , which will be used as the virtual control input of the robot arm system in the first step; is the weight of the actor and critic neural networks.
[0056] This embodiment uses the Bellman residual to construct the condition of the neural network parameter adaptive law. Define the Bellman residual as: .
[0057] Calculate the partial derivative of the HJB equation with respect to the weight of the actor neural network: ; Further, construct a positive definite function , and the weights of the actor and critic neural networks need to achieve the negative gradient descent of .
[0058] Set the Lyapunov function of the first step as: ; where is the parameter estimation error.
[0059] For derivative, combined with the requirements of positive definite function , the parameter adaptive law of can be designed as: ; ; where are the learning rates of actor and critic neural networks respectively. Substituting the virtual control signal and the parameter adaptive law into , we can get: ; where is the maximum eigenvalue of , and is the minimum eigenvalue of .
[0060] In the second step, we can get the optimal actual control signal as: .
[0061] The second continuous auxiliary function is introduced as: ; where f is the nonlinear dynamics equation in the system dynamics model, is the derivative of the command filter output .
[0062] The auxiliary function is estimated using a radial basis function neural network: ; where denotes the ideal weight vector of the neural network, denotes the number of rules of the neural network, denotes the basis function vector of the neural network, and Gaussian function is selected , the center and width of the Gaussian function respectively, and the estimation error of the neural network satisfies , where is an unknown positive constant.
[0063] The partial derivative of the performance index with respect to and the optimal actual control signal can be obtained as: ; .
[0064] Since unknown, thus the optimal virtual control signal is unavailable. An approximate solution of the HJB equation is sought using the Actor-Critic policy, which selects the partial derivatives of the virtual control signal and the performance index with respect to as the actor and critic, respectively. The generated actor and critic are where is the estimate of , is the estimate of , and are the weights of the actor and critic neural networks, respectively.
[0065] The condition for constructing the neural network parameter adaptive law using the Bellman residual is defined. The Bellman residual is defined as
[0066] The partial derivative of the HJB equation with respect to the weights of the actor neural network is calculated as
[0067] Further, a positive definite function is constructed, and the weights of the actor and critic neural networks need to implement the negative gradient descent of
[0068] This embodiment designs a novel relative threshold event-triggered policy and develops an adaptive event-triggered optimal controller. The designed event-triggered policy is expressed as where , , , are the design parameters, is a column vector of all ones with length 2; , is the designed intermediate control signal; denotes the event-triggered error, denotes the time of the i-th link's k-th control input update, is the maximum lower bound of the set; when the i-th event-triggered condition is satisfied, the control input is updated as For the new trigger time, otherwise, in the trigger interval , the control input remains .
[0069] The Lyapunov function of step 2 is set as: ; where, is the parameter estimation error.
[0070] Take the derivative of , combined with the requirements of positive definite function , the parameter adaptive law of can be designed as: ; where, is the design parameter.
[0071] This embodiment combines stability theory to prove the stability of the closed-loop system, ensuring that the tracking error can converge within the preset range, guaranteeing that all signals in the closed-loop system are bounded, and there is a lower limit to the time interval of event triggering, avoiding Zeno behavior. The specific process is: According to the results of step 2, the derivative of Lyapunov function can be obtained as: ; where, , , is the maximum eigenvalue of . is the minimum eigenvalue of , is the minimum eigenvalue of , is the minimum eigenvalue of . According to the existing lemma, we can get , that is, the compensation error signal is bounded and will converge to a compact set . At the same time, it can also be concluded that is bounded in the closed-loop system. According to the relationship , in order to obtain the boundedness of the tracking error , we need to prove that the compensation signal is bounded.
[0072] Construct the following Lypunov function:
[0073] .
[0074] Its derivative is obtained as: ; wherein, , is the minimum eigenvalue of is the minimum eigenvalue of
[0075] Therefore, the compensation signal is bounded, and the tracking error is also bounded.
[0076] This embodiment uses a real robotic arm experimental platform to conduct angle tracking experiments. The joint angles of the robotic arm are calculated by the encoder values of each joint, the angular velocities and the angular accelerations are calculated by the difference method, and the control inputs are the motor torques . The robotic arm controls the joint movement by establishing a simulation model of the controller on the host computer and compiling it to the lower computer. The experiment is conducted on the first two joints of a six-degree-of-freedom robotic arm, the initial joint angles are , the reference signal is selected as , , and the total running time is set to 60s. The experimental results are shown in Figures 2-6 .
[0077] Figure 2 show the trajectories of the system state and the reference signal . Figure 3 show the trajectory of the system tracking error . Figure 4 show the trajectory of the adaptive parameter . Figure 5 show the event-triggered intermittence and the number of triggers of the event-triggered control signal and . Figure 6 show the event-triggered intermittence and the number of triggers of the event-triggered control signal and . The experimental results show that the robotic arm adaptive preset performance control method based on reinforcement learning proposed in the invention has good tracking performance, makes the tracking error meet the preset performance, ensures the boundedness of all signals in the closed-loop system, and reduces the transmission frequency of the control signal, saving communication resources.
[0078] The present application is described in reference to the flowchart and / or block diagram of the method, device (system) and computer program product according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate a device that implements the flow Figure One one or more flows and / or blocks. Figure One one or more blocks.
[0079] These computer program instructions can also be stored in a computer readable memory that can direct the computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce a manufactured product including instruction devices that implement the flow Figure One one or more flows and / or blocks. Figure One one or more blocks.
[0080] These computer program instructions can also be loaded into a computer or other programmable data processing device, so that a series of operation steps are performed on the computer or other programmable device to produce a computer implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the flow Figure One one or more flows and / or blocks. Figure One one or more blocks.
[0081] The principles and implementation of the present application are described in the specific embodiments, and the above description of the embodiments is only for the purpose of helping the reader to understand the method of the present application and its core idea; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation and application range, and the above description of the present application should not be understood as a limitation of the present application.
[0082] Those skilled in the art will realize that the embodiments described herein are for the purpose of helping the reader to understand the principles of the present application, and should be understood as not limited to such specific statements and embodiments. Those skilled in the art can make various other specific modifications and combinations according to the technical inspiration disclosed in the present application without departing from the essence of the present application, and these modifications and combinations are still within the scope of protection of the present application.
Claims
1. A method for adaptive preset performance control of a robot arm based on reinforcement learning, characterized in that, The method comprises the following steps: A dynamic model of the multi-link robot arm with external disturbance is constructed; A tracking error variable of the multi-link robot arm is set, a preset performance control function is used to constrain the tracking error, and an equivalent error conversion is performed on the tracking error to obtain an unconstrained error variable; A compensation tracking error signal based on a backstepping method and an instruction filter is established according to the unconstrained error variable, and a Hamilton-Jacobi-Bellman equation is established by constructing a performance index related to the compensation tracking error signal, a virtual control signal and an actual control signal; A reinforcement learning method based on a neural network is used to solve the Hamilton-Jacobi-Bellman equation, and a Lyapunov function related to the compensation tracking error signal and a neural network weight estimation error is constructed in combination with a Lyapunov stability theory to determine the virtual control signal, the actual control signal and a parameter adaptive law.
2. The reinforcement learning based adaptive preset performance control method for a robot arm according to claim 1, characterized in that, The dynamic model of the multi-link robot arm with external disturbance is specifically as follows: ; wherein, is the joint angle of the robot arm, is the angular velocity of the robot arm, is the angular acceleration of the robot arm, M is a symmetric positive definite inertia matrix, C is a Coriolis matrix, G is a gravity vector, is the control torque, is an external disturbance; Let the state variable be defined as , the system output be , and the control input be , then the dynamics model of the multi-link robot arm is represented as ; wherein , .
3. The reinforcement learning based adaptive preset performance control method for a robot arm according to claim 2, wherein, The tracking error variable of the multi-link robot arm is specifically designed as follows: ; wherein, is the tracking error, is the system output, is the desired trajectory.
4. The reinforcement learning based adaptive preset performance control method for a robot arm according to claim 3, characterized in that, Tracking error Setting constraints wherein, is the number of links of the robot arm, is a preset performance control function, are positive parameters designed respectively, is a lower bound proportional coefficient, is an upper bound proportional coefficient, t is a current time, and T is a transpose symbol.
5. The reinforcement learning based adaptive preset performance control method for a robot arm according to claim 4, characterized in that, The equivalent error conversion on the tracking error is specifically as follows: ; wherein, is a conversion function, is a converted tracking error, ; Further, the error variable without constraint is obtained .
6. The reinforcement learning based adaptive preset performance control method for a robot arm according to claim 5, wherein, The compensation tracking error signal based on the backstepping method and the instruction filter is established according to the unconstrained error variable, and includes: The coordinate transformation related to the tracking error is set as follows: ; wherein, is an unconstrained tracking error variable, is a velocity error variable, is an output of the command filter, s is an unconstrained error variable; The compensation tracking error signal is set as follows: ; wherein is a compensation tracking error signal, is a compensation signal.
7. The reinforcement learning based adaptive preset performance control method for a robot arm according to claim 6, characterized in that, According to the following formula: ; The output of the instruction filter is obtained as follows: ; wherein, is the filtered virtual control output for the ith link, is the command filter bandwidth for the ith link, is the auxiliary state variable for the ith link corresponding filter, is the filtered virtual control output initial value for the ith link, is the auxiliary state variable initial value for the ith link corresponding filter, is the damping ratio for the ith link corresponding filter, , is the virtual control signal input to the filter.
8. The reinforcement learning based adaptive preset performance control method for a robot arm according to claim 7, characterized in that, The Hamilton-Jacobi-Bellman equation is established by constructing the performance index related to the compensation tracking error signal, the virtual control signal and the actual control signal, and includes: The Hamilton-Jacobi-Bellman equation is established by constructing the performance index related to the compensation tracking error signal, the virtual control signal and the actual control signal as follows: ; ; wherein, is the optimal cumulative cost of the virtual control subsystem, is the optimal cumulative cost of the actual control subsystem, is the minimum of the admissible control set, is the virtual control signal, is the actual control signal, is the optimal virtual control signal, is the optimal actual control signal, is the instantaneous cost function, , is the optimal instantaneous cost function, is the admissible control set, is the compact set containing the origin, is the integral dummy variable; The Hamilton-Jacobi-Bellman equation is established as follows: ; ; wherein is the Hamilton-Jacobi-Bellman equation for the virtual control subsystem, is the Hamilton-Jacobi-Bellman equation for the actual control subsystem, T is the transpose symbol, is the derivative of the compensation tracking error signal.
9. The reinforcement learning based adaptive preset performance control method for a robot arm according to claim 8, characterized in that, The reinforcement learning method based on the neural network is used to solve the Hamilton-Jacobi-Bellman equation, and the Lyapunov function related to the compensation tracking error signal and the neural network weight estimation error is constructed in combination with the Lyapunov stability theory to determine the virtual control signal, the actual control signal and the parameter adaptive law, and includes: Step 1 calculation : The first continuous auxiliary function is constructed, and the radial basis function neural network is adopted to estimate the continuous auxiliary function, so as to obtain the partial derivative of the performance index with respect to and the optimal virtual control signal. ; ; wherein, is a predetermined performance transformation function, , is a design parameter, is a diagonal matrix, , is an ideal weight vector of the neural network, is a basis function vector of the neural network, is an estimation error of the neural network, is a derivative of the desired trajectory, is a derivative of the predetermined performance control function; With the optimal virtual control signal and the partial derivative of the performance index with respect to as the actor and critic of the reinforcement learning method, specifically: ; ; wherein, is an estimate of the optimal virtual control signal, is an estimate of is an estimate of are weights of the actor and critic neural networks; The condition of constructing the neural network parameter adaptive law by using the Bellman residual error is set by constructing a positive definite function, and the neural network parameter adaptive law is set as follows: ; ; wherein, is the derivative of are the learning rates for the actor and critic neural networks, respectively, is the derivative of Step 2 calculation : The second continuous auxiliary function is constructed, and the radial basis function neural network is adopted to estimate the continuous auxiliary function, so as to obtain the partial derivative of the performance index with respect to and the optimal virtual control signal. ; ; wherein, is a design parameter, f is a nonlinear dynamics equation in the system dynamics model, is a derivative of the command filter output is an ideal weight vector of the neural network, is a basis function vector of the neural network, is an estimation error of the neural network; With the optimal virtual control signal and the partial derivative of the performance index with respect to as the actor and critic of the reinforcement learning method, specifically: ; ; wherein, is an estimate of the optimal virtual control signal, is an estimate of is an estimate of are weights of the actor and critic neural networks, respectively; The condition of constructing the neural network parameter adaptive law by using the Bellman residual error is set by constructing a positive definite function, and the neural network parameter adaptive law is set as follows: ; where is the derivative of is the learning rate of the actor and critic neural networks, respectively, is the derivative of 10. The reinforcement learning based adaptive preset performance control method for a robot arm according to claim 9, characterized in that, A relative threshold event-triggered strategy is used to update the signals transmitted between the actuator and the controller, and specifically as follows: ; wherein, is the designed intermediate control signal, , , , are design parameters, is the actual control signal of the ith link, denotes the time of the kth control input update of the ith link, is the maximum lower bound of the set, is the event-triggered error, .