Robust control method for mechanical arm based on event-triggered reinforcement learning

CN122463188BActive Publication Date: 2026-09-29NANJING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610947202.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-29
Publication Date
2026-09-29
Estimated Expiration
2046-06-29

AI Technical Summary

Technical Problem

但是,现有的联合驱动方法仍存在不足:(1)时序特征提取能力不足,传统的前馈神经网络自适应律通常只利用当前时刻的误差信号,缺乏对历史信息的记忆能力

Benefits of technology

[0011]本发明与现有技术相比,其显著优点是:(1)采用LSTM网络拟合评价与执行机制,克服了传统前馈神经网络在处理长距离时序依赖上的局限性,采用固定隐含层、更新输出层权值的轻量化策略,规避了LSTM耗时的随时间反向传播运算,满足微秒级硬实时约束;(2)设计了基于贝尔曼误差的事件触发机制,有效避免了控制后期的零点反复穿越问题,并大幅降低了计算负担;(3)结合强化学习和鲁棒控制处理系统的不匹配和匹配不确定性,可以获得更好的跟踪性能。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122463188B_ABST
    Figure CN122463188B_ABST
Patent Text Reader

Abstract

The application provides a mechanical arm robust control method based on event triggering reinforcement learning, and comprises the following steps: step S100, a basic robust controller based on a dynamics model of a multi-degree-of-freedom mechanical arm is designed, and a discrete-time error dynamics model is constructed; step S200, a long short-term memory execution-evaluation network is designed to approximate unknown time-varying disturbances of the multi-degree-of-freedom mechanical arm system, and is used for feedforward torque compensation; and step S300, an event triggering mechanism based on a time sequence difference Bellman error is designed, and a gradient descent method is used to update network weights.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to robotic arm servo control technology, and in particular to a robust control method for robotic arms based on event-triggered reinforcement learning. Background Technology

[0002] With the rapid development of industrial automation and intelligent manufacturing, multi-degree-of-freedom robotic arms have been widely used in complex unstructured scenarios such as machinery, medical care, deep-sea operations and aerospace maintenance. At present, robotic arm control is mainly divided into two paths: model-driven and data-driven. The traditional pure model-driven method relies on accurate dynamic parameters, but in actual working conditions, accurate modeling is often difficult to achieve, resulting in limited control accuracy; while the pure data-driven method has a lower dependence on the model, it faces the problems of huge computational load, long training time and difficulty in meeting the requirements of microsecond-level hard real-time control at the underlying level. Therefore, the joint driving method that combines the model for basic control and the data-driven method for disturbance compensation has become a current research hotspot. However, the existing joint driving method still has shortcomings: (1) Insufficient ability to extract time-series features. The traditional feedforward neural network adaptive law usually only uses the error signal at the current moment and lacks the ability to remember historical information. When the system state approaches the equilibrium point and the tracking error approaches zero, the learning mechanism often fails, resulting in poor performance of the controller when dealing with complex disturbances with long-distance time-series dependence features. (2) Computational burden and instability: Online learning algorithms based on deep neural networks usually involve time-consuming backpropagation operations, which are difficult to run in real time in embedded controllers. Continuous parameter updates can easily cause repeated zero crossings in the later stages of control, and even lead to weight divergence, threatening system stability. Summary of the Invention

[0003] The purpose of this invention is to provide a robust control method for a robotic arm based on event-triggered reinforcement learning, comprising the following steps: Step S100: Design a basic robust controller based on the dynamic model of a multi-degree-of-freedom robotic arm and construct a discrete-time error dynamic model; Step S200: Design a long short-term memory execution-evaluation network to approximate the unknown time-varying disturbances of the multi-degree-of-freedom robotic arm system for feedforward torque compensation. Step S300: Design an event triggering mechanism based on temporal difference Bellman error, and design a gradient descent method to update network weights.

[0004] Furthermore, the specific process in step S100 is as follows: Step S110: Establish a dynamic model of the multi-degree-of-freedom manipulator based on the Newton-Euler method, define state variables, and convert the established dynamic model of the multi-degree-of-freedom manipulator into state-space equations. Step S120: Based on the desired trajectory and actual state of the multi-degree-of-freedom robotic arm, design a controller that includes feedforward compensation and robust terms. Step S130: Substitute the designed robot joint driving torque into the continuous physical model of the robotic arm, perform time discretization, and obtain the discretized error dynamic equation of the closed-loop system.

[0005] Furthermore, the specific process of step S110 is as follows: Step S111: Establish the dynamic model of the multi-degree-of-freedom robotic arm's mechanical system based on the Newton-Euler method. (1) in, , , These are the joint position, velocity, and acceleration vectors, respectively. M ( q )for n Inertial matrix of a robot with degrees of freedom It is a matrix that includes centrifugal force and Coriolis force; G ( q () represents the robot's gravity vector; d This represents the actual unknown disturbances in the system. u This represents the driving torque of the robot's joints; Step S112, the robot dynamics satisfy the following properties: Inertia matrix Symmetric positive definite matrix For antisymmetric matrices, Unknown disturbance Bounded; Step S113: Based on the mechanical system dynamics model of the multi-degree-of-freedom robotic arm, select the system state. , , Then the state equation of the system is (2) Furthermore, the specific process of step S120 is as follows: Step S121, for Design virtual control variables based on the given desired position. Position compared to actual sensor feedback Design error variables , , (3) Step S122: Design the virtual control quantity according to formula (4) The virtual control quantity is obtained according to formula (5). error ,make Approaching zero (4) (5) (6) in, It is a positive definite diagonal matrix; Step S123, obtained according to formulas (4), (5), and (6), (7) Step S124: Design the joint driving torque of the system robot. make Approaching zero, the robot joint driving torque The structure is as follows: (8) in, This is the positive gain matrix of the robotic arm controller. It is a feedforward compensation control term designed based on the dynamic model of the robotic arm. For robust control terms, It is the first The reinforcement learning perturbation estimation term for each discrete control cycle.

[0006] Furthermore, the specific process of step S130 is as follows: Step S131: Establish the continuous-time closed-loop system state tracking error. The dynamic change model, (9) in, Represents the dynamic nominal error under the action of the basic feedback controller. It is the input gain matrix. The deviation between the ideal control input and the actual controller output; Step S132: Substitute the designed execution network control input and the actual system disturbance into the dynamic change model of the closed-loop system state tracking error. The actual unknown disturbance and unmodeled dynamics are completely represented by the ideal optimal network as follows: (10) Step S133: Obtain the actual controller output. (11) Step S134, by Obtain control deviation (12) Step S135 yields the continuous-time closed-loop error dynamic equation. (13) Step S136, Discretize (14) (15) in, z k For the present k The system augmented state tracking error at time t is z k+1 For the next moment k The system augmented state tracking error at time +1 T The system's discretization sampling time, To execute neural networks in k The weight estimation error at time step, h a,k To execute the temporal feature vector extracted by the hidden layer at time k in the neural network. This refers to the inherent network approximation residual generated when the ideal optimal execution neural network approximates real time-varying perturbations. u s The robust compensation control term output by the basic robust controller, I It is an identity matrix with the same dimension as the system state matrix; Step S137 yields the discretized error dynamic equations of the closed-loop system. (16) in, This represents the discretized system state matrix. It is a Herwitz matrix. This represents the discretized input matrix.

[0007] Furthermore, step S200 specifically includes: Step S210: Design an evaluation network to assess the control performance of the system based on its current operating state and fit a value function in the infinite time domain; Step S220: Design the execution network to output estimated compensation values ​​for real unknown disturbances and unmodeled dynamics based on the current operating state of the system. Both the evaluation network and the execution network use a Long Short-Term Memory (LSTM) neural network structure as hidden layers.

[0008] Furthermore, step S210 specifically includes: Step S211, define the augmented system state error as... The evaluation feature vector is obtained after processing by the hidden layer of LSTM. ,in The number of neurons in the hidden layer; Step S212: Design the instantaneous cost function for penalizing tracking error and robust control output. (17) In the formula, and These are the weight matrices for positive semi-definite and positive definite weights, respectively. Step S213, define the true infinite time domain value function. The sum of the instantaneous costs of all future discounts. r ( k Let ) be the instantaneous cost function. The discount factor is introduced; an evaluation network is introduced to approximate the true infinite-time-domain value function online; the output of the evaluation network is defined as an estimate of the infinite-time-domain value function. (18) in, To evaluate the weights of the network output layer, This is the temporal feature vector extracted from the hidden layer of the LSTM.

[0009] Further, step S220 specifically includes: defining the execution network input as... The evaluation feature vector is obtained after processing by the hidden layer of LSTM. ; Execute network output perturbation estimates: (19) in, To execute neural networks in k The estimated weights at time 1.

[0010] Furthermore, the specific process of step S300 is as follows: Step S310: Design an event triggering mechanism based on timing differential Bellman error; Step S320: Design the gradient descent method to evaluate the network weight update law; Step S330: Design the execution network weight ratio based on the Lyapunov function; Step S310 specifically includes: Step S310, based on the Bellman optimality principle, the time-difference Bellman error is defined as: (20) Step S312, define the event trigger indicator function. ,when hour, Trigger network update; otherwise Maintain the weights; ec,k For the first k The temporal difference Bellman error at time step 1. Δ T The initial constant weights for the dynamic threshold. ρ It is an exponential decay factor, 0 < ρ <1, δ The steady-state dead zone threshold; Step S320 specifically includes: Step S321: Minimize the temporal difference Bellman error, defining the instantaneous target cost function at time k. , , (twenty one) Step S322: Use gradient descent to calculate the partial derivatives of the output layer weights. ;(twenty two) Step S323, the updated law of the obtained weights is: ,(twenty three) in, α c To evaluate the network learning rate; Step S330 specifically includes: Step S331, define the Lyapunov function of the system state as follows: Find the difference between them. (25) in, V z,k The discrete-time Lyapunov function constructed for the system at the k-th discrete time. P Represents a symmetric positive definite matrix; expansion yields an error term approximated by the weights of the execution network. , Based on this, an equal negative term is generated by designing the weight update law of the execution network, thus canceling it out; Step S332, define the Lyapunov function for executing network weights as follows: Find the difference and ignore higher-order minima. (26) in, V a,k The discrete-time Lyapunov function representing the error in performing network weight estimation. represent k Time's up k The increment of neural network weights at time +1. It is a symmetric positive definite matrix; tr(*) represents the trace of a matrix in linear algebra; Step S333, let The dominant term generated and The unstable cross terms are opposites of each other, that is... ,have to (27) Step S334, when the system tends to steady state, the introduced... Modifications , for Modify the constant gain of the item and set the event trigger indicator function. It is introduced into the update law as a switch multiplier. (28) Furthermore, in step S320, a normalized least mean square is introduced, which divides the learning rate of the evaluation network by the square of the norm. Add a very small constant 1 to the event trigger indicator function. Introduced as a switch multiplier into the update law; ultimately, the network weight update rate is obtained. .

[0011] Compared with the prior art, the significant advantages of this invention are: (1) It adopts the LSTM network fitting evaluation and execution mechanism, which overcomes the limitations of traditional feedforward neural networks in handling long-distance temporal dependencies. It adopts a lightweight strategy of fixing the hidden layers and updating the output layer weights, which avoids the time-consuming backpropagation operation of LSTM and meets the microsecond-level hard real-time constraints; (2) It designs an event triggering mechanism based on Bellman error, which effectively avoids the problem of repeated zero crossings in the later stage of control and greatly reduces the computational burden; (3) By combining reinforcement learning and robust control to handle the mismatch and matching uncertainty of the system, better tracking performance can be obtained. Attached Figure Description

[0012] Figure 1 This is a schematic diagram of the method of the present invention.

[0013] Figure 2 This is a schematic diagram of the circular trajectory tracking test of the multi-degree-of-freedom robotic arm of the present invention.

[0014] Figure 3 The image shows the task space tracking effect of the robust controller for the robotic arm based on event-triggered long short-term memory neural network reinforcement learning designed in this invention.

[0015] Figure 4 This is an estimated interference waveform diagram under the action of the controller described in this invention.

[0016] Figure 5 This is a diagram of joint control torque under the action of the controller described in this invention.

[0017] Figure 6 This is a trajectory tracking error diagram of multiple joints under the action of the controller described in this invention. Detailed Implementation

[0018] Combination Figure 1 A robust control method for a robotic arm based on event-triggered long short-term memory neural network reinforcement learning includes the following steps: Step S100: Design a basic robust controller based on the dynamic model of a multi-degree-of-freedom robotic arm and construct a discrete-time error dynamic model; Step S200: Design a long short-term memory execution-evaluation network to approximate the unknown time-varying disturbances of the multi-degree-of-freedom robotic arm system for feedforward torque compensation. Step S300: Design an event triggering mechanism based on temporal difference Bellman error, and design a gradient descent method to update network weights; In step S400, the stability of the robust controller for the robotic arm based on event-triggered long short-term memory neural network reinforcement learning is analyzed using Lyapunov stability theory, and the system is found to be bounded stable.

[0019] Step S100 specifically includes the following process: Step S110: Based on the dynamic model of the multi-degree-of-freedom manipulator established by Newton-Euler method, define state variables and convert the established dynamic model of the multi-degree-of-freedom manipulator into state-space equations. Step S120: Based on the desired trajectory and actual state of the multi-degree-of-freedom robotic arm, design a controller that includes feedforward compensation and robust term to ensure the global stability of the system. Step S130: Substitute the designed robot joint driving torque into the continuous physical model of the robotic arm, and discretize it in time to obtain the discrete-time error dynamic model of the closed-loop system.

[0020] Step S110: Based on the dynamic model of the multi-degree-of-freedom manipulator established by Newton-Euler method, define state variables and convert the established dynamic model of the multi-degree-of-freedom manipulator into state-space equations. The specific process includes steps S111 to S112.

[0021] Step S111: Based on the Newton-Euler method, the dynamic model of the multi-degree-of-freedom robotic arm's mechanical system is described as follows: (1) In the formula These are the joint position, velocity, and acceleration vectors, respectively. for Inertial matrix of a robot with degrees of freedom; It is a matrix that includes centrifugal force and Coriolis force; This is the robot's gravity vector; This is an unmodeled external force vector that includes external disturbances; This represents the driving torque of the robot's joints; Assumption: Robot dynamics satisfy the following properties (1) Inertia matrix Symmetric positive definite; (2) Matrix It is an antisymmetric matrix; (3) Unknown disturbance Bounded, that is .

[0022] Step S112: Select the system state according to the dynamic model (1). , Then the state equation of the system can be described as (2) Step S120: Based on the desired trajectory and actual state of the multi-degree-of-freedom robotic arm, design a controller that includes feedforward compensation and robust term to ensure the global stability of the system. The specific process includes steps S121 to S124.

[0023] Step S121, for Design virtual control variables based on the given desired position. Position compared to actual sensor feedback Design error variables , , (3) Step S122: Design the virtual control quantity according to formula (4) The virtual control quantity is obtained according to formula (5). error ,make Approaching zero (4) (5) (6) in, It is a positive definite diagonal matrix.

[0024] Step S123, according to formulas (4), (5), and (6), can be obtained as follows: (7) Step S124, design the system robot joint drive torque input make Approaching zero, the robot joint driving torque The structure is as follows: (8) In the formula, This is the positive gain matrix of the robotic arm controller. It is a feedforward compensation control term designed based on the dynamic model of the robotic arm. For robust control terms, It is the first The reinforcement learning perturbation estimation term for each discrete control cycle.

[0025] Step S130: Substitute the designed robot joint driving torque into the continuous physical model of the robotic arm and discretize it in time to obtain the discrete-time error dynamic model of the closed-loop system. The specific process includes steps S131 to S136.

[0026] Step S131: Establish the continuous-time closed-loop system state tracking error. The dynamic change model, (9) in, This represents the nominal error dynamics under the action of the basic feedback controller; if the basic controller is properly designed, the matrix... It is a Herwitz matrix, meaning the system itself is asymptotically stable; It is the input gain matrix; This represents the deviation between the ideal control input and the actual controller output.

[0027] Step S132: Substitute the designed execution network control input and the actual system disturbance into the dynamic change model of the closed-loop system state tracking error. Here, the actual unknown disturbance and unmodeled dynamics of the system can be assumed to be completely approximated by an ideal optimal neural network, expressed as: (10) in, d This represents the actual unknown disturbances in the system. W a * To achieve the ideal optimal weights for the neural network, h a The feature vector is obtained after processing by the hidden layer of the neural network. To calculate the approximation error of the network.

[0028] Step S133, the actual controller output is: (11) Step S134, since the ideal control input should completely cancel out the real disturbance, by Obtain control deviation, (12) Step S135 yields the following equation: the continuous-time closed-loop error dynamics equation becomes... (13) Step S136: Discretize the data. (14) (15) in, z k For the present k The system augmented state tracking error at time t is z k+1 For the next moment k The system augmented state tracking error at time +1 T The system's discretization sampling time, To execute neural networks in k The weight estimation error at time step, h a,k To execute the temporal feature vector extracted by the hidden layer at time k in the neural network. This refers to the inherent network approximation residual generated when the ideal optimal execution neural network approximates real time-varying perturbations. u s The robust compensation control term output by the basic robust controller, I It is an identity matrix with the same dimension as the system state matrix.

[0029] Step S137 yields the discretized error dynamic equations of the closed-loop system. (16) in, This represents the discretized system state matrix. This represents the discretized input matrix.

[0030] The specific process of step S200 includes: Step S210: Design an evaluation network to assess the control performance of the system based on its current operating state and fit a value function in the infinite time domain; Step S220: Design the execution network to output estimated compensation values ​​for real unknown disturbances and unmodeled dynamics based on the current operating state of the system. .

[0031] Step S210, evaluate network design. The role of the evaluation network is to evaluate the control performance and fit the value function in the infinite time domain based on the current operating state of the system. The evaluation network adopts a long short-term memory neural network (LSTM). In order to adapt to the hard real-time constraints of the multi-degree-of-freedom manipulator and ensure the stability of the closed-loop physical system, the following two adjustments were made to the LSTM in this embodiment: (1) Linearization adjustment of the output layer structure: Standard LSTM usually has a complex nonlinear output layer, but in order to facilitate the rigorous Lyapunov stability proof and realize online real-time adaptation, the output layer of the evaluation network and the execution network is designed as a simple linear structure. (2) Lightweight adjustment of the learning mechanism, that is, fixed hidden layer and only updated output layer: The training of standard LSTM requires the use of an extremely time-consuming backpropagation algorithm, which is difficult to complete within a microsecond-level control cycle; In this embodiment, the network structure is adjusted, and the complex gating unit of LSTM is used only as a historical feature extractor, that is, the hidden layer weights are fixed, and only the above-mentioned linear output layer weights are updated online during the control process.

[0032] The specific design process of step S210 includes steps S211 to S213.

[0033] Step S211, define the augmented system state error as... The evaluation feature vector is obtained after processing by the hidden layer of LSTM. ,in This represents the number of neurons in the hidden layer.

[0034] Step S212: Considering that the actual disturbance is unmeasurable, in order to comprehensively reflect the tracking accuracy and energy consumption of the multi-degree-of-freedom robotic arm, an instantaneous cost function for penalizing tracking error and robust control output is designed: (17) In the formula, and These are the weight matrices for positive semi-definite and positive definite values, respectively.

[0035] Step S213, Instantaneous cost function It can only reflect the local performance of the system at a single moment. In order to achieve suboptimal global control in the long run, a true infinite time-domain value function is defined. The sum of the instantaneous costs of all future discounts. r ( k Let ) be the instantaneous cost function. As a discount factor; due to the unavailable future information in the Bellman equation, the true function Since it is difficult to solve directly, an evaluation network is introduced to approximate it online; the output of the evaluation network is defined as an estimate of the infinite time-domain value function: (18) in, To evaluate the weights of the network output layer, This is the temporal feature vector extracted from the hidden layer of the LSTM.

[0036] Step S220: Perform network design. The task of the network is to output estimated compensation values ​​for real unknown disturbances and unmodeled dynamics based on the current state of the system. Define the network input for execution as... That is, the system augmented state error vector, which is processed by the LSTM hidden layer to obtain the evaluation feature vector. Define the output disturbance estimate of the execution network: (19) Step S300: Design an event-triggered mechanism based on temporal difference Bellman error, and design a gradient descent method to update network weights, specifically including: Step S310: Design an event triggering mechanism based on timing differential Bellman error; Step S320: Design the gradient descent method to evaluate the network weight update law; Step S330: Design the execution network weight ratio based on the Lyapunov function.

[0037] Step S310: Design an event triggering mechanism based on timing differential Bellman error, specifically including steps S311 to S312.

[0038] Step S311, the infinite time-domain value function of the system is According to the Bellman optimality principle, the time-difference Bellman error is defined as: (20) Step S312, define the event trigger indicator function. ,when hour, Trigger network update; otherwise Maintain the weights; e c,k For the first k The temporal difference Bellman error at time step 1. Δ T The initial constant weights for the dynamic threshold. ρ It is an exponential decay factor, 0 < ρ <1, δ This is the steady-state dead zone threshold.

[0039] Step S320: Design the gradient descent method to evaluate the network weights. Using a lightweight strategy of fixing hidden layers and updating output layer weights, and employing normalized gradient descent, design the weight update law for evaluating the network. Specifically, this includes: Step S321: In order to train the evaluation network to accurately estimate the infinite time-domain value function and minimize the temporal difference Bellman error, the instantaneous target cost function at time k is defined. , , (twenty one) Step S322, minimizing the objective cost function Gradient descent is used to calculate the partial derivatives of the output layer weights. ;(twenty two) Step S323 yields the following update law for the weights: ,(twenty three) in, α c To evaluate the network learning rate; Step S324: To prevent the feature vector output by the LSTM hidden layer from being damaged by sudden shocks to the robotic arm. Excessively large magnitudes can lead to gradient explosion. To address this, we introduce normalized least mean square, which divides the learning rate by the energy of the input features, i.e., the norm squared. To prevent the denominator from being zero, a very small constant of 1 is added. To reduce the computational burden, the event trigger indicator function is... Introduced as a switch multiplier into the update law; ultimately, the network weight update rate is obtained. .(twenty four) Step S330, designing the execution network weight ratio based on the Lyapunov function, specifically includes: Step S331, define the Lyapunov function of the system state as follows: Find the difference between them. (25) in, V z,k The discrete-time Lyapunov function constructed for the system at the k-th discrete time. P It represents a symmetric positive definite matrix; The expansion yields a term including the approximation error of the execution network weights. , Based on this, an equal negative term is generated by designing the weight update law of the execution network, thus canceling it out; Step S332, define the Lyapunov function for executing network weights as follows: Find the difference and ignore higher-order minima. (26) in, Va,k The discrete-time Lyapunov function representing the error in performing network weight estimation. It is a symmetric positive definite matrix; tr(*) represents the trace of a matrix in linear algebra; Step S333: To make the total Lyapunov difference less than or equal to 0, let The dominant term generated and The unstable cross terms are opposites of each other, that is... ,have to (27) In step S334, when the system approaches steady state, the hidden layer input loses its continuous excitation condition. At this point, any tiny noise may be continuously integrated by the basic update law, leading to... Infinite drift divergence; introduced Modifications , for The constant gain of the modification item, and trigger the event indicator function. It is introduced into the update law as a switch multiplier. (28) Step S400: Lyapunov stability theory is applied to analyze the stability of the robust controller for the robotic arm based on event-triggered long short-term memory neural network reinforcement learning, yielding a result indicating that the system is boundedly stable. To prove the stability of the entire control system and the LSTM neural network weights, the parameter estimation error is defined as... and , The ideal optimal weights are determined by these values.

[0040] Theorem 1: Considering the dynamics system of a robotic arm, using the above event-triggered LSTM reinforcement learning control law and weight update law, if the learning rate... If the condition of a sufficiently small positive number is satisfied, then the state of the closed-loop system is... and neural network weight estimation error It is ultimately uniformly bounded.

[0041] Proof: Construct the discrete-time Lyapunov function as follows: (27) In the formula, V c,k This represents the Lyapunov function part used to evaluate the error in network weight estimation. It is a symmetric positive definite matrix.

[0042] Consider the first-order forward difference of the Lyapunov function .

[0043] Scenario 1: Triggering a network update (28) in, It is a lumped disturbance term.

[0044] Based on the discrete Lyapunov equations satisfied during the design of the basic controller. ,in It is a symmetric positive definite matrix. According to the Rayleigh quotient property of matrices: (29) According to the Cauchy-Schwarz inequality and the product property of matrix norms

[0045] (30) Using Young's inequality to handle the remainder term .

[0046] for According to Young's inequality (in ), , It can be integrated into the first item. middle, It is bounded, therefore It is a normal number.

[0047] for ,have lumped disturbance Bounded, It is a bounded positive number.

[0048] All the bounded positive constants generated by the approximation error, robust control terms, and Young's inequality scaling are summed together and collectively referred to as constants. .

[0049] (31) For execution networks, the weight update law includes the learning rate. Modifications to prevent weight divergence .because Expand : (32) Using the linearity property of the matrix trace and positive definite matrix The symmetry brought about ,get: (33) The first item expands to: (34) Using the cyclic property of matrix trace , .

[0050] Using the properties of matrix trace

[0051] (35) in, , It is a bounded constant, and it is related to the higher-order infinitesimal terms generated in the first step. Combined into a bounded constant The final result is: (36) Similarly, for those with a learning rate According to the universal approximation theorem, the evaluation network has an ideal optimal weight. This makes the ideal optimal Bellman equation satisfy: (37) In the formula, The inherent Bellman residuals resulting from the optimal network approximation, and which satisfy boundedness. Actual timing difference error Subtract this from the optimal Bellman equation and substitute in the weight estimation error. We can obtain: (38) (39) For the first item ,Will Substitute: (40) For the second item By utilizing the positive definite property and the defining property of the normalized denominator ,get : (41) We use Young's inequality to safely shrink the final cross term because of the discount factor. Hidden layer output and approximation error Both are bounded, and the cross term can be absorbing by scaling. And a containing And the bounded constant part that approximates the error boundary.

[0052] (42) In the formula For relevant positive numbers, It is a bounded constant.

[0053] (43) in, This is the sum of all approximation errors and higher-order bounded terms. This represents the cross term between the system state error and the network weight error generated during the derivation process.

[0054] By absorbing cross terms using Young's inequality, due to the physical system parameters and hidden layer output... It is bounded, encompassing all things related to and The cross terms are uniformly enlarged and defined as (in (A bounded positive constant of the whole).

[0055] (44) In the formula, This is a constant that can be freely chosen.

[0056] Substituting the bounded cross terms back into the overall difference equation and extracting the common terms, we obtain the final bounded inequality: (45) To ensure system stability, all coefficients on the right-hand side of the inequality must be positive. Definition (46) By selecting an appropriate base controller gain matrix and adjustment constant It can guarantee .

[0057] When the state error of the closed-loop system Or the neural network weight estimation error exceeds the constant term When the set is defined as compact, there always exists According to the discrete-time Lyapunov stability theorem, all signals in this closed-loop system (including system state errors) Execution network weight error And evaluate the network weight error All of these are eventually consistent. The closed-loop tracking error will eventually converge to a small neighborhood near the origin.

[0058] Scenario 2: Network update not triggered At this point, the system error is small, so we keep the network weights unchanged. .

[0059] (47) Using Young's Inequality : (48) Since no event was triggered at this time, it indicates that the system's current TD Bellman error is limited to a very small dead zone threshold, meaning that the network's current output plus the basic robust control term is within acceptable limits. It can effectively suppress the remaining lumped disturbances within the boundary. Therefore, the sum of the last two terms on the right-hand side of the above equation is a bounded positive constant, which is denoted as . .

[0060] (49) Therefore, as long as a suitable basic feedback gain is selected, Then when the system state error , there must be This ensures the eventual consistency and boundedness of the closed-loop system.

[0061] This embodiment uses the Matlab / Simulink Simscape-Multibody simulation platform to build a seven-DOF Franka robotic arm platform to verify the effectiveness of the proposed control strategy. The end effector of the multi-DOF robotic arm is required to track a given desired circular trajectory in the task space. At the same time, unmodeled friction and time-varying external unknown disturbances are added to the system to simulate the real working environment under complex unstructured scenarios.

[0062] like Figure 2 As shown, the desired trajectory in the task space is set to a circle:

[0063]

[0064] In the formula: The desired position in the task space of the robotic arm. The desired posture of the robotic arm.

[0065] The following controller is used for comparison in the simulation: A robust controller for a robotic arm based on reinforcement learning (Controller 1) is provided, with the following controller parameters:

[0066] Robust controller (Controller 2), the controller parameters are as follows:

[0067] The tracking effect of the robotic arm's task space trajectory under the action of Controller 1 is as follows: Figure 3 As shown, even with unmodeled dynamics and time-varying disturbances, the actual motion trajectory of the robotic arm's end effector (represented by the solid line in the figure) still highly coincides with the desired circular trajectory (represented by the dashed line in the figure, obscured by the solid line). This indicates that the strategy of combining basic robust control with reinforcement learning feedforward compensation significantly improves the macroscopic trajectory tracking accuracy of the system. Estimated disturbances under the action of Controller 1, such as Figure 4 As shown in the figure, the online approximation process of the evaluation-execution network for unknown time-varying perturbations is illustrated. Long Short-Term Memory (LSTM) networks possess powerful historical time-series feature extraction capabilities, and the perturbation estimates output by the network (represented by solid lines in the figure) can quickly and accurately track the actual injected time-varying perturbation signal (represented by dashed lines in the figure). It is proven that even with a lightweight strategy of fixing the hidden layers and updating the output layer weights, the network still possesses high-precision estimation capabilities under microsecond-level hard real-time constraints. The control torque of the seven joints under the action of Controller 1 is as follows... Figure 5 As shown, the output joint drive torque is smooth and has no obvious high-frequency jitter, which can effectively protect the mechanical structure of the robotic arm and reduce the computational burden on the controller.

[0068] The comparison chart of tracking errors of the seven joints under the action of Controller 1 and Controller 2 is shown below. Figure 6 As shown, the controller proposed in this invention (shown by the solid line in the figure) exhibits a significant improvement in control performance compared to the traditional robust controller (shown by the dashed line in the figure), demonstrating excellent tracking performance. Throughout the later stages of control, no repeated zero-point crossings or error divergence due to overfitting or insufficient continuous excitation occurred, verifying the effectiveness of the proposed control method. Experimentally, this also verifies the conclusion of uniformly bounded system derived in step 4 using the discrete-time Lyapunov theorem.

Claims

1. A robust control method for a robotic arm based on event-triggered reinforcement learning, characterized in that, Includes the following steps: Step S100: Design a basic robust controller based on the dynamic model of a multi-degree-of-freedom robotic arm and construct a discrete-time error dynamic model; Step S200: Design a long short-term memory execution-evaluation network to approximate the unknown time-varying disturbances of the multi-degree-of-freedom robotic arm system for feedforward torque compensation. Step S300: Design an event triggering mechanism based on temporal difference Bellman error, and design a gradient descent method to update network weights. The specific process in step S100 is as follows: Step S110: Establish a dynamic model of the multi-degree-of-freedom manipulator based on the Newton-Euler method, define state variables, and convert the established dynamic model of the multi-degree-of-freedom manipulator into state-space equations. Step S120: Based on the desired trajectory and actual state of the multi-degree-of-freedom robotic arm, design a controller that includes feedforward compensation and robust terms. Step S130: Substitute the designed robot joint driving torque into the continuous physical model of the robotic arm, perform time discretization, and obtain the discretized error dynamic equation of the closed-loop system. The specific process of step S110 is as follows: Step S111: Establish the dynamic model of the multi-degree-of-freedom robotic arm's mechanical system based on the Newton-Euler method. ,(1) in, , , These are the joint position, velocity, and acceleration vectors, respectively. M ( q )for n Inertial matrix of a robot with degrees of freedom It is a matrix that includes centrifugal force and Coriolis force; G ( q () represents the robot's gravity vector; d This represents the actual unknown disturbances in the system. u This represents the driving torque of the robot's joints; Step S112, the robot dynamics satisfy the following properties: Inertia matrix Symmetric positive definite matrix For antisymmetric matrices, Unknown disturbance Bounded; Step S113: Based on the mechanical system dynamics model of the multi-degree-of-freedom robotic arm, select the system state. , , Then the state equation of the system is (2); The specific process of step S120 is as follows: Step S121, for Design virtual control variables based on the given desired position. Position compared to actual sensor feedback Design error variables , , ;(3) Step S122: Design the virtual control quantity according to formula (4) The virtual control quantity is obtained according to formula (5). error ,make Approaching zero ,(4) ,(5) ,(6) in, This is the positive gain matrix of the robotic arm controller; Step S123, according to formulas (4), (5), and (6), we obtain... ;(7) Step S124: Design the robot joint drive torque make Approaching zero, the robot joint driving torque The structure is as follows: ,(8) in, k 1. k 2 represents the positive gain matrix of the robotic arm controller. It is a feedforward compensation control term designed based on the dynamic model of the robotic arm. For robust control terms, It is the first Reinforcement learning perturbation estimation term for each discrete control cycle; The specific process of step S130 is as follows: Step S131: Establish the continuous-time closed-loop system state tracking error. The dynamic change model, ,(9) in, Represents the dynamic nominal error under the action of the basic feedback controller. It is the input gain matrix. For ideal control input τ ideal With actual controller output τ actual Deviation between; Step S132: Substitute the designed execution network control input and the actual system disturbance into the dynamic change model of the closed-loop system state tracking error. The actual unknown disturbance and unmodeled dynamics are completely represented by the ideal optimal network as follows: ;(10) in, d This represents the actual unknown disturbances in the system. W a To perform weighting of the neural network, W a * To achieve the ideal optimal weights for the neural network, h a The feature vector is obtained after processing by the hidden layer of the neural network. To calculate the approximation error of the network; Step S133: Obtain the actual controller output. ,(11) in For W a Estimated value; Step S134, by Obtain control deviation ,(12) in, ; Step S135 yields the continuous-time closed-loop error dynamic equation. ;(13) Step S136, Discretize ,(14) ,(15) in, z k For the present k The system augmented state tracking error at time t is z k+1 For the next moment k The system augmented state tracking error at time +1 T The system's discretization sampling time, To execute neural networks in k The weight estimation error at time step, h a,k To execute the hidden layers in the neural network k The temporal feature vector extracted at each time step. This refers to the inherent network approximation residual generated when the ideal optimal execution neural network approximates real time-varying perturbations. I It is an identity matrix with the same dimension as the system state matrix; Step S137 yields the discretized error dynamic equations of the closed-loop system. ,(16) in, This represents the discretized system state matrix. It is a Herwitz matrix. This represents the discretized input matrix.

2. The method according to claim 1, characterized in that, Step S200 specifically includes: Step S210: Design an evaluation network to assess the control performance of the system based on its current operating state and fit a value function in the infinite time domain; Step S220: Design the execution network to output estimated compensation values ​​for real unknown disturbances and unmodeled dynamics based on the current operating state of the system. Both the evaluation network and the execution network use a Long Short-Term Memory (LSTM) neural network structure as hidden layers.

3. The method according to claim 2, characterized in that, Step S210 specifically includes: Step S211, define the augmented system state error Z k The evaluation feature vector is obtained after processing by the hidden layer of LSTM. ,in The number of neurons in the hidden layer; Step S212: Design the instantaneous cost function for penalizing tracking error and robust control output. ,(17) In the formula, and These are the weight matrices for positive semi-definite and positive definite weights, respectively. Step S213, define the true infinite time domain value function. This is the sum of the instantaneous costs of all future discounts. r ( k ) is the instantaneous cost function. The discount factor is introduced; an evaluation network is introduced to approximate the true infinite-time-domain value function online; the output of the evaluation network is defined as an estimate of the infinite-time-domain value function. ,(18) in, To evaluate the weights of the network output layer, This is the temporal feature vector extracted from the hidden layer of the LSTM.

4. The method according to claim 3, characterized in that, Step S220 specifically includes: obtaining the evaluation feature vector after processing by the LSTM hidden layer. ; Execution network output disturbance estimate: ,(19) in, To execute neural networks in k The estimated weights at time 1.

5. The method according to claim 4, characterized in that, The specific process of step S300 is as follows: Step S310: Design an event triggering mechanism based on timing differential Bellman error; Step S320: Design the gradient descent method to evaluate the network weight update law; Step S330: Design the execution network weight ratio based on the Lyapunov function; Step S310 specifically includes: Step S310, based on the Bellman optimality principle, the time-difference Bellman error is defined as: ,(20) Step S312, define the event trigger indicator function. ,when hour, Trigger network update; otherwise Maintain the weights; e c,k For the first k The temporal difference Bellman error at time step 1. Δ T The initial constant weights for the dynamic threshold. ρ It is an exponential decay factor, 0 < ρ <1, δ The steady-state dead zone threshold; Step S320 specifically includes: Step S321, minimize the timing difference Bellman error, defined in k Instantaneous target cost function at time 1 , , (21) Step S322: Use gradient descent to calculate the partial derivatives of the output layer weights. ;(22) Step S323, the updated law of the obtained weights is: ,(23) in, α c To evaluate the network learning rate; Step S330 specifically includes: Step S331, define the Lyapunov function of the system state as follows: Find the difference between them. ,(25) in, V z,k For the system in the first k The discrete-time Lyapunov function constructed at discrete moments. P Represents a symmetric positive definite matrix; expansion yields an error term approximated by the weights of the execution network. , Based on this, an equal negative term is generated by designing the weight update law of the execution network, thus canceling it out; Step S332, define the Lyapunov function for executing network weights as follows: Find the difference and ignore higher-order minima. ,(26) in, V a,k The discrete-time Lyapunov function representing the error in performing network weight estimation. represent k Time's up k The increment of neural network weights at time +1. It is a symmetric positive definite matrix; tr(*) represents the trace of a matrix in linear algebra; Step S333, let The dominant term generated and The unstable cross terms are opposites of each other, that is... ,have to ,(27) Step S334, when the system tends to steady state, the introduced... Modifications , for Modify the constant gain of the item and set the event trigger indicator function. It is introduced into the update law as a switch multiplier. (28)。 6. The method according to claim 5, characterized in that, In step S320, a normalized least mean square is introduced, which divides the learning rate of the evaluation network by the square of the norm. Add a very small constant 1 to the event trigger indicator function. Introduced as a switch multiplier into the update law; ultimately, the network weight update rate is obtained. 。

Citation Information

Patent Citations

  • Event triggering decentralized optimal fault-tolerant control method and system for reconfigurable mechanical arm

    CN115890650A

  • Self-adaptive evaluation control method for double-connecting-rod mechanical arm based on event triggering

    CN116604570A