Radial basis function network based iterative learning control method for robot arm
By adopting an iterative learning control method based on radial basis function networks, the problem of insufficient trajectory tracking accuracy of robotic arms under nonlinear systems and external disturbances is solved, achieving high-precision trajectory tracking and improving system stability, thus adapting to complex dynamic environments.
Patent Information
- Application Number
- CN202510201379.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-02-24
AI Technical Summary
Traditional iterative learning control methods suffer from insufficient trajectory tracking accuracy and robustness when dealing with nonlinear systems and external disturbances in robotic arms. They are particularly difficult to adapt effectively to unknown nonlinear characteristics and external disturbances in complex dynamic environments.
An iterative learning control method based on radial basis function networks is adopted. By constructing a dynamic model of the robotic arm system, designing the radial basis function network and dynamic learning gain, and combining feedback control and nonlinear compensation, the trajectory error is optimized. The stability and error convergence of the system are analyzed using Lyapunov stability theory.
It achieves high precision and rapid error convergence in robotic arm trajectory tracking, enhances the system's adaptability and robustness in complex dynamic environments, and improves trajectory tracking accuracy and control system stability.
Smart Images

Figure CN119795191B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of mechanical arm control and nonlinear control technology, and particularly relates to a mechanical arm iterative learning control method based on a radial basis function network. BACKGROUND
[0002] Efficient control of a mechanical arm is one of the core research directions in the field of automation, and the key is to achieve high-precision trajectory tracking. In this context, iterative learning control is widely used in repetitive task scenarios because it can dynamically adjust and gradually optimize control performance using historical data. However, traditional iterative learning control methods have limited performance when dealing with nonlinear systems, external disturbances, and measurement noise. In addition, for unknown nonlinear characteristics in complex dynamic systems, existing methods often lack effective modeling and control means, limiting their robustness and precision in practical applications.
[0003] Among them, the trajectory alignment problem is one of the key challenges for a mechanical arm to achieve high-precision control. Existing technologies usually rely on strictly assumed stochastic modeling methods (such as Markov chains) to describe trajectory changes and alignment conditions. However, these methods not only rely on the assumption that the probability distribution is known, but also are difficult to meet in practical applications. On the other hand, the rule design of fixed or decreasing learning gain in traditional iterative learning control methods also has limitations. Although simple to implement, this rule lacks the ability to adapt to dynamic changes in the system when faced with complex nonlinearities and external disturbances, which can lead to low learning efficiency and even cause the system to diverge, further affecting the robustness and stability of iterative learning control.
[0004] Therefore, it is necessary to study new control strategies for the trajectory alignment problem in mechanical arm systems and the limitations of traditional methods to improve the trajectory tracking accuracy and robustness of nonlinear systems under external disturbances, thereby meeting the application requirements in complex scenarios. SUMMARY
[0005] To solve the above problems in the prior art, the present application proposes a mechanical arm iterative learning control method based on a radial basis function network. Compared with the prior art, this method fully considers the influence of nonlinear dynamic behavior, external disturbances, and unmodeled dynamics of the system, aiming to overcome the limitations of traditional control methods in complex dynamic environments in terms of response performance, and to improve trajectory tracking accuracy and system adaptability.
[0006] To achieve the above technical purposes, the present application provides the following technical solutions:
[0007] The mechanical arm iterative learning control method based on a radial basis function network specifically includes the following steps:
[0008] S1, a mechanical arm system dynamics model is constructed, and the dynamics characteristics of the mechanical arm system and the applicable assumption conditions are determined;
[0009] S2, combined with the initial trajectory data of the mechanical arm, a dynamic correction strategy is adopted to optimize the reference trajectory of the mechanical arm multiple times, and the trajectory error is defined; at the same time, the starting and target states of the reference trajectory are ensured to meet the task requirements, and the smoothness and reachability of the optimized trajectory path are ensured;
[0010] S3, a radial basis function (RBF) network is designed, the center position and width parameters of the basis function are selected, and a nonlinear compensation term suitable for dynamic characteristics is constructed to compensate for the influence of unmodeled dynamics and external disturbances on the mechanical arm system;
[0011] S4, the estimated dynamic weight parameter and dynamic learning gain are designed, the response speed and output smoothness of the controller in the iteration process are optimized, the oscillation problem of the control input is reduced, and the progressive convergence of the weight error is ensured;
[0012] S5, combined with feedback control, nonlinear compensation and iterative learning method, an iterative learning controller is designed to fully utilize the historical data, adjust the control law in real time, gradually reduce the trajectory error, and enhance the adaptability of the control system;
[0013] S6, based on Lyapunov stability theory and composite energy function (CEF), the stability of the control system in complex dynamic environment is analyzed, and the global asymptotic stability of the controller and the gradual convergence of the trajectory error of the mechanical arm are tested.
[0014] Through the above technical solutions, the method proposed by the application can face the iterative learning control of the mechanical arm system with different test lengths and alignment conditions, and at least has the following beneficial effects:
[0015] 1. The application realizes high precision and fast error convergence performance of mechanical arm trajectory tracking by combining radial basis function network and dynamic adjustment mechanism.
[0016] 2. While ensuring the trajectory tracking precision, the application can effectively compensate for unmodeled dynamics and external disturbances, and provides strong adaptability and robustness support for mechanical arm control in complex dynamic environment. BRIEF DESCRIPTION OF DRAWINGS
[0017] The drawings described herein are used to provide further understanding of the application, and form a part of the application. The illustrative embodiments of the application and their descriptions are used to explain the application, and do not constitute an improper limitation on the application. In the drawings:
[0018] Figure 1 The mechanical arm trajectory change diagram in the first, third, fifth and twentieth iterations in the method proposed by the application;Figure 1 (a) in the figure is a curve showing the change of the trajectory in the first dimension. Figure 1 (b) in the figure is a curve showing the change of the trajectory in the second dimension;
[0019] Figure 2 This is a convergence curve showing the change in the mean magnitude of the trajectory error in the first 20 iterations of the method proposed in this invention.
[0020] Figure 3 The graphs show the dynamic response of the nonlinear compensation term estimated in the method proposed in this invention during the 1st, 3rd, 5th and 20th iterations.
[0021] Figure 4 The graph shows the dynamic changes of the control input in the 1st, 3rd, 5th and 20th iterations of the method proposed in this invention. Figure 4 (a) in the figure is a graph showing the change of control input in the first dimension. Figure 4 (b) in the figure is a graph showing the change of control input in the second dimension;
[0022] Figure 5 The graph shows a comparison of the mean magnitude of trajectory error between the method proposed in this invention and a variable gain iterative learning control method.
[0023] Figure 6 The graph shows a comparison of the mean magnitude of trajectory error between the method proposed in this invention and a nonlinear control method including an error compensation factor.
[0024] Figure 7 The graph shows a comparison of the mean magnitude of trajectory error between the method proposed in this invention and an improved adaptive iterative learning control method.
[0025] Figure 8 This is a comparison diagram of the control input between the method proposed in this invention and an improved adaptive iterative learning control method; Figure 8 (a) is a comparison chart of control inputs in the first dimension. Figure 8 (b) in the figure is a comparison chart of control input in the second dimension;
[0026] Figure 9 This is a comparison of the mean value of the control input modulus between the method proposed in this invention and an improved adaptive iterative learning control method.
[0027] Figure 10 This is a flowchart of the method proposed in this invention. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of this invention clearer, the following description is provided in conjunction with the appendix. Figures 1-10The present application will be further described in detail with reference to the embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the present application and are not intended to limit the present application.
[0029] The steps in the present application are arranged by using reference numerals, but are not used to limit the sequence of the steps, unless the sequence of the steps is explicitly described or the execution of a certain step needs other steps as a basis, otherwise the relative sequence of the steps can be adjusted. It can be understood that the term "and / or" used herein relates to and covers any and all possible combinations of one or more of the associated listed items.
[0030] As shown in Figure 10 The mechanical arm iterative learning control method based on the radial basis function network provided by the present application solves the problem of insufficient precision of the traditional mechanical arm control method in dealing with nonlinear dynamic compensation and disturbance influence, and the method specifically comprises the following steps:
[0031] S1, constructing a mechanical arm system dynamics model, and explicitly the dynamics characteristics of the mechanical arm system and setting applicable assumption conditions;
[0032] As a preferred embodiment, the mechanical arm system dynamics model constructed in step S1 is:
[0033]
[0034] wherein, respectively represent the actual displacement, velocity and acceleration of the mechanical arm; is the iteration learning number, t∈[0,T k ] is a time index; T k is the actual execution time of the kth iteration, T min is a constant and satisfies 0<T min ≤T k ≤T; is the inertia matrix, is the centripetal-Coriolis force matrix, is the gravity vector; τ k (t) is the torque input, i.e. the control input, d k (t) is a comprehensive disturbance vector containing unmodeled dynamics and external disturbances;
[0035] In addition, for the above formula, two points are explained here:
[0036] ①、In the present application, the actual displacement, velocity and acceleration of the robot arm correspond to the actual trajectory, the first derivative of the actual trajectory and the second derivative of the actual trajectory respectively, and can be collectively referred to as trajectory variables, and the three variables constitute the state vector of the robot arm system, which is used for subsequent trajectory error calculation and controller design; in addition, since the robot arm usually includes multiple joints, and the displacement, velocity and acceleration of each joint are independent, in the present application, the three variables are actually in the form of a multi-dimensional vector, and each dimension in the vector corresponds to the displacement, velocity and acceleration of a joint, and in the present application, each joint is not highlighted to ensure the mathematical simplicity of the dynamic equation;
[0037] ②、"τ k (t) is the torque input, that is, the control input" is because in the dynamic equation (emphasizing physical modeling) τ k (t) represents the torque acting on the joint of the robot arm, which directly affects the angular acceleration of the joint; therefore, if "torque input" is used in the dynamic equation, it is more accurate; if "control input" is used in the description of the control method and algorithm, it is more general, indicating that the control method can be applied to different actuators.
[0038] Next, for the sake of simplicity, the time and state dependent symbols of the system are omitted without causing confusion; for example, taking q k (t) as an example, some formulas in the present embodiment are simplified to q k , and the meanings expressed by the two expressions are consistent;
[0039] Based on the simplified dynamic model of the robot arm system, to further clarify the dynamic characteristics and assumption conditions of the system involved in the present application, the dynamic model should satisfy the following characteristics:
[0040] Characteristic 1, the inertia matrix M(q k ) is symmetric and positive definite, and its value remains bounded during the operation of the robot arm system;
[0041] Characteristic 2, the matrix is an anti-symmetric matrix; that is, it satisfies:
[0042]
[0043] where x is an arbitrary real matrix, x T is the transpose matrix of x; is the first derivative of the inertia matrix;
[0044] Characteristic 3, the centripetal-Coriolis force matrix and the gravity vector G(q k ) satisfy the following boundedness condition:
[0045]
[0046] G(q k )≤κ g ;
[0047] where κ c and κ g are positive parameters;
[0048] On the basis of the explicit dynamics, in order to ensure the rationality of the control strategy design and the rigor of the theoretical analysis, the following assumptions are introduced:
[0049] Assumption 1, the target trajectory and its first and second time derivatives, i.e., q d , and and the comprehensive disturbance vector d k are bounded under the condition of ;
[0050] Assumption 2, in each iterative learning control, the initial state of the manipulator system satisfies the alignment condition, i.e., q k (0)=q k-1 (T k ) and The dynamics model of the manipulator system is thus constructed.
[0051] It should be noted that in the iterative learning control of the present embodiment, the trajectory of the manipulator can be divided into the following three types:
[0052] Target trajectory q d : a pre-set ideal trajectory, i.e., the final trajectory that the manipulator hopes to achieve; q d is generally fixed and does not change with iteration;
[0053] Reference trajectory q m,k : an optimized trajectory used for control in the iteration process, which is adjusted gradually in each iteration to approximate the target trajectory q d ; q m,k is adjusted in the present application by a correction function so as to be more consistent with the actual operation of the system;
[0054] Actual trajectory q k : the trajectory actually executed by the manipulator in the kth iteration; q k is affected by the control input, un-modeled dynamics, external disturbances, etc., and generally deviates from the target trajectory q d .
[0055] S2, in combination with the initial trajectory data of the manipulator, a dynamic correction strategy is adopted to perform multiple iterative optimization on the reference trajectory of the manipulator, and a trajectory error is defined; at the same time, the starting and target states of the reference trajectory are ensured to meet the task requirements, and the smoothness and reachability of the optimized trajectory path are ensured.
[0056] As a preferred embodiment, step S2 specifically comprises:
[0057] S21、In this embodiment, under the condition of non-equal length test, in order to meet the requirements of iterative learning control, while ensuring the continuity and adaptability of trajectory dynamic adjustment, a trajectory optimization method based on correction function is proposed; the correction function is constructed to dynamically optimize the reference trajectory of the robot arm, and the formula is expressed as:
[0058]
[0059] Wherein, q m,k (t) represents the reference trajectory generated by the kth iteration optimization; And is a second-order differentiable correction function; used for dynamically adjusting the target trajectory, ensuring smooth transition of the trajectory at different time periods;
[0060] S22, in order to realize the dynamic adjustment of the trajectory and ensure the design performance of the controller, the constraint condition of the correction function needs to be set; specifically including:
[0061] The constraint condition of the correction function is set as:
[0062] When t∈[T a ,T], q (t) is a continuous function;
[0063] It is bounded on [0,T];
[0064] When t∈[T a ,T], q (t) is a continuous function;
[0065] When t∈[T a ,T], q (t) is a continuous function;
[0066] The constraint condition of the correction function is set as:
[0067] When t∈[T a ,T], q (t) is a continuous function;
[0068] It is bounded on [0,T];
[0069] When t∈[T a ,T], q (t) is a continuous function;
[0070] when t∈[T a ,T], and
[0071] where, are the first and second time derivatives of the correction function , respectively; are the first and second time derivatives of the correction function , respectively; T a is a design parameter satisfying 0 < T a ≤ T k ; the initial state of the reference trajectory q m,k (t) generated in each iteration is consistent with the final state of the last iteration, i.e., q m,k (0) = q m,k-1 (T k ) and
[0072] S23, to quantitatively describe the deviation between the actual trajectory and the reference trajectory of the system, and to provide a theoretical basis for controller design and performance analysis, the embodiment defines a trajectory error, which includes displacement error and velocity error The formula is expressed as:
[0073]
[0074] where q m,k (t) represents the reference trajectory generated by the kth iteration optimization, q k (t) represents the actual trajectory in the kth iteration process; and represent the velocity of the manipulator on the reference trajectory and the actual trajectory in the kth iteration, i.e., the first derivative of q m,k (t), q k (t).
[0075] S3, a radial basis function (RBF) network is designed, and by selecting the center position and width parameters of the basis function, a nonlinear compensation term that adapts to the dynamic characteristics is constructed to compensate for the influence of unmodeled dynamics and external disturbances on the manipulator system.
[0076] As a preferred embodiment, step S3 specifically includes:
[0077] S31, in the dynamics model of the manipulator system, the dynamic characteristics of the trajectory error may be affected by the comprehensive disturbance vector d k , the gravity vector G(q k) and control input, etc. To constrain the error dynamic behavior and guarantee the control performance of the system, the method proposed in the present application introduces a nonlinear compensation term to describe the dynamic characteristics of the trajectory error and impose constraints; first, according to the dynamic properties, the dynamic model of the manipulator system is reconstructed as:
[0078]
[0079] In combination with the optimized reference trajectory and the error definition, the above formula can be further expressed as
[0080]
[0081] And based on the nonlinear compensation term, the following constraint is established for the trajectory error:
[0082]
[0083] Wherein, α k (t) is the real nonlinear compensation term; is the velocity error, and sign(.) is the sign function; q k respectively represent the acceleration of the manipulator on the reference trajectory and the actual trajectory of the manipulator in the kth iteration process; ∈ is an error smoothing parameter; the left side of the constraint inequality reflects the relationship between the trajectory error dynamics and the inertia matrix of the system, the gravity vector G(q k ), and the comprehensive disturbance vector d k ; and by imposing constraints on the error through the real nonlinear compensation term α k (t) and the sign function sign(·), it can be ensured that the influence of the velocity error will not be too large and will gradually converge with iteration;
[0084] More specifically, according to the known inequality, for any ∈ > 0 and χ ∈ R, there is always:
[0085]
[0086] Therefore, the constraint established on the velocity error based on the nonlinear compensation term can be further divided as:
[0087]
[0088] S32, to accurately express and compensate the nonlinear dynamic behavior and disturbance of the manipulator system, the present example approximates and approaches the real nonlinear compensation term α k (t) through an RBF network, which is expressed in the formula as:
[0089] α k (t) = θ T (t) μ(Ψ(t), t) + ε;
[0090] wherein, is the input feature vector, q k respectively represent the acceleration of the robot arm on the reference trajectory and the actual trajectory of the robot arm in the kth iteration process; is the ideal weight parameter of the RBF network, but is usually unknown in actual calculation. p is the number of nodes of the base function, and ε is the approximation error. Increasing the number of base function nodes p reduces the approximation error ε to be small enough; is the base function vector of the RBF network; each base function is defined as:
[0091]
[0092] wherein, Λ i and γ i are the center and width of the base function respectively; the RBF network can dynamically adapt to the characteristics of the error and estimate the nonlinear characteristics and changes of the system by learning the base function vector and the weight parameter.
[0093] S4, the estimated dynamic weight parameter and the dynamic learning gain are designed to optimize the response speed and output smoothness of the controller in the iteration process, reduce the oscillation problem of the control input, and ensure the progressive convergence of the weight error;
[0094] As a preferred embodiment, step S4 specifically includes:
[0095] S41, the estimated dynamic weight parameter is designed to realize dynamic adjustment and step-by-step optimization of the weight parameter in the radial basis function RBF network in the iteration process; the update rule of the estimated dynamic weight parameter is
[0096]
[0097] wherein, are the estimated dynamic weight parameters in the kth and (k-1)th iteration processes respectively; the initial estimated dynamic weight parameter is set as to ensure the optimization from zero; through iterative learning, θ(t) is gradually approached to realize optimal control; μ(Ψ(t), t) as the base function vector of the RBF network significantly improves the fitting ability and flexibility of the system by embedding the error characteristics; the dynamic learning gain η k (t) > 0 is responsible for adjusting the speed of parameter update, thereby effectively controlling the speed and stability of the convergence process; tanh(.) is the hyperbolic sine function, is the speed error in the kth iteration process, ∈ is the error smoothing parameter, and here is used to limit the influence range of the trajectory error to avoid excessive error affecting the adjustment stability;
[0098] S42、The method proposed in the application is matched with the dynamic weight parameter, and a dynamic learning gain is designed, that is, a dynamic adjustment and gradual optimization mechanism is designed, the dynamic learning gain at different times in the iteration process is adjusted, and the iteration process is controlled to consider efficiency and stability;
[0099] The adjustment rule of the dynamic learning gain is:
[0100]
[0101] Wherein, η k (t) is the dynamic learning gain in the kth iteration process; η0>0 is the initial gain, Δ η is the gain amount, which can be any large value, and is used to accelerate the initial parameter learning; ζ k (t) is the gain parameter in the kth iteration process, which ensures the stability of the adjustment process;
[0102] The specific definition of the gain parameter is:
[0103]
[0104] Wherein, the initial value ζ1(t)=1, indicating that the system starts from the most basic adjustment starting point; in the dynamic adjustment process, the initial learning gain η k (t) and the lower incremental parameter ζ k (t) are used to enable the system to quickly respond to errors and achieve preliminary adjustment; in the middle and later stages, with the gradual increase of the incremental parameter ζ k (t), the learning gain η k (t) is automatically reduced, and the adjustment amplitude is gradually reduced, so as to ensure that the adjustment process of the system tends to be stable, and the robustness and accuracy of the overall control are ensured;
[0105] The method proposed in the application realizes the adaptive adjustment of the dynamic characteristics of the trajectory error by designing the estimated dynamic weight parameter and the dynamic learning gain η k (t), so that the controller can dynamically optimize the parameters according to the real-time error, and improve the trajectory tracking accuracy and adaptability to disturbances of the manipulator system.
[0106] S5, in combination with feedback control, nonlinear compensation and iterative learning method, an iterative learning controller is designed to fully utilize historical data, adjust the control law in real time, gradually reduce the trajectory error, and enhance the adaptability of the control system;
[0107] As a preferred embodiment, according to the contents described in steps S1-S4, the design of the manipulator system controller in this embodiment includes the following steps:
[0108] Firstly, by analyzing the dynamic characteristics of the manipulator system, the appropriate feedback gain matrix K P and K D are selected to adjust the dynamic response of displacement error and velocity error;
[0109] Then, the nonlinear compensation term is designed, the basis function vector is constructed by RBF network, and the network is ensured to have sufficient approximation ability by selecting appropriate basis function center and width. At the same time, the estimated dynamic weight parameter is initially optimized to zero, so as to realize dynamic adjustment in subsequent iterations;
[0110] Subsequently, the current reference trajectory is dynamically modified combined with the trajectory data in each iteration, and the starting point is ensured to be consistent with the ending state of the last round;
[0111] Finally, the control input is calculated by using the trajectory error of the current iteration, combined with the feedback term and the nonlinear compensation term, to realize the real-time update of the control law;
[0112] More specifically, the specific structure of the control law is represented as:
[0113]
[0114] Where τ k (t) is the control input in the kth iteration process, K P and K D are the positive definite feedback gain matrices of displacement error and velocity error, respectively, and represent the displacement error and velocity error of the kth iteration, respectively; is the estimated dynamic nonlinear compensation term, the hyperbolic tangent function is used to limit the amplitude of the compensation term, and ∈ is the error smoothing parameter used to limit the influence range of the trajectory error;
[0115] The estimated dynamic nonlinear compensation term is calculated by RBF network, that is:
[0116]
[0117] Where, is the estimated dynamic weight parameter of RBF network; μ(Ψ(t),t) is the basis function vector of RBF network, and Ψ(t) is the input feature vector of RBF network;
[0118] Based on the control law, the controller combines the synergistic effect of feedback control, nonlinear compensation and iterative learning, so as to dynamically adjust the trajectory error and ensure the error to gradually converge. The dynamic equation of trajectory error is represented as:
[0119]
[0120] In this embodiment, the inertia matrix M -1 (q k ), the centripetal-Coriolis force matrix The gravity term G(q k ) and the disturbance vector d k all affect the dynamic characteristics of the system. The controller adjusts the error dynamic response through feedback control, and uses the estimated dynamic nonlinear compensation term generated by the RBF network to realize real-time compensation for unmodeled dynamics and external disturbances, ultimately ensuring the step-by-step reduction of displacement error and velocity error, and improving the trajectory tracking accuracy and robustness of the system.
[0121] The controller combines feedback control and nonlinear compensation to accurately adjust the trajectory error and improve tracking performance, while dynamically responding to external disturbances and unmodeled dynamics to enhance the adaptability of the system. Through iterative learning method, historical data is fully utilized to accelerate error convergence speed, and dynamic response ability that flexibly adapts to different task requirements is exhibited, thereby ensuring the efficiency and stability of the control system.
[0122] At this point, the mechanical arm system dynamics model and control law in the method proposed by the present application have been constructed; in the design of the mechanical arm iterative learning control method, verifying its stability and error convergence is the key to ensuring the feasibility and theoretical integrity of the system; stability analysis aims to prove that the control law ensures the global asymptotic stability of the system under complex dynamic environment; error convergence analysis focuses on whether the trajectory error (displacement error and velocity error) can converge to zero as the number of iterations tends to infinity; therefore, the present application uses Lyapunov stability theory combined with composite energy function CEF to strictly derive and prove the stability and convergence of the control algorithm.
[0123] S6, based on Lyapunov stability theory and composite energy function CEF, analyze the stability of the control system under complex dynamic environment, test the global asymptotic stability of the controller and the step-by-step convergence of the trajectory error of the mechanical arm;
[0124] As a preferred embodiment, step S6 specifically comprises:
[0125] S61, construct a CEF containing displacement error, velocity error and network weight error as a Lyapunov function to analyze the overall energy change of the mechanical arm system; for simplicity of expression, the time and state dependent symbols of the mechanical arm system are omitted; then the constructed CEF is expressed by the formula:
[0126] V k (t)=V 1,k (t)+V2,k (t) + V 3,k (t) ;
[0127] wherein, V 1,k (t) represents an error energy term related to the inertia matrix M(q k ) and the velocity error , used to characterize the influence of error dynamics on the system energy; V 2,k (t) represents a displacement error term, combined with the feedback gain in the control law, used to describe the displacement error V 3,k (t) represents a weight error term, reflecting the deviation between the estimated weight and the ideal weight θ(t) in the RBF network;
[0128] The error energy term, the displacement error term and the weight error term are respectively represented by the following formulas:
[0129]
[0130] wherein, the weight error s is an auxiliary scalar, and satisfies for regulating the dynamic behavior of the composite energy function; the core design intention is to ensure the rapid decay of the energy function, weaken the cumulative influence of historical errors, and at the same time meet the convergence condition of Lyapunov function, so as to realize the global stability of the system and the effective convergence of the error;
[0131] It should be noted that in the method proposed in the present application, θ(t) is the ideal weight parameter, but it is usually unknown; the controller estimates it online through iterative learning, and adjusts so that it gradually approximates θ(t); the error term reflects the gap between the estimated weight and the ideal weight; mathematically, θ(t) and both exist at the same time, but in the estimation calculation of the controller iterative learning, only θ(t) is used as a theoretical reference; through iterative optimization, the error between the two gradually decreases, and finally converges to θ(t), ensuring the performance of the control system.
[0132] S62, combined with the dynamics model of the mechanical arm system and the control law, according to the Lyapunov stability theory, the kinetic energy error term, the displacement error term and the weight error term are respectively derived, the monotonic boundedness of the composite energy function is verified, and the stability of the mechanical arm system and the convergence of the mechanical arm trajectory error are analyzed;
[0133] The specific derivation process is as follows:
[0134] S621, analyze the energy change of the system at the kth iteration, which is defined as the difference:
[0135] AV k (t) = AV 1,k (t) + AV 2,k (t) + AV 3,k (t).
[0136] where AV 1,k (t) is the error energy term related to the inertia matrix M(q k ) and the velocity error AV 2,k (t) is the displacement error term energy change, AV 3,k (t) is the weight error term energy change.
[0137] Then, the embodiment analyzes the energy change of each error term in detail.
[0138] S622, analyze the kinetic energy error term energy change; according to the mechanical arm system dynamics model and the properties of Lyapunov function, the mathematical expression of the error energy term AV 1,k (t) is as follows:
[0139]
[0140] According to assumption 2, it is assumed in the iterative learning process that the initial state of each iteration is consistent with the end state of the last iteration, that is, and then V 1,k (0) and the end energy V 1,k-1 (T k-1 ) of the last iteration can establish a relationship, that is, V 1,k (0) = V 1,k-1 (T k-1 ).
[0141] Then, replace the derivative in the integral term of the above formula with the corresponding dynamic expression, and scale according to the velocity error constraint inequality , and finally obtain:
[0142]
[0143] where α k represents the real nonlinear compensation term, which is usually unknown, and is the nonlinear compensation term estimated by the RBF network of the controller. To evaluate the deviation between the nonlinear compensation term estimated by the RBF network and the real nonlinear compensation term α k , define the nonlinear compensation term error The error term is used for Lyapunov convergence analysis to ensure that the displacement error e is gradually reduced, so that converges to a k , ensuring control accuracy. The relationship between a k and a k is also referred to and θ(t); in addition, δ k = 0.2785n∈η k .
[0144] S623, analyze the displacement error term energy change; the displacement error term energy change ΔV 2,k (t) is as follows:
[0145] ΔV 2,k (T) = V 2,k (T k )-V 2,k-1 (T k-1 );
[0146] Similarly, according to assumption 2, V 2,k (0) = V 2,k-1 (T k-1 ), and combined with the definition of the energy function derivative, the derivation process is referred to the kinetic energy error term energy change ΔV 1,k (T); in addition, due to the limited action time of η k , the rapid decay of e -st , and the gradual reduction of the displacement error , its influence on the overall energy change gradually weakens; therefore, the secondary contribution of η k can be reasonably ignored, and the focus is placed on the core role of the displacement error and the feedback gain; finally, the derivation process of the displacement error term energy change is as follows:
[0147]
[0148] S624, analyze the weight error term energy change; the derivation of the weight error term energy change is based on the dynamic characteristics of the weight error, by gradually decomposing the integral term and combining the Lyapunov energy function constraint, using the nonlinear mapping ability of the RBF network, the influence range of the weight adjustment is constrained, and finally ΔV 3,k (T) exists an upper bound. The specific derivation process is as follows:
[0149]
[0150] S625, according to steps S621-S624, the composite energy function change ΔV k (t) satisfies the following inequality:
[0151]
[0152] From the inequality, ΔV k (t) is a non-increasing sequence; and when ΔV k (t)≤0; thus, as long as V1(t) is bounded, it can be further deduced that the compound energy function V k (t) remains bounded throughout the iteration process.
[0153] Therefore, the boundedness of V1(t) is further demonstrated in this embodiment; the specific process includes:
[0154] 1) define the derivative of V1(t) as:
[0155]
[0156] Substitute the control law and feedback gain, and rearrange to obtain its upper bound form:
[0157]
[0158] 2) analyze V1(t) using integral inequality; at the same time, considering that V1(0) = 0, it can be further deduced that:
[0159]
[0160] From this, a parameter M is introduced, representing the maximum energy change of the system, such that:
[0161]
[0162] Thus, it is proven that V1(t) is bounded within , i.e.:
[0163]
[0164] According to the inequality satisfied by the compound energy function change in step S625, we obtain:
[0165]
[0166] Due to the non-negativity of V k (t), we further have:
[0167]
[0168] This inequality shows that as the iteration number k increases, the accumulated trajectory error energy gradually tends to zero, i.e. converges. Therefore, it can be further deduced that Based on the continuity of error dynamics, the L2 norm of displacement error will also tend to zero with the number of iterations, that is Thus, the stability and convergence of the control method proposed in the present application are verified.
[0169] In order to enable those skilled in the art to better understand the implementation of the present application, the present embodiment will use Matlab software to perform simulation experiments to verify the reliability of the method proposed in the present application.
[0170] The specific information of the simulation software is as follows:
[0171] Software name: MATLAB;
[0172] Version information: 9.8.0.1380330 (R2020a) Update 2;
[0173] License number: 919961;
[0174] Operating system: Microsoft Windows 11 Chinese version Version 11.0 (Build22621.2861);
[0175] Java version: Java 1.8.0_202-b08 with Oracle Corporation Java HotSpot(TM) 64-Bit Server VM mixed mode;
[0176] Special toolbox: Statistics and Machine Learning Toolbox-11.7 (R2020a), Aircraft Control Toolbox-1.0;
[0177] In the present embodiment, in order to verify the effectiveness of the mechanical arm iterative learning control method based on the radial basis function network proposed in the present application and its adaptability in complex dynamic environment, a plurality of simulation experiments are performed, and a comparative analysis is performed with three existing methods.
[0178] In order to verify the effectiveness of the method proposed and its adaptability in complex dynamic environment, a plurality of simulation experiments based on a two-degree-of-freedom mechanical arm system are performed. The classical dynamic model is used in the experiment, and the specific forms of the inertia matrix M, the centripetal-Corriolis force matrix C and the gravity vector G are as follows:
[0179]
[0180] The comprehensive disturbance vector d = [d1, d2] Td1 = d2 = rand(k) sin(t) is set for simulating unknown disturbance in complex dynamic environment, wherein rand(k) is a random function with a value range of [0, 1]. The target trajectory q d is composed of q d,1 and q d,2 , that is The execution time T = 1 s is set, and the initial state of the trajectory of the manipulator in the first iteration is q1(0) = [0, 0] T . The non-equal length test condition is introduced, and it is assumed that T k is uniformly distributed in the interval [0.7, 1] and the probability
[0181] The control law adopts the iterative learning control algorithm proposed in the application, and the parameters include an initial learning gain η0 = 10, a gain amount Δ η = 1000, a feedback gain matrix K P of displacement error = diag{500, 500}, a positive definite feedback gain matrix K D of velocity error = diag{150, 500}, and an error smoothing parameter ∈ = 0.1.
[0182] The correction functions are respectively: wherein t ∈ [0, T a ), T a = 0.1 s.
[0183] The RBF neural network includes 6 base functions, and the center parameters Λ i are respectively Λ1 = [1, 1, 1, 1] T , Λ2 = [1, 1, -1, -1] T , Λ3 = [1, -1, 1, -1] T , Λ4 = [-1, -1, -1, -1] T , Λ5 = [-1, -1, 1, 1] T , and Λ6 = [-1, 1, -1, 1] T ; and the width parameter is γ i = 6.
[0184] In addition, φ e represents the average value of the displacement error module at all sampling points, and reflects the trajectory tracking accuracy of the manipulator system; and φ τ represents the average value of the control input module ||τ k (t)|| at all sampling points, and is used for evaluating the smoothness of the control input; and the sampling period is set as T = 0.001 s.
[0185] The trajectories q k(t)(q k,1 , q k,2 represent the trajectory of the first dimension and the second dimension, respectively) and the control input τ k (t) are recorded, and the corresponding variation trend graphs are drawn, the results are shown in Figure 1 and Figure 4 , respectively. In addition, Figure 2 shows the variation curve of φ e in the first 20 iterations, Figure 3 shows the variation trend of the estimated nonlinear compensation term . From Figure 1 and Figure 2 , it can be seen that the trajectory error converges rapidly, indicating that the system has high tracking performance; Figure 3 and Figure 4 indicate that the estimated nonlinear compensation term and the control input gradually stabilize, and the control law has no obvious chattering phenomenon. The above simulation results verify the effectiveness of the control method proposed in the present application. By introducing the dynamic gain adjustment mechanism and the nonlinear compensation term, the tracking accuracy and control stability of the system are effectively improved, and the system is adapted to complex dynamic environments.
[0186] In order to verify the superiority of the control method proposed in the present application, it is compared with three existing methods for simulation analysis, which includes the following contents:
[0187] Existing method one: a variable gain iterative learning control method
[0188] This method uses a variable gain-based iterative learning control method under the initial condition q k (0) = [0, 1] T . Its control law expression is:
[0189]
[0190] Where S t is the sampling time index, ζ k (S t ) is the dynamic gain parameter. The simulation results are shown in Figure 5 . By comparing the average value of the trajectory error modulus φ e in the first 20 iterations, the method of the present application shows faster convergence speed and lower error level.
[0191] Existing method two: a nonlinear control method containing error compensation factor
[0192] For non-equal length test conditions, a nonlinear control method containing error compensation factor is used, and its control law is designed based on the following form:
[0193]
[0194] where, is part of the system dynamics, including the inertia matrix M(q k ), the centripetal-Coriolis matrix and the gravity vector G(q k (t)). The update rule of the compensation factor is as follows:
[0195]
[0196] The simulation results are shown in Figure 6 , and the average value of the trajectory error modulus φ e of the first 20 iterations is compared. The method of the application is significantly better than method two, showing higher trajectory tracking accuracy and dynamic adaptability.
[0197] Existing method three: an improved adaptive iterative learning control method
[0198] Assuming that the operation can be completely performed, an improved adaptive iterative learning control method is adopted, and the control law is:
[0199]
[0200] where, The update rule of the compensation factor
[0201]
[0202] Figure 7 The average value of the trajectory error modulus φ e of the first 20 iterations of different methods is shown. The results show that the convergence of the trajectory error modulus and the final error of the method of the application are better than those of the existing method three, showing faster convergence speed and lower error level. Figure 8 The control input τ k (t) of different methods in the 20th iteration is shown. The results show that by introducing the hyperbolic tangent function, the method of the application can effectively suppress the input chattering phenomenon, thereby significantly improving the smoothness of the control input. Figure 9 The average value of the control input modulus φ τ of different methods in the first 20 iterations is shown. The results show that the method of the application successfully limits the amplitude of the control input amplitude through the dynamic learning gain dynamic adjustment mechanism, avoiding the problem of over-learning.
[0203] In summary, the control method proposed in the application has superior performance in trajectory tracking accuracy, control input smoothness and control input amplitude optimization, and also has effectiveness and adaptability in complex dynamic environments.
[0204] It will be obvious to a person skilled in the art that the application is not limited to the details of the foregoing exemplary embodiments and can be implemented in other concrete forms without departing from the spirit or essential characteristics of the application. The embodiments are therefore to be considered in all respects as illustrative and not restrictive, the scope of the application being indicated by the appended claims rather than by the foregoing description, and all changes which come within the meaning and range of equivalency of the claims are therefore intended to be embraced therein. No reference signs in the claims should be considered as limiting the scope of the claims to the identity of the reference signs therein.
[0205] Furthermore, it should be understood that although the description is made on the basis of the embodiments, not every embodiment contains only one independent technical solution, and the description of the specification is only for the sake of clarity, and those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that those skilled in the art can understand.
Claims
1. An iterative learning control method for a robotic arm based on radial basis function networks, characterized in that, Specifically, the following steps are included: S1. Construct a dynamic model of the robotic arm system, define its dynamic characteristics, and set applicable assumptions. The constructed dynamic model of the robotic arm system is as follows: in, These represent the actual displacement, velocity, and acceleration of the robotic arm, respectively. Let t be the number of iterations for learning, where t∈[0,T] k ] represents the time index; T k Let T be the actual execution time of the k-th iteration, and satisfy 0 < T. min ≤T k ≤T, T min T represents the shortest execution time for a single iteration, and T represents the maximum execution time for a single iteration. The inertia matrix, The centripetal-Coriolis force matrix, τ is the gravity vector; k (t) represents the torque input, which is also the control input, and d k (t) is the combined disturbance vector that includes unmodeled dynamics and external disturbances; Defining the dynamic characteristics of the robotic arm system and setting applicable assumptions specifically includes: To simplify the expression, time and state dependency symbols are omitted; therefore, the dynamic model of the robotic arm system constructed in step S1 must satisfy the following dynamic characteristics: Characteristic 1, Inertia matrix M(q) k The value of is symmetric and positive definite, and remains bounded during the operation of the robotic arm system. Feature 2, Matrix For antisymmetric matrices; where, The first derivative of the inertia matrix; Characteristic 3: Centripetal-Coriolis force matrix and gravity vector G(q) k It satisfies the following boundedness conditions: G(q k )≤κ g ; Among them, κ c and κ g It is a positive parameter; Based on a clear understanding of the dynamic characteristics, the following assumptions are introduced: Assumption 1, the target trajectory and its first and second time derivatives, i.e., q d , and and the combined perturbation vector d k exist Under all conditions, they are bounded; Assumption 2: In each iteration of learning control, the initial state of the robotic arm system satisfies the alignment condition, i.e., q k (0)=q k-1 (T k )and S2. Combining the initial trajectory data of the robotic arm, a dynamic correction strategy is adopted to iteratively optimize the reference trajectory of the robotic arm multiple times, and the trajectory error is defined; at the same time, it is ensured that the initial and target states of the reference trajectory meet the task requirements, and the smoothness and reachability of the trajectory path are optimized; step S2 specifically includes the following steps: S21. Construct a correction function to dynamically optimize the reference trajectory of the robotic arm. The formula is expressed as: Where, q m,k (t) represents the reference trajectory generated in the k-th iteration optimization; and It is a twice differentiable correction function; S22. Set the constraints for the correction function; specifically including: Set correction function The constraints are: When t∈[T] a When [T], and It is bounded on [0,T]; When t∈[T] a When [T], and When t∈[T] a When [T], and Set correction function The constraints are: When t∈[T] a When [T], and It is bounded on [0,T]; When t∈[T] a When [T], and When t∈[T] a When [T], and in, Correction functions The first and second time derivatives; Correction functions First-order and second-order time derivatives; T a For design parameters, satisfy 0 < T a ≤T k The reference trajectory q generated in each iteration m,k The starting state of (t) is consistent with the ending state of the previous iteration, i.e., q m,k (0)=q m,k-1 (T k )and S23. Define trajectory error; trajectory error includes displacement error. and speed error The formula is expressed as: Where, q m,k q(t) represents the reference trajectory generated in the k-th iteration optimization. k (t) represents the actual trajectory during the k-th iteration; and Let q represent the velocities of the robotic arm on the reference trajectory and the actual trajectory in the k-th iteration, respectively. m,k (t), q k The first derivative of (t); S3. Design a radial basis function (RBF) network. By selecting the center position and width parameters of the basis functions, construct a nonlinear compensation term that adapts to the dynamic characteristics to compensate for the influence of unmodeled dynamics and external disturbances on the robotic arm system. S4. Design the estimated dynamic weight parameters and dynamic learning gain to optimize the controller's response speed and output smoothness during the iteration process, reduce the oscillation problem of the control input, and ensure the asymptotic convergence of the weight error. S5. Combining feedback control, nonlinear compensation, and iterative learning methods, an iterative learning controller is designed to fully utilize historical data, adjust the control law in real time, gradually reduce trajectory errors, and enhance the adaptability of the control system. The specific structure of the control law is expressed as follows: Where, τ k (t) represents the control input during the k-th iteration, K P and K D These are the positive definite feedback gain matrices for displacement error and velocity error, respectively. and Let represent the displacement error and velocity error of the k-th iteration, respectively; For the estimated dynamic nonlinear compensation term, the hyperbolic tangent function Used to limit the magnitude of the compensation term, ∈ is the error smoothing parameter, used to limit the range of influence of trajectory error; The estimated dynamic nonlinear compensation term is calculated by the RBF network, i.e.: in, For the estimated dynamic weight parameters, μ(Ψ(t),t) is the basis function vector of the RBF network, and Ψ(t) is the input feature vector of the RBF network. The trajectory error of the robotic arm is dynamically adjusted by a control law. The dynamic equation of the trajectory error is expressed as: S6. Based on Lyapunov stability theory and composite energy function CEF, analyze the stability of the control system under complex dynamic environment, and verify the global asymptotic stability of the controller and the gradual convergence of the robot arm trajectory error.
2. The iterative learning control method for a robotic arm based on radial basis function networks according to claim 1, characterized in that, Step S3 specifically includes: S31. Introducing a nonlinear compensation term to describe the dynamic characteristics of the trajectory error and imposing constraints; omitting the system's time and state dependency symbols to simplify the expression, the following constraints are established for the trajectory error based on the nonlinear compensation term: Where, α k (t) represents the actual nonlinear compensation term; For velocity error, sign(.) is the sign function; q k Let represent the acceleration of the robotic arm on the reference trajectory and the actual trajectory of the robotic arm during the k-th iteration, respectively; ∈ is the error smoothing parameter; S32. Applying an RBF network to the actual nonlinear compensation term α k (t) is approximated by the formula: a k (t)=θ T (t)μ(Ψ(t),t)+ε; in, For the input feature vector, q k Let X represent the acceleration of the robotic arm on the reference trajectory and the actual trajectory of the robotic arm during the k-th iteration, respectively. θ represents the ideal weight parameters for the RBF network. T (t) is the transpose of θ(t); p is the number of basis function nodes, and ε is the approximation error. Here is the basis function vector of the RBF network; each basis function is defined as: Among them, Λ i γ i These represent the center and width of the basis functions, respectively.
3. The iterative learning control method for a robotic arm based on radial basis function networks according to claim 2, characterized in that, Step S4 specifically includes: S41. Design the estimated dynamic weight parameters to achieve dynamic adjustment and stepwise optimization of the weight parameters in the radial basis function (RBF) network during the iteration process; the update rule for the estimated dynamic weight parameters is as follows: in, These are the dynamic weight parameters estimated during the k-th and (k-1)-th iterations, respectively; the initial values of the estimated dynamic weight parameters are set to... Through iterative learning, The optimal control is achieved by gradually approximating θ(t); μ(Ψ(t),t) is the basis function vector of the RBF network; the dynamic learning gain η is used. k (t) > 0 is responsible for adjusting the speed of parameter updates; tanh(.) is a hyperbolic sine function. denoted as the velocity error during the k-th iteration, and ∈ is the error smoothing parameter used to limit the influence range of the trajectory error; S42. Design a dynamic adjustment and stepwise optimization mechanism; by adjusting the dynamic learning gain at different times in the iteration process, control the iteration process to balance efficiency and stability. The adjustment rule for dynamic learning gain is as follows: Where, η k (t) represents the dynamic learning gain during the k-th iteration; η0>0 is the initial gain, Δ η ζ is the gain quantity. k (t) represents the gain parameter during the k-th iteration; The specific definition of the gain parameter is: The initial value ζ1(t) = 1 indicates that the system starts from the most basic adjustment point.
4. The iterative learning control method for a robotic arm based on radial basis function networks according to claim 1, characterized in that, Step S6 specifically includes: S61. Construct a CEF (Compound Energy Function) that includes displacement error, velocity error, and network weight error as a Lyapunov function to analyze the overall energy change of the robotic arm system; for simplification, the time and state dependency symbols of the robotic arm system are omitted; the constructed composite energy function is expressed by the formula: V k (t)=V 1,k (t)+V 2,k (t)+V 3,k (t); Among them, V 1,k (t) represents the relationship with the inertia matrix M(q) k and speed error The relevant error energy term is used to characterize the effect of error dynamics on the system energy; V 2,k (t) represents the displacement error term, which, combined with the feedback gain in the control law, is used to describe the displacement error. V 3,k (t) represents the weight error term, reflecting the estimated dynamic weight parameters in the RBF network. The deviation between the ideal weight parameter θ(t); The error energy term, displacement error term, and weighted error term are expressed by the following formulas: in, s is an auxiliary scalar, and satisfies Δ η η k Let be the gain and the dynamic learning gain during the k-th iteration, respectively; S62. Combining the dynamic model and control law of the robotic arm system, and based on Lyapunov stability theory, the kinetic energy error term, displacement error term, and weighted error term are derived respectively. The monotonically bounded nature of the composite energy function is verified, and the stability of the robotic arm system and the convergence of the robotic arm trajectory error are analyzed.
Citation Information
Patent Citations
Plane mechanical arm trajectory tracking control method based on iterative learning
CN114995144A