An adaptive dynamic programming optimal sliding mode control method for multiple manipulators

Through adaptive dynamic programming and sliding mode control methods, combined with specified time performance functions and self-adjusting functions, an adaptive dynamic programming optimal sliding mode controller for a multi-manipulator system is designed. This solves the shortcomings of the multi-manipulator system in performance constraints, control accuracy and trajectory tracking, achieves flexible constraints and improved robustness, reduces the amount of computation and achieves optimal control effects.

CN119644756BActive Publication Date: 2025-09-30GUANGDONG OCEAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510002984.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-09-30
Estimated Expiration
2045-01-02

AI Technical Summary

Technical Problem

Existing control methods for multi-manipulator systems have shortcomings in performance constraints, control accuracy, and trajectory tracking, especially in the face of system dynamic uncertainty and complexity, making it difficult to achieve efficient optimal control.

Method used

An adaptive dynamic programming method is used, combined with the specified time performance function, self-adjusting function and sliding mode control, to design an adaptive dynamic programming optimal sliding mode controller for a multi-manipulator system. A sliding mode control method with flexible constraint capability is designed through the self-adjusting performance function and inverse tangent function, and a critical neural network dynamic programming controller is constructed using the gradient descent method and parallel learning technology.

Benefits of technology

The flexible performance constraints of the multi-manipulator system are realized, the robustness and control accuracy of the system are improved, the calculation amount is reduced, and the optimal control effect is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119644756B_ABST
    Figure CN119644756B_ABST
Patent Text Reader

Abstract

This invention discloses an optimal sliding mode control method for multiple robotic arms using adaptive dynamic programming. This method relates to the field of trajectory tracking control for multiple robotic arms, including: establishing a two-degree-of-freedom dynamic model for the multiple robotic arm system; defining the two-degree-of-freedom position error of the multiple robotic arm system; designing a self-adjusting performance control method for the multiple robotic arm system based on a specified time performance function and a self-adjusting function; designing a sliding mode control method for the multiple robotic arm system with flexible constraints based on the sliding mode control principle; and designing an optimal sliding mode controller for the multiple robotic arm system based on adaptive dynamic programming. This invention addresses the performance constraint and limited control accuracy issues inherent in current control methods for multiple robotic arm systems, as well as the optimal control problem for trajectory tracking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of trajectory tracking control of multiple robotic arms, and in particular to an adaptive dynamic programming optimal sliding mode control method for multiple robotic arms. Background Art

[0002] With the continuous advancement of technologies such as intelligence, big data, and cloud computing, robotic arms are expanding from single functions to multiple functions, and from fixed scenarios to complex ones. Their applications in smart manufacturing, healthcare, logistics, and other fields will continue to expand. Therefore, multi-robot systems will develop towards greater intelligence and sophistication. Due to the complexity and interdependence of multi-robot systems, control performance is particularly important in this development, and this has been a key research focus in recent years on multi-robot collaborative control.

[0003] Specified performance control is an effective method to achieve good performance. The principle of specified performance control is to provide constraints on the tracking error by specifying the performance function, ensuring that the convergence of the error meets the predetermined conditions, thereby achieving the required transient and steady-state performance. In recent years, through the unremitting efforts of researchers, many specified performance control methods have been proposed. Using specified performance control, the tracking error converges to a predetermined area, ensuring the transient and steady-state performance of the system. However, traditional performance control methods have a priori conditions, that is, the initial state must be within a specified range. At the same time, traditional performance control does not consider the threshold requirements when the system enters the steady-state stage, and may be too conservative in parameter selection, affecting control efficiency. Therefore, designing a more flexible specified performance control method for multi-manipulator systems is crucial to ensuring the performance of multi-manipulator collaborative control.

[0004] Sliding mode control has gained increasing attention due to its advantages such as fast response speed, strong resistance to external interference, and simple physical implementation. Therefore, the research on sliding mode control is very practical, and many methods have been proposed to solve control problems. The boundary layer method has a good effect in improving robustness. However, this method makes it difficult to ensure the existence of sliding modes within the boundary layer, which means that the method has to sacrifice a certain degree of tracking accuracy. High-order sliding modes can improve the robustness of the system without affecting the tracking performance. However, the complex time derivatives of high-order sliding modes need to be considered. Therefore, designing a suitable sliding mode controller for multi-manipulator collaborative control remains a challenging research.

[0005] The goal of optimal control is to minimize the performance indicator function while pursuing system stability, thereby achieving a balance between system performance and control resources. Reinforcement learning has attracted widespread attention from researchers due to its success in solving optimal control problems. Given that the system dynamics are known, an action-critic framework of reinforcement learning is constructed based on neural networks or fuzzy logic systems to solve the optimal control problem of the system. However, the system often has uncertain dynamics, which can be estimated through identifiers. However, using the identifier-action-critic framework in a multi-manipulator system will greatly increase the computational intensity of the system. Dynamic programming is considered to be an effective and classic tool for solving the optimal control problem of nonlinear systems. In the 1970s, researchers developed adaptive dynamic programming through neural networks or fuzzy logic systems. In recent decades, the research on adaptive dynamic programming has achieved great success and has attracted widespread attention from researchers. Therefore, one of the objectives of the present invention is to solve the optimal control problem of multi-manipulators through adaptive dynamic programming. Summary of the Invention

[0006] In response to the above-mentioned deficiencies in the prior art, the present invention provides a multi-manipulator adaptive dynamic programming optimal sliding mode control method that solves the performance constraint problems, limited control accuracy problems, and optimal control problems of trajectory tracking that exist in the application of current multi-manipulator system control methods.

[0007] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention is: a multi-manipulator adaptive dynamic programming optimal sliding mode control method, comprising the following steps:

[0008] S1: Establish a 2-DOF dynamic model of the multi-manipulator system;

[0009] S2: Based on the 2-DOF dynamic model, define the 2-DOF position error of the multi-manipulator system;

[0010] S3: Based on the 2-DOF position error, design a self-tuning performance control method for a multi-manipulator system with a specified time performance function and a self-tuning function;

[0011] S4: Based on the self-adjusting performance control method of the multi-manipulator system and the sliding mode control principle, a sliding mode control method for the multi-manipulator system with flexible constraints is designed;

[0012] S5: Based on the sliding mode control method and adaptive dynamic programming of the multi-manipulator system, design the optimal sliding mode controller of the multi-manipulator system and complete the optimal sliding mode control of the multi-manipulator adaptive dynamic programming.

[0013] Furthermore, the 2-DOF dynamic model in S1 is:

[0014]

[0015] For simplicity, let The expression is rewritten as:

[0016]

[0017] Let H ti Represents the matrix H i The diagonal matrix of , we get:

[0018] H i =H ti +H Ti

[0019] therefore, Rewritten as:

[0020]

[0021] Among them, M i (·) is the inertia matrix function of the i-th manipulator, C i (·) is the vector function of the centripetal force and Coriolis force of the i-th manipulator, G i (·) is the gravity vector function of the i-th manipulator, q i is the position vector of the i-th robotic arm, is the velocity vector of the i-th robotic arm, is the acceleration vector of the i-th robotic arm, τ i is the input torque vector of the i-th manipulator, d is the unknown bounded disturbance vector, Ψ i =H Ti τ i +H i D i +H i d;

[0022]

[0023] in is the inertia matrix of the i-th robotic arm, C i is the centripetal force and Coriolis force vector of the i-th manipulator, is the velocity vector of the i-th manipulator in the first degree of freedom, q i,2 is the position vector of the i-th manipulator in the second degree of freedom, is the velocity vector of the i-th manipulator in the second degree of freedom, J i,1 is the moment of inertia of the i-th manipulator in the first degree of freedom, J i,2 is the moment of inertia of the i-th manipulator in the second degree of freedom, m i,1 is the mass of the i-th manipulator in the first degree of freedom, m i,2is the mass of the i-th manipulator in the second degree of freedom, R i,1 is the center of mass of the i-th manipulator in the first degree of freedom, R i,2 is the center of mass of the i-th manipulator in the second degree of freedom, r i,1 is the length of the i-th manipulator in the first degree of freedom.

[0024] Furthermore, the 2-DOF position error of the multi-manipulator system is defined in S2, including the following sub-steps:

[0025] S21: Use a directed graph χ = (K, E, A) to represent the relationship between the manipulators in exchanging state information;

[0026] A=[a ij ]∈R n×n

[0027] when a ij =1 means that robot i can receive information from robot j, otherwise a ij =0;

[0028] K={k1,k2,...,k n}

[0029] Among them, A is the communication between the manipulators, R represents the set of real numbers, R n×n is an n×n dimensional Euclidean space, n is a constant, K is the set of followers, k1, k2,…, k n is the number of followers, E is the position error;

[0030] S22: Based on the relationship between the state information exchanged between the manipulators, the 2-DOF position error of the multi-manipulator system is obtained as:

[0031]

[0032] Its derivative is:

[0033]

[0034] Among them, E i,s is the position error of the i-th manipulator in the s-th degree of freedom, q i,s is the position vector of the i-th manipulator in the s-th degree of freedom, q j,s is the position vector of the j-th manipulator in the s-th degree of freedom, b i,s Indicates that the i-th manipulator receives information from the leader in the s-th degree of freedom. When the i-th manipulator can receive information from the leader in the s-th degree of freedom, b i,s =1, otherwise b i,s =0,q L,sis the position of the reference manipulator in the sth degree of freedom, N is the total number of manipulators, For E i,s The first derivative of a ij,s When the i-th manipulator can receive information from the j-th manipulator in the s-th degree of freedom, a ij,s =1, otherwise a ij,s =0, q i,s The first derivative of q j,s The first derivative of q L,s The first derivative of .

[0035] Furthermore, the step S3 includes the following sub-steps:

[0036] S31: The design specifies the time performance function as follows:

[0037]

[0038] in, is the specified time performance function, is the initial value of the specified time performance function, is the predetermined domain of the specified time performance function, T a is the specified time, and t is the current time;

[0039] S32: Design the self-adjusting function as:

[0040]

[0041] Its derivative is:

[0042]

[0043] The self-tuning parameters are designed as:

[0044]

[0045] in, is a self-adjusting function, is a constant, is the self-tuning parameter, T b For the specified time, for The first derivative of B m is the positive parameter of the design, E i (T a ) is T b Error in time;

[0046] S33: Based on the specified time performance function and the self-adjustment function, a self-adjustment performance control method for the multi-manipulator system is obtained as follows:

[0047]

[0048] Among them, γ i,s (·)Specify performance functions for self-tuning of multi-manipulator systems.

[0049] Furthermore, the sliding mode control method for the multi-manipulator system with flexible constraint capability designed in S4 is:

[0050]

[0051] in, is the sliding mode variable, a1, a2, c1, c2 and c3 are constants, sig(·) is the symbol, γ i,s is a self-tuning performance function;

[0052] Its derivative is:

[0053]

[0054] in F i,s =a1(d i +b i )H ti , is a sliding mode controller, for The first derivative of For E i,s The second derivative of for b i,s The first derivative of is the acceleration vector of the j-th robotic arm, is the acceleration vector of the reference trajectory.

[0055] Furthermore, the S5 includes the following sub-steps:

[0056] S51: Based on the sliding mode control method of the multi-manipulator system, the Bellman equation is derived using Leibniz's law:

[0057]

[0058]

[0059] in, is the gradient, r represents the integral variable, w is the attenuation coefficient, N i,s (·) is the cost function, is the partial derivative;

[0060] S52: Based on the Bellman equation, the Hamiltonian function is obtained as:

[0061]

[0062] Among them, H i,s (·) is the Hamiltonian function;

[0063] S53: Based on the Hamiltonian function and the optimal cost function, the Hamilton-Jacobi-Bellman equation is obtained;

[0064] S54: Based on the Hamilton-Jacobi-Bellman equation, by solving Define the optimal sliding mode controller;

[0065] S55: Use the critical neural network to approximate the defined optimal sliding mode controller, obtain the approximate optimal sliding mode controller, and complete the design of the optimal sliding mode controller for the multi-manipulator system.

[0066] Furthermore, the optimal cost function in S53 is:

[0067]

[0068] Among them, the superscript * indicates the optimal, τ i,s is the control rate, min is the minimum value, Ω(A) is the allowed control set for A, and r is the integral variable;

[0069] The Hamilton-Jacobi-Bellman equation is:

[0070]

[0071] Furthermore, the optimal sliding mode controller is defined in S54 as:

[0072]

[0073] Among them, Γ i,s is a positive number, is an unknown function.

[0074] Furthermore, the criticism neural network in S55 is:

[0075]

[0076] in, For unknown functions The estimated value of is the activation function vector, the superscript T is the transpose of the matrix, To criticize the neural network weight θ i,s estimated value of;

[0077] The approximately optimal sliding mode controller is:

[0078]

[0079] Furthermore, in said S55, reinforcement learning Bellman residual and experience replay technology are used to update the critic neural network in real time;

[0080] The reinforcement learning Bellman residual p i,s for:

[0081]

[0082] t1=tT w

[0083]

[0084] Among them, t1 is a certain moment in the past, T w is a positive constant, C i,s The estimated value of (·), and is the difference;

[0085] To minimize the Bellman residual of reinforcement learning, the following update rules are obtained using gradient descent, parallel learning technology, and experience replay technology:

[0086]

[0087] in, for The first derivative of 0,s is the learning rate, for In the past l The state of the moment, For p i,s In the past l The state at the moment, l is the number of indexes, and L is the total number of indexes;

[0088] according to get:

[0089]

[0090] in, for The first derivative of F i,s In the past l The state at the moment, ξ i,s is the approximation error of the cost function, ξ i,s (t1) is ξ i,s At state t1, for ξ i,s In the past state, for State at t1.

[0091] The beneficial effects of the present invention are:

[0092] (1) The present invention combines traditional specified performance functions and self-adjustment functions to design a self-adjustment performance function for multiple robotic arms, and enables the convergence boundary of the robotic arm system to be automatically adjusted according to the system state, thereby ensuring better control performance of the multiple robotic arms.

[0093] (2) The present invention designs a sliding mode control method with constraint capability based on self-adjusting performance and inverse tangent function, which not only realizes flexible performance constraints but also enhances the robustness of the multi-manipulator system.

[0094] (3) The present invention combines the gradient descent method, parallel learning technology and experience replay technology to construct a critical neural network dynamic programming controller, which not only reduces the amount of calculation but also achieves optimal control. BRIEF DESCRIPTION OF THE DRAWINGS

[0095] Figure 1 This is a flow chart of an optimal sliding mode control method for adaptive dynamic programming of multiple manipulators.

[0096] Figure 2 Schematic diagram of the interaction of two-degree-of-freedom multi-manipulators.

[0097] Figure 3 Trajectory tracking diagram of multiple robotic arms.

[0098] Figure 4 Error diagram of multiple robotic arms.

[0099] Figure 5 Estimated weight map of a 1-critic neural network for a multi-robot arm joint.

[0100] Figure 6 Estimated weight graph of a 2-critic neural network for a multi-robot arm joint.

[0101] Figure 7 This is the error diagram of the multi-manipulator in 35-300s using the method of the present invention.

[0102] Figure 8 This is the input diagram of the multi-manipulator in 35-300s under the method of the present invention.

[0103] Figure 9 This is the error diagram of multiple robotic arms in 35-300s under the existing method.

[0104] Figure 10 This is the input image of multiple robotic arms in 35-300s under the existing method. DETAILED DESCRIPTION

[0105] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.

[0106] like Figure 1 As shown, a multi-manipulator adaptive dynamic programming optimal sliding mode control method includes the following steps:

[0107] S1: Establish a 2-DOF dynamic model of the multi-manipulator system;

[0108] S2: Based on the 2-DOF dynamic model, define the 2-DOF position error of the multi-manipulator system;

[0109] S3: Based on the 2-DOF position error, design a self-tuning performance control method for a multi-manipulator system with a specified time performance function and a self-tuning function;

[0110] S4: Based on the self-adjusting performance control method of the multi-manipulator system and the sliding mode control principle, a sliding mode control method for the multi-manipulator system with flexible constraints is designed;

[0111] S5: Based on the sliding mode control method and adaptive dynamic programming of the multi-manipulator system, design the optimal sliding mode controller of the multi-manipulator system and complete the optimal sliding mode control of the multi-manipulator adaptive dynamic programming.

[0112] The 2-DOF dynamic model in S1 is:

[0113]

[0114] Among them, M i (·) is the inertia matrix function of the i-th manipulator, C i (·) is the vector function of the centripetal force and Coriolis force of the i-th manipulator, G i (·) is the gravity vector function of the i-th manipulator, q i is the position vector of the i-th robotic arm, is the velocity vector of the i-th robotic arm, is the acceleration vector of the i-th robotic arm, τ i is the input torque vector of the i-th manipulator, and d is the unknown bounded disturbance vector;

[0115] Combining the dynamic equations of the multi-manipulator, we can conclude that:

[0116]

[0117] For simplicity, let The expression can be rewritten as Let H ti Represents the matrix H iThe diagonal matrix of H i =H ti +H Ti ,therefore, Rewritten as:

[0118]

[0119] Where Ψ i =H Ti τ i +H i D i +H i d;

[0120] Among them, H Ti =H i -H ti ,Ψ i is a nonlinear function;

[0121] The dynamic model of the i-th robotic arm is:

[0122]

[0123] in P i,2 =m i,2 r i,1 R i,2 , is the inertia matrix of the i-th robotic arm, C i is the centripetal force and Coriolis force vector of the i-th manipulator, is the velocity vector of the i-th manipulator in the first degree of freedom, q i,2 is the position vector of the i-th manipulator in the second degree of freedom, is the velocity vector of the i-th manipulator in the second degree of freedom, J i,1 is the moment of inertia of the i-th manipulator in the first degree of freedom, J i,2 is the moment of inertia of the i-th manipulator in the second degree of freedom, m i,1 is the mass of the i-th manipulator in the first degree of freedom, m i,2 is the mass of the i-th manipulator in the second degree of freedom, R i,1 is the center of mass of the i-th manipulator in the first degree of freedom, R i,2 is the center of mass of the i-th manipulator in the second degree of freedom, r i,1 is the length of the i-th manipulator in the first degree of freedom.

[0124] The 2-DOF position error of the multi-manipulator system is defined in S2, including the following steps:

[0125] S21: Use a directed graph χ = (K, E, A) to represent the relationship between the manipulators in exchanging state information;

[0126] A=[a ij ]∈R n×n

[0127] when a ij =1 means that robot i can receive information from robot j, otherwise a ij =0;

[0128] K={k1,k2,...,k n}

[0129] Among them, A is the communication between the manipulators, R represents the set of real numbers, R n×n is an n×n dimensional Euclidean space, n is a constant, K is the set of followers, k1, k2,…, k n is the number of followers, E is the position error;

[0130] S22: Based on the relationship between the state information exchanged between the manipulators, the 2-DOF position error of the multi-manipulator system is obtained as:

[0131]

[0132] Its derivative is:

[0133]

[0134] Among them, E i,s is the position error of the i-th manipulator in the s-th degree of freedom, q i,s is the position vector of the i-th manipulator in the s-th degree of freedom, q j,s is the position vector of the j-th manipulator in the s-th degree of freedom, b i,s Indicates that the i-th manipulator receives information from the leader in the s-th degree of freedom. When the i-th manipulator can receive information from the leader in the s-th degree of freedom, b i,s =1, otherwise b i,s =0,q L,s is the position of the reference manipulator in the sth degree of freedom, N is the total number of manipulators, For E i,s The first derivative of a ij,s When the i-th manipulator can receive information from the j-th manipulator in the s-th degree of freedom, a ij,s =1, otherwise a ij,s =0, q i,s The first derivative of q j,sThe first derivative of q L,s The first derivative of .

[0135] The S3 includes the following steps:

[0136] S31: The design specifies the time performance function as follows:

[0137]

[0138] in, is the specified time performance function, is the initial value of the specified time performance function, is the predetermined domain of the specified time performance function, T a is the specified time, and t is the current time;

[0139] S32: Design the self-adjusting function as:

[0140]

[0141] Its derivative is:

[0142]

[0143] The self-tuning parameters are designed as:

[0144]

[0145] in, is a self-adjusting function, is a constant, is the self-tuning parameter, T b For the specified time, for The first derivative of B m is the positive parameter of the design, E i (T a ) is T b Error in time;

[0146] S33: Based on the specified time performance function and the self-adjustment function, a self-adjustment performance control method for the multi-manipulator system is obtained as follows:

[0147]

[0148] Among them, γ i,s (·)Specify performance functions for self-tuning of multi-manipulator systems.

[0149] The sliding mode control method of the multi-manipulator system with flexible constraint capability designed in S4 is:

[0150]

[0151] in is the sliding mode variable, a1, a2, c1, c2 and c3 are constants, sig(·) is the symbol, γ i,s is a self-tuning performance function;

[0152] Its derivative is:

[0153]

[0154] in F i,s =a1(d i +b i )H ti , is a sliding mode controller, for The first derivative of For E i,s The second derivative of for b i,s The first derivative of is the acceleration vector of the j-th robotic arm, is the acceleration vector of the reference trajectory.

[0155] The S5 includes the following sub-steps:

[0156] S51: Based on the sliding mode control method of the multi-manipulator system, the Bellman equation is derived using Leibniz's law:

[0157]

[0158] in, is the gradient, r represents the integral variable, w is the attenuation coefficient, N i,s (·) is the cost function, is the partial derivative;

[0159] S52: Based on the Bellman equation, the Hamiltonian function is obtained as:

[0160]

[0161] Among them, H i,s (·) is the Hamiltonian function;

[0162] S53: Based on the Hamiltonian function and the optimal cost function, the Hamilton-Jacobi-Bellman equation is obtained;

[0163] S54: Based on the Hamilton-Jacobi-Bellman equation, by solving Define the optimal sliding mode controller;

[0164] S55: Use the critical neural network to approximate the defined optimal sliding mode controller, obtain the approximate optimal sliding mode controller, and complete the design of the optimal sliding mode controller for the multi-manipulator system.

[0165] The optimal cost function in S53 is:

[0166]

[0167] Among them, the superscript * is optimal, τ i,s is the control rate, min is the minimum value, Ω(A) is the allowed control set for A, where A∈R n represents a compact set (in mathematics, a compact set refers to a set in which any open cover can find a finite subset covering the entire set. Simply put, a compact set is a set that can be covered by a finite number of open sets. For example, a finite interval or a finite curve is a compact set), r is the integral variable;

[0168] The Hamilton-Jacobi-Bellman equation is:

[0169]

[0170] Considered to be the best By solving The optimal controller can be defined as:

[0171]

[0172] Obviously, due to There are complex nonlinearities in the system, and it is very difficult to directly calculate the optimal controller. In order to achieve the best performance control goal, is constructed as where Γ i,s Is a positive number.

[0173] Therefore, we can get:

[0174]

[0175] therefore, can be rewritten as:

[0176]

[0177] The optimal sliding mode controller defined in S54 is:

[0178]

[0179] Among them, Γ i,s is a positive number, is an unknown function.

[0180] Taking into account and are unknown functions, so a critical neural network is used to approximate them. The unknown function is approximated as and where θ i,s is the ideal neural network parameter vector, is the activation function vector, ξ i,s represents the approximation error and is a constant.

[0181] The neural network criticized in S55 is:

[0182]

[0183] in, For unknown functions The estimated value of is the activation function vector, the superscript T is the transpose of the matrix, To criticize the neural network weight θ i,s estimated value of;

[0184] The approximately optimal sliding mode controller is:

[0185]

[0186] In said S55, reinforcement learning Bellman residual and experience replay technology are used to update the critic neural network in real time;

[0187] By using reinforcement learning algorithms, we can obtain:

[0188]

[0189] Therefore, the reinforcement learning Bellman residual p i,s for:

[0190]

[0191] t1=tT w

[0192]

[0193] Among them, t1 is a certain moment in the past, T w is a positive constant, C i,s The estimated value of (·), and is the difference;

[0194] To minimize the Bellman residual of reinforcement learning, the following update rules are obtained using gradient descent, parallel learning technology, and experience replay technology:

[0195]

[0196] in, for The first derivative of 0,s is the learning rate, for In the past l The state of the moment, For p i,s In the past l The state at the moment, l is the number of indexes, and L is the total number of indexes;

[0197] according to get:

[0198]

[0199] in, for The first derivative of F i,s In the past l The state at the moment, ξ i,s is the approximation error of the cost function, ξ i,s (t1) is ξ i,s At state t1, for ξ i,s In the past state, for State at t1.

[0200] In one embodiment of the present invention, the stability of the multi-manipulator adaptive dynamic programming optimal sliding mode control method is proved based on steps S4 and S5, including the following steps:

[0201] S61: According to and V s Taking the derivative, we can get:

[0202]

[0203] Assumptions is a candidate for a continuously differentiable Lyapunov function that satisfies:

[0204]

[0205] in, It is V s Along gradient;

[0206] Therefore, there exists a positive definite matrix ψ that satisfies the following conditionsi,s ∈R n×n :

[0207]

[0208] Among them, η min (·) represents the minimum eigenvalue of the matrix.

[0209] You can get:

[0210]

[0211] S62: According to θ i,s and Taking the derivative of V1, we can get:

[0212]

[0213] According to Young's inequality, we can get:

[0214]

[0215] You can get:

[0216]

[0217]

[0218] You can get:

[0219]

[0220] S63: According to and Taking the derivative of V2, we can get:

[0221]

[0222] S64: Define V = V s +V1+V2, taking the derivative of V, we can get:

[0223]

[0224] Finally, we can get:

[0225]

[0226] According to the standard Lyapunov extension theorem, the trajectories of a multi-manipulator system can be considered to be eventually uniformly bounded.

[0227] In one embodiment of the present invention, Figure 2As shown in Figure 1, a two-DOF multi-manipulator system consisting of a leader and five followers is used to verify the effectiveness of the proposed method. Table 1 shows the initial state and parameters of the multi-manipulator system, and Table 2 shows the multi-manipulator model parameters:

[0228] Table 1 Initial state and parameters of the multi-manipulator system

[0229]

[0230] Table 2 Multi-manipulator model parameters

[0231]

[0232] from Figure 3 and Figure 4 It can be seen that the controller designed by the present invention has good tracking performance, and the tracking error of the system can converge to the predetermined area within the predetermined time without violating the constraints, highlighting the independence of the predefined control scheme to the initial conditions and controller parameters and the robustness of the system. Figure 5 and Figure 6 The learning process of the multi-arm critic neural network’s estimated weights is shown, which eventually becomes stable.

[0233] Specifically, the performance constraint method shown in the present invention enhances the flexibility of the multi-manipulator state and has better control performance. First, in the initial period, due to the properties of the inverse tangent function, it can be seen that the system error is in an unconstrained state. After a period of time, when When the system error begins to be Constraint. a Later, under the action of the self-adjusting function, the constraint boundary of the system error becomes smaller again, which makes the system error smaller again. b Finally, the constraint boundary of the system error is designed by humans. and self-adjusting The structure makes the system error enter a smaller constraint range, thereby achieving better control performance.

[0234] In order to highlight the advantages of the sliding mode control method proposed in the present invention, this embodiment will be compared with the existing sliding mode control method. The existing sliding mode control method is as follows:

[0235]

[0236] Where a2, a3 and c4 are constants.

[0237] Figure 7 and Figure 8 They respectively represent the error graph and control input graph within 35-300s under the control method of the present invention, Figure 9 and Figure 10They represent the error graph and control input graph within 35-300s under the existing control method. By comparison, it can be seen that the control method proposed in the present invention can keep the multi-manipulator system stable, while the multi-manipulator system under the existing control method has serious jitter and poor control performance. In order to comprehensively evaluate the effectiveness of the two methods, three indicators are used to quantify their performance. They are the integral of the absolute value of the error, the integral of the time multiplied by the absolute value of the error, and the integral of the square of the control input, which are defined as follows:

[0238]

[0239] In order to facilitate the comparison of the two methods, the average values ​​of the corresponding indicators of the five robotic arms are directly taken. The control performance indicators of the five robotic arms are shown in Table 3. As can be seen from Table 3, the integral of the absolute value of the error and the integral of the time multiplied by the absolute value of the error of the control method proposed by the present invention are both the smallest, indicating that the tracking error of the control method proposed by the present invention is smaller than that of the existing control method, and the dynamic tracking accuracy is higher than that of the existing method; the integral index of the square value of the control input of the control method proposed by the present invention is the smallest, indicating that the control cost of the controller proposed by the present invention is the smallest. Therefore, compared with the existing control method, the control method proposed by the present invention not only has better tracking performance, but also saves more costs.

[0240] Table 3 Comparison of three control performance indicators of the two methods

[0241]

[0242] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and it should be understood that the scope of protection of the present invention is not limited to such specific descriptions and embodiments. Those skilled in the art can make various other specific variations and combinations based on the technical teachings disclosed in the present invention without departing from the essence of the present invention, and such variations and combinations are still within the scope of protection of the invention.

Claims

1. A multi-manipulator adaptive dynamic programming optimal sliding mode control method, characterized in that: The following steps are involved: S1: Establish a 2-DOF dynamic model of the multi-manipulator system; S2: Based on the 2-DOF dynamic model, define the 2-DOF position error of the multi-manipulator system; S3: Based on the 2-DOF position error, design a self-tuning performance control method for a multi-manipulator system with a specified time performance function and a self-tuning function; S4: Based on the self-adjusting performance control method of the multi-manipulator system and the sliding mode control principle, a sliding mode control method for the multi-manipulator system with flexible constraints is designed; S5: Based on the sliding mode control method and adaptive dynamic programming of the multi-manipulator system, design the optimal sliding mode controller of the multi-manipulator system and complete the optimal sliding mode control of the multi-manipulator system through adaptive dynamic programming; The sliding mode control method of the multi-manipulator system with flexible constraint capability designed in S4 is: ; in , is the sliding mode variable, 、 、 、 and is a constant, is a symbol, is the self-tuning performance function, Indicates that the i-th manipulator receives information from the leader at the s-th degree of freedom, is the position error of the i-th manipulator in the s-th degree of freedom, for The first derivative of ; Its derivative is: ; in , , , It means that the i-th robot can receive information from the j-th robot in the s-th degree of freedom. ,on the contrary , is the total number of robotic arms, Represents a robotic arm Able to receive data from the robotic arm information, otherwise , for The first derivative of for The second derivative of for The first derivative of For the The acceleration vector of the manipulator, is the acceleration vector of the reference trajectory, is the control rate, Representation matrix The diagonal matrix of ; The S5 includes the following sub-steps: S51: Based on the sliding mode control method of the multi-manipulator system, the Bellman equation is derived using Leibniz's law: ; ; ; in, is the gradient, , represents the integration variable, is the attenuation coefficient, is the cost function, is the partial derivative, is the current time; S52: Based on the Bellman equation, the Hamiltonian function is obtained as: ; in, is the Hamiltonian function; S53: Based on the Hamiltonian function and the optimal cost function, the Hamilton-Jacobi-Bellman equation is obtained; S54: Based on the Hamilton-Jacobi-Bellman equation, by solving , defines the optimal sliding mode controller, where the superscript Indicates optimal; S55: Use the critical neural network to approximate the defined optimal sliding mode controller, obtain the approximate optimal sliding mode controller, and complete the design of the optimal sliding mode controller for the multi-manipulator system.

2. The multi-manipulator adaptive dynamic programming optimal sliding mode control method according to claim 1, characterized in that: The 2-DOF dynamic model in S1 is: ; For simplicity, let , , The expression is rewritten as: ; set up Representation matrix The diagonal matrix of , we get: ; therefore, Rewritten as: ; in, is the inertia matrix function of the i-th robotic arm, is the centripetal force and Coriolis force vector function of the i-th manipulator, is the gravity vector function of the i-th manipulator, is the position vector of the i-th robotic arm, is the velocity vector of the i-th robotic arm, is the acceleration vector of the i-th robotic arm, Input torque vector for the i-th manipulator, is the unknown bounded perturbation vector, ; ; ; in , , , is the inertia matrix of the i-th robotic arm, is the centripetal force and Coriolis force vector of the i-th manipulator, is the velocity vector of the i-th manipulator in the first degree of freedom, is the position vector of the i-th manipulator in the second degree of freedom, is the velocity vector of the i-th manipulator in the second degree of freedom, is the moment of inertia of the i-th manipulator in the first degree of freedom, is the moment of inertia of the i-th manipulator in the second degree of freedom, is the mass of the i-th manipulator in the first degree of freedom, is the mass of the i-th manipulator in the second degree of freedom, is the center of mass of the i-th manipulator in the first degree of freedom, is the center of mass of the i-th manipulator in the second degree of freedom, is the length of the i-th manipulator in the first degree of freedom.

3. The multi-manipulator adaptive dynamic programming optimal sliding mode control method according to claim 1, characterized in that: The 2-DOF position error of the multi-manipulator system is defined in S2, including the following steps: S21: Using directed graphs Represents the relationship between the manipulators in exchanging state information; ; when Represents a robotic arm Able to receive data from the robotic arm information, otherwise ; ; in, For communication between robotic arms, represents the set of real numbers, for Dimensional European space, is a constant, For the collection of followers, is the number of followers, is the position error; S22: Based on the relationship between the state information exchanged between the manipulators, the 2-DOF position error of the multi-manipulator system is obtained as: ; Its derivative is: ; in, is the position error of the i-th manipulator in the s-th degree of freedom, is the position vector of the i-th manipulator in the s-th degree of freedom, is the position vector of the j-th manipulator in the s-th degree of freedom, Indicates that the i-th manipulator receives information from the leader in the s-th degree of freedom. When the i-th manipulator can receive information from the leader in the s-th degree of freedom ,otherwise , is the position of the reference manipulator in the sth degree of freedom, is the total number of robotic arms, for The first derivative of , It means that the i-th robot can receive information from the j-th robot in the s-th degree of freedom. ,on the contrary , for The first derivative of for The first derivative of for The first derivative of .

4. The multi-manipulator adaptive dynamic programming optimal sliding mode control method according to claim 3, characterized in that: The S3 includes the following steps: S31: Design the specified time performance function as: ; in, is the specified time performance function, is the initial value of the specified time performance function, is the predetermined domain of the specified time performance function, For the specified time, is the current time; S32: Design the self-adjusting function as: ; Its derivative is: ; The self-tuning parameters are designed as: ; in, is a self-adjusting function, is a constant, is a self-tuning parameter, For the specified time, for The first derivative of is the positive parameter of the design, for Error in time; S33: Based on the specified time performance function and the self-adjustment function, a self-adjustment performance control method for the multi-manipulator system is obtained as follows: ; in, Specify performance functions for self-tuning of multi-manipulator systems.

5. The multi-manipulator adaptive dynamic programming optimal sliding mode control method according to claim 4, characterized in that: The optimal cost function in S53 is: ; Among them, the superscript Indicates the best, is the control rate, is the minimum value, For The permission control set, is the integral variable; The Hamilton-Jacobi-Bellman equation is: 。 6. The multi-manipulator adaptive dynamic programming optimal sliding mode control method according to claim 5, characterized in that: The optimal sliding mode controller defined in S54 is: ; in, is a positive number, is an unknown function.

7. The multi-manipulator adaptive dynamic programming optimal sliding mode control method according to claim 6, characterized in that: The neural network criticized in S55 is: ; in, For unknown functions The estimated value of is the activation function vector, the superscript is the transpose of the matrix, To criticize neural network weights estimated value of; The approximately optimal sliding mode controller is: 。 8. The multi-manipulator adaptive dynamic programming optimal sliding mode control method according to claim 7, characterized in that: In said S55, reinforcement learning Bellman residual and experience replay technology are used to update the critic neural network in real time; The reinforcement learning Bellman residual for: ; ; ; ; in, For a moment in the past, is a positive constant, for The estimated value of and is the difference; To minimize the Bellman residual of reinforcement learning, the following update rules are obtained using gradient descent, parallel learning technology, and experience replay technology: ; in, for The first derivative of is the learning rate, for in the past The state of the moment, for in the past The state of the moment, is the number of indexes, is the total number of indexes; according to ,get: ; ; ; in, for The first derivative of for in the past The state of the moment, is the approximation error of the cost function, for exist status, for In the past state, for exist status.

Citation Information

Patent Citations

  • Space mechanical arm coordination control method based on self-adaption dynamic programming Nash game

    CN109108964A

  • Optimal design method of adaptive fuzzy controller of multi degree of freedom manipulator system

    CN110450156A