Flexible joint mechanical arm tracking control method based on reinforcement learning and optimization control

By combining reinforcement learning and optimization control methods, the weight matrix is dynamically optimized, and the control accuracy and stability of flexible joint robot arms in dynamic environments is solved, and the precise tracking of motion trajectory and the smoothness of the control input sequence is achieved.

CN120276252APending Publication Date: 2025-07-08XIAN INST OF OPTICS & PRECISION MECHANICS CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510394258.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing flexible joint robotic arm control methods are difficult to adapt to the uncertainty and external disturbances of the dynamic environment, resulting in poor control accuracy and stability.

Method used

Using a method based on reinforcement learning and optimization control, the weight matrix is dynamically optimized to achieve precise tracking control of flexible joint robot arms by setting optimization goals and learning rates, defining reward functions and advantage functions, combining optimization controllers and augmentation systems.

Benefits of technology

It improves the control accuracy and stability of flexible joint robot arms in dynamic environments, reduces the complexity of manual adjustment, improves the utilization rate of control data and the stability of weight matrix, adapts to different environments, and realizes accurate tracking of motion trajectories and smoothness of control input sequences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276252A_ABST
    Figure CN120276252A_ABST
Patent Text Reader

Abstract

The invention relates to a flexible joint mechanical arm tracking control method, in particular to a flexible joint mechanical arm tracking control method based on reinforcement learning and optimization control, and solves the problem that an existing flexible joint mechanical arm tracking control method is difficult to adapt to uncertainty of a dynamic environment or is sensitive to disturbance. And the control precision and the stability are poor. According to the method, the optimization controller is adopted to process the output constraint, and the weight matrix of the extended cost function in the optimization controller is optimized by maximizing the accumulated reward in combination with the reinforcement learning algorithm, so that the optimization controller can still keep good performance when facing model errors and external interference, the complexity of manual adjustment is reduced, and the optimization efficiency is improved. And a better weight matrix can be found, so that a faster and more accurate control effect is achieved, and the method has universality for different environments and different control objects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a tracking control method for a flexible joint manipulator, and more particularly to a tracking control method for a flexible joint manipulator based on reinforcement learning and optimal control. Background Art

[0002] With the continuous progress of space technology and the increasing complexity of space exploration missions, spacecraft payloads are developing towards being larger and lighter. In the space environment, space exploration missions require manipulators to be able to cover a large working space to facilitate the execution of various tasks. The flexible joint manipulator can be folded and unfolded during storage and transportation, occupying a small space, and can be quickly unfolded during operation, covering the working area more effectively. The folding and unfolding characteristics and flexible joint design of the flexible joint manipulator enable it to adapt to the microgravity environment and complex space obstacles, improving the efficiency and safety of space exploration missions.

[0003] The flexible joint manipulator is a multi-variable, strongly coupled control system with some constraint conditions, and usually adopts control strategies such as model predictive controller (MPC) or linear quadratic regulator (LQR). MPC can naturally handle multi-input multi-output systems, comprehensively consider multiple control variables and process variables for control, thereby optimizing the overall performance. At the same time, MPC can directly consider various physical constraints during the optimization process. By introducing physical constraints, it ensures that the controlled object operates within a reasonable range.

[0004] As Figure 1 shown, MPC is modeled in the form of a state space equation. Using the established mathematical model, it predicts the state evolution of the controlled object over a period of time in the future. This prediction usually involves simulating the influence of a series of future control inputs to predict the response of the controlled object. The optimization problem of MPC usually includes minimizing the performance index of the controlled object (such as error, energy consumption, etc.), and considering the constraints of the control input. According to the optimized control input sequence, the first one is selected as the control instruction for the current control period and applied to the system. After the controlled object responds, observe the actual state and update the state of the controlled object using the new measurement values.

[0005] The setting of the weight matrix and the accuracy of the controlled object model information are crucial to the control performance of MPC. However, for the controlled object, especially the flexible joint manipulator, there are usually some unknown parameters; at the same time, traditional MPC relies on manual adjustment of the weight matrix, which is not only time-consuming and laborious, but also difficult to adapt to the uncertainties in the dynamic environment.

[0006] To solve the above problems, the controlled object is usually regarded as an unknown model, and the reinforcement learning algorithm is used to dynamically optimize the control input sequence of the controlled object. Commonly used reinforcement learning algorithms include Proximal Policy Optimization (PPO), Deep Deterministic Policy Gradient (DDPG), Twin Delayed Deep Deterministic Policy Gradient (TD3), and Soft Actor-Critic (SAC). However, the method of using the reinforcement learning algorithm to control the whole controlled object cannot make full use of the known parameters, requires a large amount of data in training, has a long training time, is sensitive to external environmental disturbances, and has poor control effects on flexible joint manipulators in different environments. Summary of the Invention

[0007] The object of the present invention is to solve the technical problem that the existing control methods for flexible joint manipulators are difficult to adapt to the uncertainty of dynamic environments or are sensitive to disturbances, resulting in poor control accuracy and stability, and to provide a tracking control method for flexible joint manipulators based on reinforcement learning and optimal control.

[0008] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0009] A tracking control method for a flexible joint manipulator based on reinforcement learning and optimal control, characterized by including the following steps:

[0010] Step 1: Set the optimization objective of the optimal controller, the learning rate and the number of training times of the reinforcement learning algorithm, define the reward function of the reinforcement learning algorithm based on the optimization objective, then define the cumulative reward according to the reward function, and then define the advantage function according to the cumulative reward;

[0011] Step 2: Establish a prediction model of the optimal controller according to the state equation of the flexible joint manipulator;

[0012] Step 3: Construct an augmented system for tracking the motion trajectory of the flexible joint manipulator according to the prediction model;

[0013] Step 4: Define the extended cost function of the augmented system and set the initial weight matrix;

[0014] Step 5: Take the minimization of the extended cost function of the augmented system as the objective and the augmented system as the constraint to obtain the initial optimization problem of the optimal controller;

[0015] Step 6: Solve the initial optimization problem to obtain the control input sequence, and feedback the control input sequence to the prediction model of the optimal controller established in Step 2 to perform rolling optimization on the control input sequence, then control the motion of the flexible joint manipulator according to the rolled-optimized control input sequence, and use the augmented system constructed in Step 3 to track the motion trajectory of the flexible joint manipulator to obtain the motion state and tracking state of the flexible joint manipulator in the current control cycle.

[0016] Step 7: According to the motion state and tracking state of the flexible joint manipulator in the current control period, and the cumulative reward and advantage function defined in Step 1, calculate the cumulative reward value and the advantage function value of the current control period respectively, then return to Step 4. According to the advantage function value of the current control period and the learning rate of the reinforcement learning algorithm set in Step 1, update the initial weight matrix, and then reduce the learning rate of the reinforcement learning algorithm to complete one training. Until the number of training times set in Step 1 is reached, take the initial weight matrix corresponding to the maximum cumulative reward value as the weight matrix of the extended cost function;

[0017] Step 8: Substitute the weight matrix back into the extended cost function, obtain the optimization problem of the optimal controller according to the method in Step 5, and solve it to obtain the actual control input sequence. The optimal controller realizes the tracking control of the flexible joint manipulator according to the actual control input sequence.

[0018] Further, in Step 2, the optimal controller is MPC or LQR.

[0019] Further, when the optimal controller in Step 2 is MPC, the prediction model is:

[0020] x(k + 1) = Ax(k) + Bu(k)

[0021] where x(k) and x(k + 1) respectively represent the state of the flexible joint manipulator in the k-th control period and the (k + 1)-th control period of the optimal controller, k represents the control period after time discretization, 1 ≤ k ≤ M, M is the number of training times, and each training is a control period; A represents the state matrix of the prediction model, u(k) represents the control input sequence of the optimal controller in the k-th control period, and B represents the input matrix of the prediction model.

[0022] Further, in Step 3, the augmented system is:

[0023] x a (k + 1) = A a x a (k) + B a u(k)

[0024] where,

[0025] x a (k) and x a (k + 1) respectively represent the extended state of the flexible joint manipulator in the k-th control period and the extended state in the (k + 1)-th control period of the optimal controller, A a is the state matrix, B a is the control input matrix;

[0026] x d (k) and x d (k + 1) represent the tracking states of the flexible - joint manipulator in the k - th control period and the (k + 1) - th control period of the optimal controller respectively, and A D represents the state - transition matrix of the augmented system.

[0027] Furthermore, in step 4, the extended cost function of the augmented system is:

[0028]

[0029] where J(k) represents the extended cost function of the augmented system in the k - th control period of the optimal controller, and x a (i|k) represents the predicted extended state of the flexible - joint manipulator at the i - th step in the k - th control period of the optimal controller, 1 ≤ i ≤ N, N represents the prediction horizon of the optimal controller, u(i - 1|k) represents the (i - 1) - th step of the control - input sequence in the k - th control period of the optimal controller, X(k) represents all the extended states of the flexible - joint manipulator under the prediction horizon of the optimal controller, U(k) represents all the control - input sequences under the prediction horizon of the optimal controller, and X T (k), U T (k) are the transpose matrices of X(k) and U(k) respectively, Q is the weight matrix of the flexible - joint manipulator state, R is the weight matrix of the control - input sequence of the optimal controller; |||| represents taking the norm.

[0030] Furthermore, in step 5, the initial optimization problem is:

[0031]

[0032] subject to x(i|k) = Ax(i - 1|k)+Bu(i - 1|k)

[0033] x d (i|k) = A D x d (i - 1|k)

[0034] u min ≤u(i - 1|k)≤u max

[0035] where x(i|k) represents the predicted state of the flexible - joint manipulator at the i - th step in the k - th control period of the optimal controller, x d (i|k) represents the predicted tracking state of the flexible - joint manipulator at the i - th step in the k - th control period of the optimal controller, x(i - 1|k) represents the predicted state of the flexible - joint manipulator at the (i - 1) - th step in the k - th control period of the optimal controller, xd (i - 1|k) represents the predicted tracking state of the flexible - joint manipulator at the (i - 1)-th step in the k - th control period of the optimized controller, u min is the lower limit of the control input of the optimized controller, u max is the upper limit of the control input of the optimized controller.

[0036] Furthermore, in step 1, the reinforcement learning algorithm is the PPO algorithm;

[0037] In step 1, the optimization objectives are performance metrics, stability, and robustness, or optimization efficiency.

[0038] Furthermore, in step 1, when the optimization objective is performance metrics, the reward function is:

[0039] r k =-||e k || - 0.1||u k ||

[0040] e k =x(k)-x d (k)

[0041] where r k is the reward function of the optimized controller in the k - th control period, e k is the tracking error in the k - th control period, u k is the control input sequence of the optimized controller in the k - th control period;

[0042] x(k) is the state of the flexible - joint manipulator in the k - th control period of the optimized controller, x d (k) is the tracking state of the flexible - joint manipulator in the k - th control period of the optimized controller;

[0043] Then, the cumulative reward is:

[0044]

[0045] where reward is the cumulative reward for M times of training according to the control input sequence of the optimized controller in the k - th control period, γ is the discount factor, 0 ≤ γ < 1;

[0046] The advantage function is:

[0047] advantage = reward - mean(reward past )

[0048] where advantage is the advantage function, mean(reward past) is the average cumulative reward for M times of training according to the control input sequences in the 1st to (k - 1)th control cycles of the optimized controller, reward past is the sum of the cumulative rewards for M times of training according to the control input sequences in the 1st to (k - 1)th control cycles of the optimized controller.

[0049] Furthermore, in step 7, update the initial weight matrix through the following formula:

[0050]

[0051] where Q new is the weight matrix of the state of the flexible-joint manipulator after update, R new is the weight matrix of the control input sequence after update, α is the learning rate of the reinforcement learning algorithm, is the gradient of R with respect to Q, is the gradient of R with respect to R.

[0052] Furthermore, in step 7, the maximizing cumulative reward value means: the cumulative reward value corresponding to when the reward function converges.

[0053] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0054] 1. The flexible-joint manipulator tracking control method based on reinforcement learning and optimal control provided by the present invention uses an optimized controller to handle output constraints, and uses a reinforcement learning algorithm to optimize the weight matrix of the extended cost function by maximizing the cumulative reward, enabling the optimized controller to still maintain good performance in the face of model errors and external disturbances, not only reducing the complexity of manual adjustment, but also being able to find a better weight matrix, thereby achieving a faster and more accurate control effect;

[0055] 2. The flexible-joint manipulator tracking control method based on reinforcement learning and optimal control provided by the present invention uses the control data of the optimized controller by the reinforcement learning algorithm to update the weight matrix multiple times, improving the utilization rate of the control data of the optimized controller; at the same time, by reducing the learning rate of the reinforcement learning algorithm to limit the update amplitude of the weight matrix, it can ensure that the weight matrix tends to be stable in the later stage of training, avoiding excessive changes in the weight matrix during the reinforcement learning process, thereby improving the stability of the reinforcement learning algorithm;

[0056] 3. The flexible-joint manipulator tracking control method based on reinforcement learning and optimal control provided by the present invention uses a reinforcement learning algorithm to dynamically optimize the weight matrix, which is universal for different environments and different control objects;

[0057] 4. The flexible joint manipulator tracking control method based on reinforcement learning and optimal control provided by the present invention has a reward function that encourages minimizing the tracking error and the control input sequence, enabling precise tracking of the motion trajectory of the flexible joint manipulator and maintaining the smoothness of the control input sequence.

[0058] 5. The flexible joint manipulator tracking control method based on reinforcement learning and optimal control provided by the present invention takes into account the discount of future rewards in the cumulative reward, making the reinforcement learning algorithm pay more attention to recent rewards, thereby improving the stability of the algorithm.

[0059] 6. The flexible joint manipulator tracking control method based on reinforcement learning and optimal control provided by the present invention defines the advantage function as the difference between the cumulative reward in the current control period in M trainings and the average cumulative reward in the past several control periods in M trainings, which can evaluate the pros and cons of the current control strategy relative to the past control strategy. Description of the Drawings

[0060] Figure 1 It is the control schematic diagram of the existing model predictive controller;

[0061] Figure 2 It is the control schematic diagram of the embodiment of the present invention;

[0062] Figure 3 It is the structural diagram of the flexible joint manipulator controlled by the embodiment of the present invention;

[0063] Figure 4 It is the change curve diagram of the reward function during the iteration process of step 7 in the embodiment of the present invention;

[0064] Figure 5 It is the comparison diagram of the state variable curves of the flexible joint manipulator controlled by the embodiment of the present invention. Among them, (a) is the comparison diagram of the change curve of the system state variable x1, (b) is the comparison diagram of the change curve of the system state variable x2, (c) is the comparison diagram of the change curve of the system state variable x3, and (d) is the comparison diagram of the change curve of the system state variable x4;

[0065] Figure 6 It is the error curve diagram of the system state variables of the flexible key manipulator system controlled by the embodiment of the present invention. Detailed Embodiments

[0066] The following further elaborates in detail a flexible joint manipulator tracking control method based on reinforcement learning and optimal control, taking the case where the optimization controller adopts a model predictive controller (MPC) and the reinforcement learning algorithm adopts a proximal policy optimization algorithm (PPO), in combination with the accompanying drawings and specific implementation manners. Those skilled in the art should understand that these implementation manners are only used to explain the technical principle of the present invention, and the purpose is not to limit the protection scope of the present invention.

[0067] A flexible joint manipulator tracking control method based on reinforcement learning and optimal control, as Figure 2 shown, includes the following steps:

[0068] Step 1: Set the performance index of the MPC and the learning rate and number of training times of the PPO algorithm, define the reward function of the PPO algorithm based on the performance index, then define the cumulative reward according to the reward function, and then define the advantage function according to the cumulative reward. Among them, the reward function is:

[0069] r k =-||e k ||-0.1||u k ||

[0070] e k =x(k)-x d (k)

[0071] Among them, r k is the reward function of the k-th control period of the MPC, k represents the control period after time discretization, 1≤k≤M, M is the number of training times, and each training is a control period; e k is the tracking error of the k-th control period, u k is the control input sequence of the k-th control period of the MPC;

[0072] x(k) is the state of the flexible joint manipulator in the k-th control period of the MPC, x d (k) is the tracking state of the flexible joint manipulator in the k-th control period of the MPC.

[0073] The reward function in this embodiment encourages minimizing the tracking error and the control input sequence, so as to achieve precise tracking of the motion trajectory of the flexible joint manipulator while maintaining the smoothness of the control input sequence.

[0074] To improve the stability of the PPO algorithm, the cumulative reward sets a discount factor for the rewards of future control periods, making the PPO algorithm pay more attention to recent rewards. The cumulative reward is shown in the following formula:

[0075]

[0076] where reward is the cumulative reward for M training runs with the control input sequence for the k-th control period of MPC, γ is the discount factor, and 0 ≤ γ < 1.

[0077] To enable PPO to evaluate the goodness of the current control period relative to past control periods, the advantage function is defined as the difference between the cumulative reward of the current control period in M training runs and the average cumulative reward of the past few control periods in M training runs, as shown in the following formula:

[0078] advantage = reward - mean(reward past )

[0079] where advantage is the advantage function, mean(reward past ) is the average cumulative reward for M training runs with the control input sequences for the 1st to (k - 1)-th control periods of MPC, and reward past is the sum of the cumulative rewards for M training runs with the control input sequences for the 1st to (k - 1)-th control periods of MPC.

[0080] In other embodiments, the reward function, cumulative reward, and advantage function of the PPO algorithm can also be defined after obtaining the initial optimization problem of MPC.

[0081] Step 2: Establish a prediction model for MPC based on the state equation of the flexible-joint manipulator. The state equation of the flexible-joint manipulator is:

[0082]

[0083] y = b(x)

[0084] where x is the state of the flexible-joint manipulator, u is the control input, d is the external disturbance, y is the output of the flexible-joint manipulator, f(x) is the state equation of the flexible-joint manipulator, g1(x) is the input gain, g2(x) is the disturbance gain, b(x) is the output equation of the flexible-joint manipulator, and f(x), g1(x), g2(x), and b(x) are all continuous non-linear functions of x.

[0085] Then, the prediction model is:

[0086] x(k + 1) = Ax(k) + Bu(k)

[0087] where x(k) and x(k + 1) represent the state of the flexible-joint manipulator in the k-th control period and the (k + 1)-th control period of MPC, respectively; A represents the state matrix of the prediction model, u(k) represents the control input sequence for the k-th control period of MPC, and B represents the input matrix of the prediction model.

[0088] Step 3: Construct an augmented system for tracking the motion trajectory of the flexible-joint manipulator based on the prediction model:

[0089]

[0090] where x d (k) and x d (k + 1) represent the tracking states of the flexible-joint manipulator at the k-th control period and the (k + 1)-th control period of the optimal controller respectively, and A D represents the state transition matrix of the augmented system.

[0091] Abbreviate the above formula as:

[0092] x a (k + 1) = A a x a (k) + B a u(k)

[0093] where x a (k) and x a (k + 1) represent the extended states of the flexible-joint manipulator at the k-th control period and the (k + 1)-th control period of the MPC respectively, A a is the state matrix, and B a is the control input matrix.

[0094] Step 4: Define the extended cost function of the augmented system and set the initial weight matrix. Define the extended cost function of the augmented system as a quadratic form of the predicted state and the control input, as shown in the following formula:

[0095]

[0096] where J(k) represents the extended cost function of the augmented system at the k-th control period of the MPC, x a (i|k) represents the predicted extended state of the flexible-joint manipulator at the i-th step under the k-th control period of the MPC, 1 ≤ i ≤ N, N represents the prediction horizon of the MPC, u(i - 1|k) represents the (i - 1)-th step of the control input sequence under the k-th control period of the MPC, X(k) represents all the extended states of the flexible-joint manipulator under the MPC prediction horizon, U(k) represents all the control input sequences under the MPC prediction horizon, X T (k), U T (k) are the transpose matrices of X(k) and U(k) respectively, Q is the weight matrix of the flexible-joint manipulator state, R is the weight matrix of the MPC control input sequence; |||| represents taking the norm.

[0097] Step 5: Taking the minimization of the cost function of the augmented system as the objective and the augmented system as the constraint, the initial optimization problem of MPC is obtained, as shown in the following formula:

[0098]

[0099] subject to x(i|k)=Ax(i - 1|k)+Bu(i - 1|k)

[0100] x d (i|k)=A D x d (i - 1|k)

[0101] u min ≤u(i - 1|k)≤u max

[0102] where x(i|k) represents the predicted state of the flexible - joint manipulator at the i - th step in the k - th control period of MPC, x d (i|k) represents the predicted tracking state of the flexible - joint manipulator at the i - th step in the k - th control period of MPC, x(i - 1|k) represents the predicted state of the flexible - joint manipulator at the (i - 1) - th step in the k - th control period of MPC, x d (i - 1|k) represents the predicted tracking state of the flexible - joint manipulator at the (i - 1) - th step in the k - th control period of MPC, u min is the lower limit of the MPC control input, and u max is the upper limit of the MPC control input.

[0103] Due to the influence of external disturbances and model uncertainties in the augmented model, the following steps use the PPO algorithm to optimize the weight matrix of the augmented - model extended cost function.

[0104] Step 6: Solve the initial optimization problem to obtain the control - input sequence, and feedback the control - input sequence to the prediction model of MPC established in Step 2 for rolling optimization. Then, control the motion of the flexible - joint manipulator according to the rolled - optimized control - input sequence, and use the augmented system constructed in Step 3 to track the motion trajectory of the flexible - joint manipulator to obtain the motion state and tracking state of the flexible - joint manipulator in the current control period.

[0105] Step 7: According to the motion state and tracking state of the flexible-joint manipulator in the current control period, as well as the cumulative reward and advantage function defined in Step 1, calculate the cumulative reward value and advantage function value of the current control period respectively. Then return to Step 4, update the initial weight matrix according to the advantage function value of the current control period and the learning rate of the reinforcement learning algorithm set in Step 1, and then decrease the learning rate of the reinforcement learning algorithm to complete one training. Repeat this process until the number of training times set in Step 1 is reached. Take the initial weight matrix corresponding to the maximum cumulative reward value as the weight matrix of the extended cost function. Here, the maximum cumulative reward value refers to the cumulative reward value corresponding to the convergence of the reward function.

[0106] In this step, the initial weight matrix is updated by the following formula:

[0107]

[0108] where, Q new is the weight matrix of the updated state of the flexible-joint manipulator, R new is the weight matrix of the updated control input sequence, and α is the learning rate of the PPO algorithm; is the gradient of R with respect to Q, is the gradient of R with respect to R, both of which are obtained by differentiating the extended cost function.

[0109] Step 8: Substitute the weight matrix back into the extended cost function, obtain the optimization problem of MPC according to the method in Step 5 and solve it to get the actual control input sequence. MPC realizes the tracking control of the flexible-joint manipulator according to the actual control input sequence.

[0110] To verify the application effect of the tracking control method for the flexible-joint manipulator based on reinforcement learning and optimal control provided in this embodiment, numerical simulation is carried out through Matlab. First, refine the state equation in Step 2. For example, Figure 3 in the flexible-joint manipulator shown, Motor refers to the motor, B refers to the damper, J refers to the manipulator, and K refers to the joint. Then the composition of the flexible-joint manipulator can be expressed by the following formula:

[0111]

[0112] where, b2 = -b1, a4 = -a3;

[0113] t is the control time, u(t) is the control input, θ is the angular velocity on the motor side, q is the angular velocity on the joint side; R m is the armature resistance, K m is the motor back electromagnetic constant, K t is the motor torque constant, Jeq is the total inertia of the joint, J arm is the total inertia of the robotic arm, K g is the gear ratio, K s is the joint stiffness, B eq is the equivalent damping coefficient, η g is the transmission efficiency, η m is the motor efficiency.

[0114] Secondly, define the system state variables x1, x2, x3, and x4. After organizing the above formulas, we get:

[0115]

[0116] where x1 is the angular velocity θ on the motor side, and x2 is the angular acceleration on the motor side x3 is the angular velocity q on the joint side, and x4 is the angular acceleration on the joint side

[0117] Finally, taking the final position of the flexible-joint robotic arm as the optimization objective, set it as θ = y d = 100sin(0.5*pi*t) rad, conduct Matlab numerical simulation on the method provided in this embodiment, and obtain the Matlab simulation results as shown Figures 4 - 6 It can be seen that during the training process, the reward function gradually converges, and the error between the predicted curve and the actual curve of the system state variables becomes smaller and smaller, indicating that the optimized controller is effectively learning and optimizing the control strategy, and can better adapt to the environment and obtain higher rewards.

Claims

1. A tracking control method for a flexible-joint manipulator based on reinforcement learning and optimal control, characterized in that, It includes the following steps: Step 1: Set the optimization objectives of the optimization controller, the learning rate and the number of training times of the reinforcement learning algorithm. Define the reward function of the reinforcement learning algorithm based on the optimization objectives, then define the cumulative reward according to the reward function, and then define the advantage function according to the cumulative reward; Step 2: Establish a prediction model of the optimization controller according to the state equation of the flexible-joint manipulator; Step 3: Construct an augmented system for tracking the motion trajectory of the flexible-joint manipulator according to the prediction model; Step 4: Define the extended cost function of the augmented system and set the initial weight matrix; Step 5: With the goal of minimizing the extended cost function of the augmented system and taking the augmented system as a constraint, obtain the initial optimization problem of the optimization controller; Step 6: Solve the initial optimization problem to obtain the control input sequence, and feedback the control input sequence to the prediction model of the optimization controller established in Step 2 to perform rolling optimization on the control input sequence. Then, control the motion of the flexible-joint manipulator according to the rolled-optimized control input sequence, and use the augmented system constructed in Step 3 to track the motion trajectory of the flexible-joint manipulator to obtain the motion state and tracking state of the flexible-joint manipulator in the current control period; Step 7: According to the motion state and tracking state of the flexible-joint manipulator in the current control period, as well as the cumulative reward and advantage function defined in Step 1, calculate the cumulative reward value and the advantage function value of the current control period respectively. Then return to Step 4, update the initial weight matrix according to the advantage function value of the current control period and the learning rate of the reinforcement learning algorithm set in Step 1, and then reduce the learning rate of the reinforcement learning algorithm to complete one training until the number of training times set in Step 1 is reached. Take the initial weight matrix corresponding to the maximized cumulative reward value as the weight matrix of the extended cost function; Step 8: Substitute the weight matrix back into the extended cost function, obtain the optimization problem of the optimization controller according to the method in Step 5, and solve it to obtain the actual control input sequence. The optimization controller realizes the tracking control of the flexible-joint manipulator according to the actual control input sequence.

2. The flexible joint manipulator tracking control method based on reinforcement learning and optimal control according to claim 1, characterized in that: In Step 2, the optimization controller is MPC or LQR.

3. The flexible joint manipulator tracking control method based on reinforcement learning and optimal control according to claim 2, characterized in that In Step 2, the optimization controller is MPC, and the prediction model is: x(k + 1) = Ax(k) + Bu(k) where x(k) and x(k + 1) respectively represent the state of the flexible-joint manipulator in the k-th control period and the (k + 1)-th control period of the optimization controller. k represents the control period after time discretization, 1 ≤ k ≤ M, and M is the number of training times. Each training is a control period; A represents the state matrix of the prediction model, u(k) represents the control input sequence of the optimization controller in the k-th control period, and B represents the input matrix of the prediction model.

4. The flexible joint manipulator tracking control method based on reinforcement learning and optimal control according to claim 3, characterized in that, In Step 3, the augmented system is: x a (k + 1)= A a x a (k)+ B a u(k) Among them, x a (k) and x a (k + 1) respectively represent the extended state of the flexible joint manipulator in the k-th control cycle of the optimized controller and the extended state in the (k + 1)-th control cycle. A a is the state matrix, and B a is the control input matrix; x d (k) and x d (k + 1) represent the tracking states of the flexible-joint manipulator in the k-th control period and the (k + 1)-th control period of the optimized controller, respectively. A D denotes the state transition matrix of the augmented system.

5. The tracking control method for a flexible joint manipulator based on reinforcement learning and optimal control according to claim 4, characterized in that In Step 4, the extended cost function of the augmented system is: Among them, J(k) represents the augmented cost function of the augmented system at the k-th control period of the optimal controller, and x a (i|k) represents the predicted augmented state of the flexible-joint manipulator at the i-th step under the k-th control period of the optimal controller, where 1 ≤ i ≤ N and N represents the prediction horizon of the optimal controller. u(i - 1|k) represents the (i - 1)-th step of the control input sequence under the k-th control period of the optimal controller. X(k) represents all the augmented states of the flexible-joint manipulator under the prediction horizon of the optimal controller. U(k) represents all the control input sequences under the prediction horizon of the optimal controller. X T (k) and U T (k) are the transpose matrices of X(k) and U(k) respectively. Q is the weight matrix of the flexible-joint manipulator state, and R is the weight matrix of the control input sequence of the optimal controller; || || represents taking the norm.

6. The tracking control method for a flexible joint manipulator based on reinforcement learning and optimal control according to claim 5, wherein In Step 5, the initial optimization problem is: subject to x(i|k) = Ax(i - 1|k) + Bu(i - 1|k) x d (i|k) = A D x d (i - 1|k) u min u(i - 1|k) ≤ u max where, \(x(i|k)\) represents the predicted state of the flexible-joint manipulator at the \(i\)-th step in the \(k\)-th control period of the optimal controller, \(x\) d \((i|k)\) represents the predicted tracking state of the flexible-joint manipulator at the \(i\)-th step in the \(k\)-th control period of the optimal controller, \(x(i - 1|k)\) represents the predicted state of the flexible-joint manipulator at the \((i - 1)\)-th step in the \(k\)-th control period of the optimal controller, \(x\) d \((i - 1|k)\) represents the predicted tracking state of the flexible-joint manipulator at the \((i - 1)\)-th step in the \(k\)-th control period of the optimal controller, \(u\) min is the lower limit of the control input of the optimal controller, \(u\) max is the upper limit of the control input of the optimal controller.

7. The tracking control method for a flexible-joint manipulator based on reinforcement learning and optimal control according to any one of claims 1-6, characterized in that, In Step 1, the reinforcement learning algorithm is the PPO algorithm; In Step 5, the optimization objectives are performance metrics, stability and robustness, or optimization efficiency.

8. The tracking control method for a flexible joint manipulator based on reinforcement learning and optimal control according to claim 7, wherein In step 1, the optimization objective is a performance metric, and the reward function is as follows: r k = -‖e k ‖ - 0.1‖u k ‖ e k = x(k) - x d (k) where, r k is the reward function of the k-th control period of the optimization controller, e k is the tracking error of the k-th control period, and u k is the control input sequence of the k-th control period of the optimization controller; x(k) is the state of the flexible-joint manipulator at the k-th control cycle of the optimal controller, and x d (k) is the tracking state of the flexible-joint manipulator at the k-th control cycle of the optimal controller; Then, the cumulative reward is: where reward is the cumulative reward obtained from M training sessions using the control input sequence of the k-th control cycle of the optimization controller, and γ is the discount factor with 0 ≤ γ < 1; The advantage function is: advantage=reward-mean(reward past ) Among them, advantage is the advantage function, and mean(reward past ) is the average cumulative reward obtained by performing M trainings according to the control input sequence in the 1st to (k - 1)th control cycles of the optimized controller, and reward past is the sum of the cumulative rewards obtained by performing M trainings according to the control input sequence in the 1st to (k - 1)th control cycles of the optimized controller.

9. The flexible joint manipulator tracking control method based on reinforcement learning and optimal control according to claim 8, characterized in that, In step 7, the initial weight matrix is updated using the following formula: Among them, Q new is the weight matrix of the updated state of the flexible joint manipulator, R new is the weight matrix of the updated control input sequence, α is the learning rate of the reinforcement learning algorithm, is the gradient of R with respect to Q, is the gradient of R with respect to R.

10. The flexible joint manipulator tracking control method based on reinforcement learning and optimal control according to claim 9, characterized in that, In step 7, the maximization of the cumulative reward value refers to the cumulative reward value corresponding to when the reward function converges.