A mechanical arm trajectory tracking method, system and electronic device

By combining the DDPG algorithm with the STSMC DDPG-STSMC control system, the problems of low accuracy and insufficient robustness in the trajectory tracking control of the robotic arm are solved, and high-precision and stable trajectory tracking effect is achieved.

CN116394258BActive Publication Date: 2026-05-19HUAQIAO UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUAQIAO UNIVERSITY
Filing Date
2023-05-18
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing robotic arm trajectory tracking control technology suffers from low trajectory tracking accuracy, severe chattering, and insufficient robustness when faced with model uncertainties and external disturbances.

Method used

By combining the Deep Deterministic Policy Gradient (DDPG) algorithm with the Superspiral Sliding Mode Controller (STSMC), a DDPG-STSMC control system is generated by constructing a superspiral sliding surface and designing a superspiral sliding mode control law. The control parameter set is then adjusted to achieve high-precision trajectory tracking of the robotic arm.

Benefits of technology

It achieves high-precision trajectory tracking of the robotic arm under conditions of uncertainty and interference, improves the robustness and stability of the system, reduces chattering, and ensures high-precision movement of the end effector.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116394258B_ABST
    Figure CN116394258B_ABST
Patent Text Reader

Abstract

The application discloses a mechanical arm trajectory tracking method and system and electronic equipment, and relates to the technical field of mechanical arm trajectory tracking. The application combines a deep deterministic policy gradient algorithm and a super-spiral sliding mode controller to establish a mechanical arm control mode. Through adjustment of parameters of multiple super-spiral sliding mode controllers by the deep deterministic policy gradient algorithm, the application can ensure that the reinforcement learning training result converges to the optimal or suboptimal value of the control parameter set of the super-spiral sliding mode controller, and ensure high-precision trajectory tracking of a mechanical arm end effector.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robotic arm trajectory tracking technology, and in particular to a robotic arm trajectory tracking method, system and electronic device. Background Technology

[0002] In recent years, the application of industrial robotic arms in various industrial production fields has been on the rise, leading to a continuous increase in the demand for improved production quality. To meet the requirements of high-precision manufacturing, robotic arms must possess excellent dynamic characteristics and tracking accuracy. However, despite progress in key technologies, China's industrial robotic arm control technology still lags behind, especially in the field of high-precision trajectory tracking. Therefore, researching and developing high-precision tracking control methods for robotic arms is of great significance for improving the level of robotic arm control technology.

[0003] Sliding mode control is a special nonlinear control method commonly used in robotic arm control. It offers advantages such as ease of design, no need for precise models, and theoretical robustness to changes in system parameters and external disturbances. However, the discontinuous switching characteristics of sliding mode control can lead to high-frequency chattering, affecting the trajectory tracking accuracy of the robotic arm and even causing instability and physical damage. Furthermore, the dynamic performance of the system largely depends on the assessment of upper bounds for system parameter errors and external disturbances, which is often difficult to perform accurately.

[0004] Meanwhile, considering the limitations of single control methods in achieving accurate trajectory tracking under model uncertainties and external disturbances, and given the increasing use of combined control methods for robotic arm trajectory tracking by scholars both domestically and internationally in recent years, sliding mode control (SMD) has become a research hotspot for high-precision trajectory tracking control of robotic arms, thanks to its advantages such as ease of design, rapid response, and strong robustness. Therefore, there is an urgent need to provide a new, high-precision robotic arm trajectory tracking method to achieve high-precision trajectory tracking control of robotic arms. Summary of the Invention

[0005] To address the aforementioned problems in the existing technology, this invention provides a robotic arm trajectory tracking method, system, and electronic device.

[0006] To achieve the above objectives, the present invention provides the following solution:

[0007] A robotic arm trajectory tracking method, comprising:

[0008] Based on the dynamic system model of the robotic arm, a super-helical sliding surface is constructed;

[0009] Based on the aforementioned superhelical sliding surface, a superhelical sliding mode control law is designed, and a set of control parameters for the superhelical sliding mode controller is generated.

[0010] Based on the deep deterministic policy gradient algorithm, the DDPG-STSMC control system is constructed by combining the superspiral sliding mode controller; the DDPG-STSMC control system includes an action network and an action value network.

[0011] The constructed DDPG-STSMC control system is used to adjust the set of control parameters to complete the tracking control of the robotic arm trajectory.

[0012] Optionally, based on the dynamic system model of the robotic arm, a super-helical sliding surface is constructed, specifically including:

[0013] Obtain the dynamic system model of the robotic arm;

[0014] Based on the dynamic system model of the robotic arm, the joint angle tracking error and angular velocity tracking error in the robotic arm tracking control process are determined to generate the joint angle tracking error set and the angular velocity tracking error set;

[0015] The super-helical sliding surface is constructed based on the joint angle tracking error set and the angular velocity tracking error set.

[0016] Optionally, the super-helical sliding surface is:

[0017]

[0018] In the formula, s is the superspiral sliding surface, Λ is the sliding surface parameter set, and e q For the joint angle tracking error set, This is the angular velocity tracking error set.

[0019] Optionally, a superhelical sliding mode control law is designed based on the superhelical sliding mode surface, and a control parameter set for the superhelical sliding mode controller is generated, specifically including:

[0020] Determine the control error between the actual joint angular velocity information set of the robotic arm and the superhelical sliding surface, as well as the control error between the actual joint angular acceleration information set of the robotic arm and the superhelical sliding surface;

[0021] The superhelical sliding mode control law is determined based on the control error between the actual joint angular velocity information set of the robotic arm and the superhelical sliding surface, the control error between the actual joint angular acceleration information set of the robotic arm and the superhelical sliding surface, and the dynamic system model of the robotic arm; the superhelical sliding mode control law is as follows:

[0022] τ=τ m +τ c +τ s

[0023] in,

[0024]

[0025] In the formula, τ m For torque control term, τ c For sliding mode control, τ s To compensate for nonlinear dynamic model errors and external disturbances, a robust super-helical sliding mode control term, k c =diag[k ci ], k c For the sliding mode control parameter set, and k represents the upper and lower limits for adjusting the parameters of the sliding mode control item. ci k is a parameter for sliding mode control. w and k r All are robust control term parameter sets, k w =diag[k wi ], k r =diag[k ri ], and These are the upper and lower limits for adjusting the robust control term parameters, k. wi and k ri All are robust control term parameters, and M0(q) is the mass inertia matrix. The control error between the actual joint angular acceleration information set of the robotic arm and the super-helical sliding surface. Let the Coriolis force and centrifugal force be vectors. G0(q) represents the control error between the actual joint angular velocity information set of the robotic arm and the super-helical sliding surface, where G0(q) is the friction force vector, s is the super-helical sliding surface, and sat(s) is the saturation function value.

[0026] The control parameter set for generating the superspiral sliding mode controller is [Λ,k c ,k w ,k r ].

[0027] Optionally, the constructed DDPG-STSMC control system is used to adjust the control parameter set to complete the tracking control of the robotic arm trajectory, specifically including:

[0028] Obtain the current state information of the robotic arm to obtain the state space information set of the robotic arm;

[0029] The control parameter set is adjusted by training the robot arm's state space information set using the deep deterministic policy gradient algorithm.

[0030] Optionally, the action network consists of an input layer, three hidden layers, and an output layer; wherein the hidden layers of the action network include fully connected layers and activation layers.

[0031] Optionally, the action value network includes an input layer, four hidden layers, and an output layer; wherein the hidden layers of the action value network include a fully connected layer, an activation layer, a stacking layer, and an activation layer.

[0032] According to specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0033] The robotic arm trajectory tracking method provided by this invention combines the Deep Deterministic Policy Gradient (DDPG) algorithm with the Super Twisting Sliding Mode Control (STSMC) to establish a robotic arm control mode. By adjusting the parameters of multiple super-twisting sliding mode controllers through the DDPG algorithm, it can ensure that the reinforcement learning training results converge to the optimal or suboptimal values ​​of the STSMC control parameter set, and guarantee high-precision trajectory tracking of the robotic arm end effector.

[0034] The present invention also provides the following implementation architecture:

[0035] A robotic arm trajectory tracking system is provided, applied to the robotic arm trajectory tracking method described above; the system includes:

[0036] The super-helical sliding surface construction module is used to construct super-helical sliding surfaces based on the dynamic system model of the robotic arm;

[0037] The control parameter set generation module is used to design a superhelical sliding mode control law based on the superhelical sliding mode surface and generate a control parameter set for the superhelical sliding mode controller.

[0038] The DDPG-STSMC control system construction module is used to construct the DDPG-STSMC control system based on the deep deterministic policy gradient algorithm and combined with the superspiral sliding mode controller; the DDPG-STSMC control system includes an action network and an action value network.

[0039] The tracking control module is used to adjust the set of control parameters using the constructed DDPG-STSMC control system to complete the tracking control of the robotic arm trajectory.

[0040] An electronic device, comprising:

[0041] Memory, used to store computer programs;

[0042] A processor, connected to the memory, is used to retrieve and execute the computer program to implement the robotic arm trajectory tracking method provided above.

[0043] A computer-readable storage medium storing a computer program for implementing the robotic arm trajectory tracking method provided above.

[0044] Since the technical effects achieved by the three implementation architectures provided above are the same as those achieved by the robotic arm trajectory tracking method provided above, they will not be described again here. Attached Figure Description

[0045] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0046] Figure 1 A flowchart of the robotic arm trajectory tracking method provided by the present invention;

[0047] Figure 2 A schematic diagram of the DDPG-STSMC control system for the robotic arm provided in an embodiment of the present invention;

[0048] Figure 3 This is a structural diagram of the DDPG-STSMC control system provided in an embodiment of the present invention;

[0049] Figure 4 The Matlab simulation model diagram of DDPG-STSMC provided in the embodiments of the present invention;

[0050] Figure 5 This is a schematic diagram of the action network structure provided in an embodiment of the present invention;

[0051] Figure 6 This is a schematic diagram of the action value network structure provided in an embodiment of the present invention;

[0052] Figure 7 The reward function graph provided for embodiments of the present invention;

[0053] Figure 8 A comparison diagram of circular trajectory tracking provided in an embodiment of the present invention;

[0054] Figure 9 This is a comparison diagram of tracking results in the x-direction without perturbation, provided by an embodiment of the present invention.

[0055] Figure 10 This is a comparison diagram of tracking results in the z-direction without perturbation, provided by an embodiment of the present invention.

[0056] Figure 11 This is a comparison diagram of tracking results in the x-direction under the presence of disturbance, provided by an embodiment of the present invention.

[0057] Figure 12 The image shows a comparison of tracking results in the z-direction under perturbation conditions, as provided in an embodiment of the present invention. Detailed Implementation

[0058] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0059] The purpose of this invention is to provide a robotic arm trajectory tracking method, system, and electronic device that can achieve high-precision trajectory tracking control of the robotic arm.

[0060] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0061] like Figure 1 As shown, the robotic arm trajectory tracking method provided by the present invention includes:

[0062] Step 1: Based on the dynamic system model of the robotic arm, construct the super-helical sliding surface. The implementation process for this step can be as follows:

[0063] Step 1-1: Obtain the dynamic system model of the robotic arm.

[0064] Step 1-2: Based on the dynamic system model of the robotic arm, determine the joint angle tracking error and angular velocity tracking error in the robotic arm tracking control process to generate the joint angle tracking error set and angular velocity tracking error set.

[0065] Steps 1-3: Construct a superhelical sliding surface based on the joint angle tracking error set and the angular velocity tracking error set. The superhelical sliding surface is as follows:

[0066]

[0067] In the formula, s is the superspiral sliding surface, Λ is the sliding surface parameter set, and e q For the joint angle tracking error set, This is the angular velocity tracking error set.

[0068] Step 2: Design the superhelical sliding mode control law based on the superhelical sliding mode surface, and generate the control parameter set for the superhelical sliding mode controller. This step includes:

[0069] Step 2-1: Determine the control error between the actual joint angular velocity information set of the robotic arm and the superhelical sliding surface, as well as the control error between the actual joint angular acceleration information set of the robotic arm and the superhelical sliding surface.

[0070] Step 2-2: Determine the superhelical sliding mode control law based on the control error between the actual joint angular velocity information set of the robotic arm and the superhelical sliding mode surface, the control error between the actual joint angular acceleration information set of the robotic arm and the superhelical sliding mode surface, and the dynamic system model of the robotic arm. The superhelical sliding mode control law is:

[0071] τ=τ m +τ c +τ s

[0072] in,

[0073]

[0074] In the formula, τ m For torque control term, τ c For sliding mode control, τ s To compensate for nonlinear dynamic model errors and external disturbances, a robust super-helical sliding mode control term, k c =diag[k ci ], and k is the upper and lower limits for adjusting the sliding mode control parameter set. w =diag[k wi ], k r =diag[k ri ], and These represent the upper and lower limits for adjusting the robust control term parameters, respectively, and M0(q) is the mass inertia matrix. The control error between the actual joint angular acceleration information set of the robotic arm and the superhelical sliding surface. Let the Coriolis force and centrifugal force be vectors. G0(q) represents the control error between the actual joint angular velocity information set of the robotic arm and the superhelical sliding surface, G0(q) is the friction force vector, s is the superhelical sliding surface, and sat(s) is the saturation function value.

[0075] Step 2-3: Generate the control parameter set of the superspiral sliding mode controller as [Λ,k c ,k w ,k r ].

[0076] Step 3: Based on the deep deterministic policy gradient algorithm, the DDPG-STSMC control system is constructed by combining it with a superspiral sliding mode controller. The DDPG-STSMC control system includes an action network and an action value network. The action network consists of one input layer, three hidden layers, and one output layer. The hidden layers of the action network include fully connected layers and activation layers. The action value network consists of one input layer, four hidden layers, and one output layer. The hidden layers of the action value network include fully connected layers, activation layers, stacking layers, and activation layers.

[0077] Step 4: Adjust the control parameter set using the constructed DDPG-STSMC control system to complete the tracking control of the robotic arm trajectory. This step can be implemented as follows:

[0078] Step 4-1: Obtain the current state information of the robotic arm and obtain the state space information set of the robotic arm.

[0079] Step 4-2: Use the state space information set of the robotic arm to train the deep deterministic policy gradient algorithm to achieve the purpose of adjusting the control parameter set.

[0080] The following example, using the tracking control of a rigid robotic arm with n rotary joints, illustrates the specific implementation process and control advantages of the robotic arm trajectory tracking method provided by the present invention.

[0081] S1. For a rigid robotic arm dynamics system with n rotary joints, establish a super-helical sliding surface s and design a super-helical sliding control law.

[0082] S11, the dynamic system model of a rigid robotic arm with n degrees of freedom and model uncertainty and external disturbance, is as follows:

[0083]

[0084] In the formula, q is the joint position vector of the robotic arm. It is the velocity vector of the robotic arm. M(q) is the acceleration vector of the robotic arm, and M(q) is the mass inertia matrix. Let G(q) be the vector of the Coriolis force and the centrifugal force, and G(q) be the vector of gravity. It is the friction force vector, τ d Let τ be the time-varying external disturbance, and τ be the torque vector acting on the joint.

[0085] Due to the complexity of the robotic arm structure, environmental changes, and measurement errors during actual operation, obtaining an accurate dynamic model is difficult. Therefore, the above-mentioned rigid robotic arm dynamic system model is written as an accurate model part and model uncertainty terms, as shown below:

[0086]

[0087] In the formula, M0(q), G0(q) are all exact model terms, where M0(q) is the mass inertia matrix. Let G0(q) be the vector of Coriolis force and centrifugal force, and G0(q) be the vector of frictional force. The rest are uncertain terms, and the total internal and external disturbances of the system are reduced to: Where, ΔE M (q) represents the uncertainty term of the mass inertia matrix. The uncertainty term for the Coriolis force and centrifugal force vectors, ΔE G (q) represents the uncertainty term of the mass inertia matrix.

[0088] Then equation (1) can be expressed as:

[0089]

[0090] S12, in the trajectory tracking control of the robotic arm, calculate the joint tracking error, establish the superhelical sliding surface s, and design its control law. Specifically, output the actual tracking angle and actual joint angular velocity of the robotic arm, subtract the actual tracking angle and actual tracking angular velocity from the desired information, and then establish the superhelical sliding surface.

[0091] Traditional sliding mode control strategies achieve high tracking accuracy when interference is not considered, but their robustness is weak in the presence of interference, making them prone to chattering and reducing trajectory tracking accuracy. Therefore, a better sliding mode control strategy is needed to improve robustness and suppress chattering.

[0092] Design the lower mold surface as follows:

[0093]

[0094] Where Λ = diag[Λ i ], i = 1…n, is the set of sliding surface parameters, Λ i Let e ​​be the parameter of the i-th sliding surface. q =[e q1 ,e q2 ,…,e qn ]and These are the joint angle tracking error set and the angular velocity tracking error set of the robotic arm, respectively, where e q =qq d q d =[q d1 ,q d2 ,…,q dn ] represents the set of expected joint angle information for the robotic arm, q = [q1, q2, ..., q n[Represents the actual joint angle information set of the robotic arm] This represents the set of expected joint angular velocities of the robotic arm. This represents the set of information on the actual joint angular velocities of the robotic arm.

[0095] S13, Design of a super-spiral sliding mode control strategy:

[0096] Define the control error between the actual joint angular acceleration information set of the robotic arm and the superhelical sliding surface as:

[0097] The control error between the actual joint angular velocity information set of the robotic arm and the superhelical sliding surface is

[0098] Combining equation (4), we can see that

[0099] Based on this, equation (3) can be written as:

[0100]

[0101] Therefore, its control law can be designed as follows:

[0102] τ=τ m +τ c +τ s (6)

[0103] in,

[0104]

[0105] In the formula, τ m For torque control term, τ c For sliding mode control, τ s To compensate for nonlinear dynamic model errors and external disturbances, a robust super-helical sliding mode control term, k c =diag[k ci ], and k is the upper and lower limits for adjusting the sliding mode control parameter set. w =diag[k wi ], k r =diag[k ri ], and These represent the upper and lower limits for adjusting the robust control term parameters. Let the STSMC control parameter set be represented as [Λ, k]. c ,k w ,k r ], take τ s The saturation function sat in the equation is as follows:

[0106]

[0107] In the formula, ρ > 0 represents the width of the sliding surface boundary layer, which can be equivalent to a slope of ρ. The choice of ρ directly affects the chattering phenomenon of the system. When ρ is small, the slope is steeper, resulting in more pronounced chattering. Conversely, when ρ is large, the slope is gentler, and the chattering phenomenon is weaker. To balance accuracy and chattering, an appropriate value of ρ needs to be selected and adjusted based on experimental data.

[0108] S2, to adjust the control parameter set [Λ,k] c ,k w ,k r Design the DDPG-STSMC control system and establish the Actor network and Critic network of the DDPG-STSMC control system.

[0109] The DDPG-STSMC control system is a control system designed based on the DDPG algorithm and incorporating the concept of a superspiral sliding mode controller. In the DDPG-STSMC control system, the DDPG algorithm provides high-performance control, while STSMC ensures robustness and adaptability. Specifically, the DDPG algorithm generates control commands by learning the action policy function and value function, while STSMC calculates the corresponding control input based on the control commands and the current state, thereby achieving control of the robotic arm. A schematic diagram of the DDPG-STSMC control system for the robotic arm is shown below. Figure 2 As shown.

[0110] The DDPG-STSMC control system uses the DDPG algorithm to adjust the STSMC controller parameters. The DDPG algorithm collects the current state information of the robotic arm, learns from experience, and adapts to the set of error information of the robotic arm. Where e T =[e x ,e z ] represents the error information set of the robotic arm's end effector trajectory, e x For the error information of the x-direction trajectory of the robotic arm's end effector, e z This provides error information for the z-direction trajectory of the robotic arm's end effector.

[0111] Subsequently, using this information from the robotic arm error information set, the DDPG algorithm is used to train and adjust the parameter vector set [Λ,k] of the STSMC in real time. c ,k w ,k rThis allows the torque vector τ acting on the joint to be adaptively adjusted, thus ensuring that the DDPG-STSMC control system maintains stability and robustness in the face of uncertainties and disturbances, thereby ensuring the stability and optimal or suboptimal performance of the control system.

[0112] In detail, the DDPG-STSMC control system includes μ(S|θ) μ Online Actor network denoted by μ'(S|θ) μ' The target Actor network is represented by Q(S,a|θ). Q The online Critic network is represented by Q'(S,a|θ) and the network is represented by Q'(S,a|θ). Q′ ) represents the target Critic network, where θ μ θ Q θ μ' and θ Q′ These are the parameters for the online Actor network, the online Critic network, the target Actor network, and the target Critic network, respectively. The direction of the Q-value gradient caused by the action policy. Indicates the state under which action μ(S) occurs. j The change in Q value caused by ) This represents the current policy gradient direction. The specific process of adjusting the DDPG algorithm of STSMC is shown in pseudocode in Table 1.

[0113] Table 1 DDPG Algorithm Flowchart

[0114]

[0115] S3. Design appropriate state values, action values, and reward functions for the DDPG-STSMC control system.

[0116] S31, the state space and motion space are designed as follows:

[0117]

[0118] In this design, Represents the information set of the robotic arm, T = [T x ,T y ] represents the set of tracking trajectory information at the end of the robotic arm, and q represents the joint angle of the robotic arm.

[0119] During the tracking process, it is hoped that the STSMC controller parameter vector set [Λ,k] can be appropriately adjusted by monitoring the joint status of the robotic arm in real time. c ,k w ,k r To improve the tracking accuracy of the robotic arm's end effector and enhance the robustness of the control system, the motion space is designed as follows:

[0120] A = [Λ,k] c ,k w ,k r (10)

[0121] S32, by combining the tracking error of the robotic arm's end effector with the joint angle settings, a reward feedback can be provided for each time step of the robotic arm's exploration. The ultimate goal of the reward is to improve the tracking accuracy of the robotic arm's end effector. Therefore, in the DDPG-STSMC control system, the reward function needs to make a reasonable evaluation based on the error between the robotic arm's end effector tracking trajectory and the desired trajectory, thereby guiding the agent's decision-making process to generate sufficient positive data, enabling the agent to be fully trained and complete the high-precision trajectory tracking task of the robotic arm. Therefore, a reward function combining arithmetic progression and Euclidean distance is designed, as follows:

[0122] reward = λ a r1+λ b r2 (11)

[0123] Where, λ a , λ b Both represent reward allocation coefficients, where r1 represents the setting of the reward for the tracking accuracy of the robotic arm's end effector, and r2 represents the safety reward for the joint angle of the robotic arm.

[0124]

[0125] With appropriate initial term a1 and tolerance d selected, the adjustment factor can be used. To meet tracking requirements and update synchronously as accuracy changes, among which... and They are the adjustment factors l b The upper and lower bounds of the adjustment. In the formula, the number of terms is... It is the upper limit of an arithmetic sequence. The tracking error e at the robotic arm's end effector... T >l b ×10 -1 When this occurs, a reward method with smaller increments should be adopted. And when the tracking error at the robotic arm's end effector satisfies e... T ∈[l b ×10 v-1 l b ×10 -v When the accuracy is continuously improved, an incremental reward mechanism based on arithmetic sequence and Euclidean distance is designed. The corresponding reward value gradually increases.

[0126] In the DDPG-STSMC control system, if the joint angle q of the robotic arm iIf the limit δ is exceeded, a penalty value is given. And this round will be stopped.

[0127]

[0128] S4. By adjusting the relevant parameters of the DDPG-STSMC control system, ensure that the system can output the optimal or suboptimal torque, thereby achieving high-precision trajectory tracking.

[0129] The DDPG-STSMC control system designed based on the above steps is trained by adjusting relevant parameters to achieve the high-precision trajectory tracking control requirements of the robotic arm.

[0130] S41, the DDPG algorithm utilizes a neural network to optimize the policy and value function. By continuously updating and iterating the neural network parameters, it eventually learns a set of weight parameters, effectively defining the network's learning strategy. The Actor network consists of one input layer and three hidden layers, with the hidden layers including fully connected layers and activation layers. The output layer uses a parameter vector a. t The Critic network consists of one input layer and four hidden layers, where each hidden layer comprises a fully connected layer, an activation layer, a stacking layer, and another activation layer. The output layer corresponds to the Q-value of the action.

[0131] S42, Set the sliding surface parameter set Λ and the sliding control parameter set k c and robust control parameter set k w and k r The adaptive range of change. Set the number of exploration rounds M and the total number of steps H per round. Set the adjustment factor l of the reward function. b The parameters, such as the number of terms v, are used for training.

[0132] S43, for the DDPG-STSMC control system, precisely configures the Actor and Critic networks through steps S41 and S42, further optimizing key DDPG parameters including training set size, maximum number of steps, termination criterion, buffer pool size, smoothing factor, discount factor, and sampling time. Based on this, training the DDPG-STSMC control system ensures high-precision trajectory tracking of the robotic arm while maintaining good robustness.

[0133] Furthermore, the control structure of the DDPG-STSMC control system is as follows: Figure 3 As shown, to verify the effectiveness of the proposed control strategy in high-precision trajectory tracking of the robotic arm, a DDPG-STSMC simulation model was built using the MATLAB / Simulink platform, as follows. Figure 4 As shown, the design of its Actor network and Critic network is as follows: Figure 5and Figure 6 As shown in Table 2, the Kinova Jaco2 robotic arm from Quanser Inc. of Canada was used as the simulation object. For ease of research, this invention applied the method to control the second and third joints of the robotic arm in the simulation experiment. In the experiment, the trajectory part adopted a circular trajectory with a radius of 0.2m. Other relevant experimental parameters were set as shown in Table 2 below.

[0134] Table 2 Experimental Parameter Settings

[0135]

[0136] During training, you may encounter issues such as the training process being too slow. To speed up the training process, we first set the maximum number of training rounds to 800 and the maximum number of steps per round to T. s / T f The range of motion is restricted to an easily trainable range using the scaling function in the Matlab library. The restricted action value is then multiplied by a decay factor to normalize the action value. After setting the parameters, the DDPG-STSMC is trained for robotic arm trajectory tracking. Figure 7 As shown, the reward function graph of the DDPG-STSMC control system is presented. In the initial learning phase, the agent's reward value is low because the exploration process is still ongoing. However, as the number of learning iterations increases, Figure 7 The reward value gradually increases and then stabilizes, indicating that the proposed reward function can effectively guide DDPG to perform reasonable training and effectively prevent deep neural networks from getting trapped in local optima.

[0137] The trajectory tracking effect of the robotic arm end effector is as follows Figure 8 As shown, the DDPG-STSMC control system can achieve high-precision trajectory tracking control of the robotic arm. Further detailed analysis of the tracking accuracy is as follows... Figure 9 As shown, without any disturbance, the tracking accuracy of the robotic arm's end effector in the x-direction can reach 10. -8 m, which has two orders of magnitude higher tracking accuracy than traditional sliding mode, such as Figure 10 As shown, the tracking accuracy of the robotic arm's end effector in the z-direction can also reach 10. - 8 m, which is an order of magnitude higher in tracking accuracy than traditional sliding mode. When perturbations are introduced, such as... Figure 11 As shown, the tracking accuracy of the robotic arm's end effector in the x-direction can reach 10. -6 m, which has two orders of magnitude higher tracking accuracy than traditional sliding mode, such as Figure 12 As shown, the tracking accuracy of the robotic arm's end effector in the z-direction can reach 10. -7The tracking accuracy is three orders of magnitude higher than that of traditional sliding mode, demonstrating that the DDPG-STSMC control system can achieve high-precision trajectory tracking control for the robotic arm. Experiments show that proper adjustment of the STSMC control parameter set and reward function can further improve the trajectory tracking accuracy of the robotic arm.

[0138] Furthermore, the invention also provides the following implementation architecture:

[0139] A robotic arm trajectory tracking system is provided, applied to the robotic arm trajectory tracking method described above. The system includes:

[0140] The Superhelical Sliding Surface Construction Module is used to construct superhelical sliding surfaces based on the dynamic system model of the robotic arm.

[0141] The control parameter set generation module is used to design the superhelical sliding mode control law based on the superhelical sliding mode surface and generate the control parameter set of the superhelical sliding mode controller.

[0142] The DDPG-STSMC control system construction module is used to construct a DDPG-STSMC control system based on the deep deterministic policy gradient algorithm and combined with a superspiral sliding mode controller. The DDPG-STSMC control system includes an action network and an action value network.

[0143] The tracking control module is used to adjust the control parameter set using the constructed DDPG-STSMC control system to complete the tracking control of the robotic arm trajectory.

[0144] An electronic device, comprising:

[0145] Memory is used to store computer programs.

[0146] The processor, connected to the memory, is used to retrieve and execute computer programs to implement the robotic arm trajectory tracking method provided above.

[0147] Furthermore, when the computer program in the aforementioned memory is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks.

[0148] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple; relevant parts can be referred to the method section.

[0149] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A method for tracking the trajectory of a robotic arm, characterized in that, include: Based on the dynamic system model of the robotic arm, a super-helical sliding surface is constructed; Based on the aforementioned superhelical sliding surface, a superhelical sliding mode control law is designed, and a set of control parameters for the superhelical sliding mode controller is generated. Based on the deep deterministic policy gradient algorithm, the DDPG-STSMC control system is constructed by combining the superspiral sliding mode controller; the DDPG-STSMC control system includes an action network and an action value network. The constructed DDPG-STSMC control system is used to adjust the set of control parameters to complete the tracking control of the robotic arm trajectory; Based on the aforementioned superspiral sliding surface, a superspiral sliding mode control law is designed, and a set of control parameters for the superspiral sliding mode controller is generated, specifically including: Determine the control error between the actual joint angular velocity information set of the robotic arm and the superhelical sliding surface, as well as the control error between the actual joint angular acceleration information set of the robotic arm and the superhelical sliding surface; The superhelical sliding mode control law is determined based on the control error between the actual joint angular velocity information set of the robotic arm and the superhelical sliding surface, the control error between the actual joint angular acceleration information set of the robotic arm and the superhelical sliding surface, and the dynamic system model of the robotic arm; the superhelical sliding mode control law is as follows: in, In the formula, For torque control items, For sliding mode control items, To compensate for nonlinear dynamic model errors and external disturbances, a super-helical sliding mode robust control term is implemented. , For the sliding mode control parameter set, , and These are the upper and lower limits for adjusting the parameters of the sliding mode control item. These are parameters for sliding mode control. and All are robust control term parameter sets. , , , , , and , These are the upper and lower limits for adjusting the robust control term parameters, respectively. and All are robust control parameters. Here is the mass inertia matrix. The control error between the actual joint angular acceleration information set of the robotic arm and the super-helical sliding surface. Let the Coriolis force and centrifugal force be vectors. The control error between the actual joint angular velocity information set of the robotic arm and the super-helical sliding surface. Let the friction force vector be... It is a super-spiral sliding surface. sat ( s ) represents the saturation function value; The control parameter set for generating the superspiral sliding mode controller is as follows .

2. The robotic arm trajectory tracking method according to claim 1, characterized in that, Based on the dynamic system model of the robotic arm, a super-helical sliding surface is constructed, specifically including: Obtain the dynamic system model of the robotic arm; Based on the dynamic system model of the robotic arm, the joint angle tracking error and angular velocity tracking error in the robotic arm tracking control process are determined to generate the joint angle tracking error set and the angular velocity tracking error set; The super-helical sliding surface is constructed based on the joint angle tracking error set and the angular velocity tracking error set.

3. The robotic arm trajectory tracking method according to claim 2, characterized in that, The superspiral sliding surface is: In the formula, It is a super-spiral sliding surface. For the sliding surface parameter set, For the joint angle tracking error set, This is the angular velocity tracking error set.

4. The robotic arm trajectory tracking method according to claim 1, characterized in that, The constructed DDPG-STSMC control system is used to adjust the control parameter set to complete the tracking control of the robotic arm trajectory, specifically including: Obtain the current state information of the robotic arm to obtain the state space information set of the robotic arm; The control parameter set is adjusted by training the robot arm's state space information set using the deep deterministic policy gradient algorithm.

5. The robotic arm trajectory tracking method according to claim 1, characterized in that, The action network consists of an input layer, three hidden layers, and an output layer; wherein the hidden layers of the action network include fully connected layers and activation layers.

6. The robotic arm trajectory tracking method according to claim 1, characterized in that, The action value network includes an input layer, four hidden layers, and an output layer; wherein, the hidden layers of the action value network include a fully connected layer, an activation layer, a stacking layer, and an activation layer.

7. A robotic arm trajectory tracking system, characterized in that, The system is applied to the robotic arm trajectory tracking method as described in any one of claims 1-6; the system comprises: The super-helical sliding surface construction module is used to construct super-helical sliding surfaces based on the dynamic system model of the robotic arm; The control parameter set generation module is used to design a superhelical sliding mode control law based on the superhelical sliding mode surface and generate a control parameter set for the superhelical sliding mode controller. The DDPG-STSMC control system construction module is used to construct the DDPG-STSMC control system based on the deep deterministic policy gradient algorithm and combined with the superspiral sliding mode controller; the DDPG-STSMC control system includes an action network and an action value network. The tracking control module is used to adjust the set of control parameters using the constructed DDPG-STSMC control system to complete the tracking control of the robotic arm trajectory.

8. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, connected to the memory, is configured to retrieve and execute the computer program to implement the robotic arm trajectory tracking method as described in any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The system contains a computer program for implementing the robotic arm trajectory tracking method as described in any one of claims 1-6.