A dynamics simulation method for a chain multi-link flexible joint robot arm based on deep reinforcement learning

By combining deep reinforcement learning and mathematical models, a dynamic simulation method for chain-type multi-link flexible joint robotic arms is established, which solves the problem of large trajectory tracking errors in existing technologies and achieves high-precision trajectory tracking results.

CN116394247BActive Publication Date: 2026-05-12NANJING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NANJING UNIV OF SCI & TECH
Filing Date
2023-04-14
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

In existing technologies, the trajectory tracking control of chain-type multi-link flexible joint robotic arms has large errors, and deep reinforcement learning methods that are detached from mathematical models are difficult to achieve high-precision trajectory tracking.

Method used

By combining deep reinforcement learning with the mathematical model of a chain-type multi-link flexible joint robotic arm, a physical model and dynamic equations are established, and a deep neural network is used to learn the joint output parameters. A loss function is then constructed for trajectory tracking, achieving high-precision trajectory tracking.

Benefits of technology

It achieves high-precision trajectory tracking, reduces sample data requirements and simulation time, and improves the accuracy of trajectory tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116394247B_ABST
    Figure CN116394247B_ABST
Patent Text Reader

Abstract

The application discloses a kind of chain multi-link flexible joint mechanical arm dynamics simulation method based on deep reinforcement learning, and it is a kind of model-driven method.First, based on deep reinforcement learning and chain multi-link flexible joint dynamics theory, then the rigid-flexible coupling dynamics model of chain multi-link mechanical arm considering joint flexibility and joint damping is established by floating coordinate system method, finally, trajectory tracking problem is studied based on the mathematical model in combination with deep learning.The chain flexible joint mechanical arm dynamics simulation method based on deep reinforcement learning has good trajectory tracking effect, and has important value for actual chain multi-link flexible joint trajectory tracking control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of multibody system dynamics control, specifically a dynamics simulation method for a chain-type multi-link flexible joint robotic arm. Background Technology

[0002] Robotic arms, as the most commonly used robots or robot components, are widely used in important fields such as industrial manufacturing, aerospace, and medical services. Considering the safety of human-robot collaboration, a new type of robotic arm—the flexible joint robotic arm—has been proposed. The accurate control of chain-link flexible joint robotic arms is a practical and important topic. High-precision trajectory tracking control technology for chain-link flexible joint robotic arms is key to evaluating their performance and is also a focus and challenge in the field of robotic arm research. Due to the uncertainty of the chain-link flexible joint robotic arm model, since the last century, scholars have begun to apply reinforcement learning and deep learning to robot motion control. Deep reinforcement learning combines the perception capabilities of deep learning with the decision-making capabilities of reinforcement learning, mitigating some of their shortcomings and leveraging their respective advantages. Currently, robotic arm control methods using deep reinforcement learning are generally data-driven, which, when separated from the mathematical model of the robotic arm, easily leads to large trajectory tracking errors. Therefore, combining the mathematical model of the chain-link flexible joint robotic arm with deep reinforcement learning methods to achieve model-based control has significant practical value. Summary of the Invention

[0003] This invention is based on the theoretical theory of deep reinforcement learning and the mathematical model of chain-link flexible joint manipulators. Its purpose is to provide a dynamic simulation method for chain-link flexible joint manipulators based on deep reinforcement learning, so as to achieve high-precision trajectory tracking of chain-link flexible joint manipulators.

[0004] The technical solution to achieve the purpose of this invention is: a dynamic simulation method for a chain-type multi-link flexible joint robotic arm based on deep reinforcement learning, comprising the following steps:

[0005] Step 1: Establish the physical model of the chain-link flexible joint manipulator and set the parameters of the chain-link flexible joint manipulator.

[0006] Step 2: In the floating coordinate system, first establish the mathematical model of the flexible joint in the chain-type multi-link flexible joint manipulator, and then use the second kind of Lagrange equation to establish the rigid-flexible coupling dynamic equation of the chain-type multi-link flexible joint manipulator.

[0007] Step 3: Set the desired end effector trajectory of the chain-link flexible joint robotic arm and divide the trajectory equally according to the total time required to complete the trajectory.

[0008] Step 4: Construct a deep reinforcement learning system. Learn the output rotation angle, motor rotation angle, output rotation angular velocity, output rotation angular acceleration, and motor output torque at each joint through a deep neural network. Substitute these parameters into the dynamic equation of the chain-type flexible joint manipulator to calculate the imbalance term. Construct a loss function based on the imbalance term and compare it with a set threshold to determine whether the training termination condition is met, thereby completing the trajectory tracking of the chain-type multi-link flexible joint manipulator.

[0009] A dynamic simulation system for a chain-link flexible joint manipulator based on deep reinforcement learning is provided. Based on the dynamic simulation method for the chain-link flexible joint manipulator, the system realizes the dynamic simulation of the chain-link flexible joint manipulator based on deep reinforcement learning.

[0010] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it performs a dynamic simulation of a chain-link flexible joint manipulator based on the aforementioned dynamic simulation method for a chain-link flexible joint manipulator, thereby achieving a dynamic simulation of the chain-link flexible joint manipulator based on deep reinforcement learning.

[0011] A computer-readable storage medium storing a computer program, which, when executed by a processor, performs a dynamic simulation of a chain-link flexible joint manipulator based on the aforementioned dynamic simulation method for a chain-link flexible joint manipulator, thereby achieving a deep reinforcement learning-based dynamic simulation of the chain-link flexible joint manipulator.

[0012] Compared with the prior art, the present invention has the following significant advantages: (1) It combines deep reinforcement learning methods with the mathematical model of a chain-type multi-link flexible joint manipulator, requiring less sample data and shorter simulation time. (2) It can combine deep learning with mathematical models to achieve model-based driving and realize high-precision trajectory tracking. Attached Figure Description

[0013] Figure 1 This is a flowchart of the present invention.

[0014] Figure 2 This is a schematic diagram of a two-link flexible joint robotic arm.

[0015] Figure 3 This is a schematic diagram of a three-bar flexible joint robotic arm.

[0016] Figure 4 Simplified model of a flexible joint.

[0017] Figure 5 This represents the desired end effector trajectory of the robotic arm.

[0018] Figure 6 The desired joint rotation angle for a two-bar flexible joint robotic arm.

[0019] Figure 7 Let represent the simulation steps for the i-th time period.

[0020] Figure 8 The error between the output angle and the desired angle of the two-bar flexible joint robotic arm.

[0021] Figure 9 The error between the output trajectory and the desired trajectory of the two-bar flexible joint robotic arm.

[0022] Figure 10 The error between the output trajectory and the desired trajectory of the three-bar flexible joint robotic arm. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. This invention provides a simulation method for a chain-type flexible joint robotic arm based on deep reinforcement learning, comprising the following steps:

[0024] like Figure 1 As shown, the present invention provides a dynamic simulation method for a chain-type multi-link flexible joint robotic arm based on deep reinforcement learning, comprising the following steps:

[0025] Step 1: Establish the physical model of the chain-link flexible joint robotic arm and set its parameters.

[0026] (1) Physical model of chain-type multi-link flexible joint robotic arm

[0027] Chain-link flexible joint robotic arms consist of flexible joints and rigid links.

[0028] The mathematical model of the flexible joint in the chain-type multi-link flexible joint robotic arm consists of two parts: a rigid reduction device, which is simplified from a harmonic reduction gear; and a flexible torsion bar, which is simplified from a series of torsion springs. The left end of the entire flexible joint is the motor side, the right end is the connecting rod side, and the middle consists of the harmonic reduction gear and the torsion spring.

[0029] The center of gravity of the rigid linkage is at the end of the link. The first link can rotate around one end (the origin) and its end is connected to a flexible joint. The second link receives the flexible joint at the end of the first link, can rotate around that flexible joint, and its end is connected to another flexible joint. The next link is connected in the same manner. The entire chain-link flexible joint robotic arm performs planar motion.

[0030] (2) Physical parameters of chain-type multi-link flexible joint robotic arm

[0031] Physical parameters of a chain-type multi-link flexible joint robotic arm: length L and mass m of the rigid link, stiffness K and damping coefficient c of the flexible joint, and moment of inertia J of the motor at the joint.

[0032] Step 2: In the floating coordinate system, first establish the mathematical model of the flexible joint in the chain-type multi-link flexible joint manipulator, and then use the second kind of Lagrange equation to establish the rigid-flexible coupling dynamic equation of the chain-type multi-link flexible joint manipulator.

[0033] (1) Mathematical model of the flexible joint of the chain-type multi-link flexible joint robotic arm

[0034] The motor reduction ratio is N, and the relationship between the motor's moment of inertia and the moment of inertia after passing through the harmonic reduction end is J = N. 2 J p The relationship between the output torque and the output torque after the harmonic deceleration end is τ. m =Nτ pm The relationship between the output angle and the output angle after the harmonic deceleration end is q = q p / N. For ease of explanation, the values ​​mentioned below are all after harmonic deceleration.

[0035] Based on the simplified model and assumptions of the above flexible joint, the flexible deformation of the joint is derived.

[0036] The torsional force of the spring in the joint is:

[0037] τ m =K(q-θ) (1)

[0038] Where, τ m —Spring torsional force and motor output torque;

[0039] K—Joint stiffness;

[0040] q—motor angular displacement;

[0041] θ — angular displacement of the connecting rod.

[0042] The spring damping force of the joint is:

[0043]

[0044] Where, τ K —Spring damping force;

[0045] c—Spring damping coefficient;

[0046] —Connecting rod angular velocity;

[0047] —Motor angular velocity.

[0048] (2) Dynamic equations of a chain-type multi-link flexible joint robotic arm

[0049] Using the second type of Lagrange equations, we establish the rigid-flexible coupling dynamic equations for a chain-type multi-link flexible joint manipulator.

[0050] The total kinetic energy E of the robotic arm k It can be divided into the rotational kinetic energy of the motor and the rotational kinetic energy of the connecting rod:

[0051]

[0052] Where J is the motor's moment of inertia matrix;

[0053] J l —Link rotational inertia matrix;

[0054] —Motor angular velocity vector;

[0055] —Link angular velocity vector.

[0056] The total potential energy E of the robotic arm p It can be divided into the elastic potential energy of flexible joints and the gravitational potential energy of connecting rods:

[0057]

[0058] Where m is the mass vector of the connecting rod;

[0059] g—acceleration due to gravity;

[0060] h—the column vector of distances from the center of gravity of the link to the zero potential energy surface.

[0061] The second kind of Lagrange equation can generally be written as:

[0062]

[0063] Considering joint damping, and E k and E p Substituting into equation (5), we obtain the dynamic equations of the chain-type multi-link flexible joint robotic arm:

[0064]

[0065] Where, τ m —Motor output torque vector;

[0066] —Motor output angular acceleration vector;

[0067] K—Joint stiffness coefficient matrix;

[0068] θ—Link angular displacement vector;

[0069] τ—Joint output torque matrix;

[0070] M(θ) — Link inertia matrix;

[0071] —Joint output angular acceleration vector;

[0072] c—Spring damping coefficient matrix;

[0073] — Vectors of centrifugal force and Coriolis force;

[0074] G(θ) — Gravity term.

[0075] Step 3: Set the desired end effector trajectory for the chain-link flexible joint robotic arm.

[0076] Define the desired end effector of the chain-linked flexible joint robotic arm and divide the trajectory equally according to the total time required to complete the trajectory.

[0077] Step 4: Construct a deep reinforcement learning system. Learn the output rotation angle, motor rotation angle, output rotation angular velocity, output rotation angular acceleration, and motor output torque at each joint through a deep neural network. Substitute these parameters into the dynamic equation of the chain-type flexible joint manipulator to calculate the imbalance term. Construct a loss function based on the imbalance term and compare it with a set threshold to determine whether the training termination condition is met, thereby completing the trajectory tracking of the chain-type multi-link flexible joint manipulator.

[0078] The deep reinforcement learning module consists of two deep neural networks (DNNs). M time points selected within each time period serve as the inputs to both DNNs. The first DNN has three layers: an input layer with 32 neurons, a hidden layer with 64 neurons, and an output layer with 32 neurons. It outputs the rotation angle of each joint, the rotation angle of the motor, the angular velocity of the motor, the angular acceleration of the motor, and the torque of the motor at each joint. The second DNN has one layer, serving as both an input and output layer, with 32 neurons. Its output value is the standard deviation.

[0079] The desired trajectory of the robotic arm's end effector is the simulation target. Based on the output information of the first DNN and the dynamic equations of the chain-linked flexible joint robotic arm, the imbalance terms of the dynamic equations, as well as the imbalance terms between the expected and output values ​​of each joint, are calculated. A Gaussian probabilistic strategy and its corresponding loss function are obtained from the imbalance terms and standard deviation (ShiyinWei, Xiaowei Jin, Hui Li. General solutions for nonlinear differential equations: a rule-based self-learning approach using deep reinforcement learning. Comput Mech 64, 1361–1374 (2019).). If the calculated value of the loss function is greater than the error threshold, the Adam optimizer is used to minimize the system's loss function, updating the parameters in the two DNNs and iterating again. If the calculated value of the loss function is less than the error threshold, the simulation value of the current time step is output.

[0080] Example

[0081] To verify the effectiveness of the present invention, a simulation calculation of a chain-type flexible joint robotic arm was performed using PyCharm. The specific method is as follows:

[0082] Step 1: In this embodiment, the chain-type flexible joint robotic arm is a two- or three-bar flexible joint robotic arm. The physical model of the two- or three-bar flexible joint robotic arm is as follows: Figure 2 , 3 As shown in Tables 1 and 2, the physical parameters of the two-bar and three-bar flexible joint robotic arm are shown in Table 1 and 2. Proceed to step 2.

[0083] Table 1 Physical parameters of the two-bar flexible joint robotic arm

[0084]

[0085] Table 2 Physical parameters of the three-bar flexible joint robotic arm

[0086]

[0087] Step 2, Simplified model of flexible joint as follows Figure 4 Based on the known physical model of the two- and three-bar flexible joint manipulator, the dynamic equations of the two- and three-bar flexible joint manipulator are established using the second kind of Lagrange equations, and then proceed to step 3.

[0088] In the dynamic equations of the two-bar flexible joint manipulator model, the specific relevant matrices are represented as follows:

[0089] Among them, the moment of inertia matrix of the link

[0090]

[0091] Centrifugal force term and Coriolis force term vector

[0092]

[0093] Gravity vector

[0094]

[0095] Motor rotational inertia matrix

[0096]

[0097] Joint output torque

[0098] τ=[0 0] T (11)

[0099] Motor output torque

[0100] τ m =[τ1 τ2] T (12)

[0101] In the formula, K(q-θ)=[K1(q1-θ1) K2(q2-θ2)] T , c1=cos(θ11), c2=cos(θ2), s1=cos(θ1), s2=cos(θ2), c 12 =cos(θ11-θ2).

[0102] In the formula, θ1 and θ2 are the angular displacements of rods 1 and 2, and q1 and q2 are the motor output angular displacements of joints 1 and 2. Let ω be the angular velocity of rods 1 and 2. The angular velocities of the motors at joints 1 and 2 are... Let be the angular accelerations of rods 1 and 2. Motor angular acceleration of joints 1 and 2.

[0103] In the dynamic equations of the three-bar flexible joint manipulator model, the specific relevant matrices are represented as follows:

[0104] Among them, the moment of inertia matrix of the link

[0105]

[0106] Centrifugal force term and Coriolis force term vector

[0107]

[0108] Gravity vector

[0109]

[0110] Motor rotational inertia matrix

[0111]

[0112] Joint output torque

[0113] τ=[0 0 0] T (17)

[0114] Motor output torque

[0115] τ m =[τ1 τ2 τ3] T (18)

[0116] In the formula:

[0117] K(q-θ)=[K1(q1-θ1) K2(q2-θ2) K3(q3-θ3)] T ,

[0118]

[0119] a3 = (m3L2 + m3L2)L1, a5=m3L3L1, a6=m3L3L2, c1=cos(θ1), c2=cos(θ2), c3=cos(θ3), s1=cos(θ1), s2=cos(θ2), s3=cos(θ3), c 12 =cos(θ1-θ2), c 13 =cos(θ1-θ3), c 23 =cos(θ2-θ3), s 12 =cos(θ1-θ2), s 13 =cos(θ1-θ3), s 23 =cos(θ2-θ3), c 123 =cos(θ1+θ2+θ3).

[0120] In the formula, θ1, θ2, and θ3 represent the angular displacements of rods 1, 2, and 3, and q1, q2, and q3 represent the motor output angular displacements of joints 1, 2, and 3. Let ω be the angular velocities of rods 1, 2, and 3. The motor angular velocities of joints 1, 2, and 3 are... Let be the angular accelerations of rods 1, 2, and 3. Motor angular acceleration of joints 1, 2, and 3.

[0121] Step 3: Based on the known dynamic equations of the two- and three-bar flexible joint manipulators, set the desired end effector trajectory of the two- and three-bar flexible joint manipulators as (x-1). 2 +(y-1) 2 =0.36, such as Figure 5 As shown in (a), the entire tracking process is expected to be completed within 2 seconds. Dividing the time into 1000 equal parts, the expected displacement of the terminal trajectory in the x and y directions in polar coordinates is:

[0122]

[0123] like Figure 5 As shown in (b) and (c), where i = 1, 2, ..., 1000. Next, proceed to step 4.

[0124] Step 4: Based on the known mathematical model and desired end effector trajectory of the two- and three-link flexible joint manipulator, deep reinforcement learning is used to simulate the two- and three-link flexible joint manipulator. The parameters of the two DNNs in the deep reinforcement learning are set as follows:

[0125] Table 3 DNN parameter settings

[0126]

[0127] In the simulation method of two-bar and three-bar flexible joint manipulators, the desired angles θ1 and θ2 of joints 1 and 2 can be calculated based on the inverse kinematics of the two-bar flexible joint manipulator, such as... Figure 6 As shown, in the simulation method for two- and three-bar flexible joint robotic arms, the specific simulation steps for the i-th time period are as follows: Figure 7 As shown.

[0128] In the simulation method of the two-bar flexible joint robotic arm, the first DNN outputs the rotation angles [μ] of the two joints respectively. x1 ,μ x2 The motor rotation angle of the two joints [μ] x3 ,μ x4 The output rotational angular velocity of the two joints [μ] y1 ,μ y2 The output rotational angular acceleration of the two joints Motor output torque at both joints [μ] τ1 ,μ τ2 ], μ x =[μ x1 ,μ x2 ,μ x3 ,μ x4 ]T , μ y =[μ y1 ,μ y2 ,0,0] T , μ τ =[0,0,μ τ1 ,μ τ2 ] T At this point, the imbalance term of the two-bar flexible joint robotic arm is:

[0129]

[0130] In the above equation, r1 and r2 represent the imbalance in the dynamic equation of the two-bar flexible joint manipulator, and the other symbols refer to the dynamic equation of the two-bar flexible joint manipulator. r3 and r4 represent the imbalance between the expected and output of the two-bar flexible joint manipulator (r3 is the sum of squared errors between the expected joint angle and the simulated output angle, and r4 is the sum of squared errors between the expected end position and the simulated output position), where x 实际 y 实际 To output the displacement of the end trajectory in the x and y directions:

[0131] x 实际 =L1cos(μ x1 )+L2cos(μ x1 +μ x2 ) (twenty one)

[0132] y 实际 =L1sin(μ x1 )+L2sin(μ x1 +μ x2 ) (twenty two)

[0133] In the simulation method of a three-bar flexible joint robotic arm, the first DNN output is the output rotation angle [μ] of the three joints. x1 ,μ x2 ,μ x3 The motor rotation angles of the three joints [μ] x4 ,μ x5 ,μ x6 The output rotational angular velocity of the three joints [μ] y1 ,μ y2 ,μ y3 The output rotational angular acceleration of the three joints Motor output torque at three joints [μ] τ1 ,μ τ2 ,μ τ3 ], μ x =[μ x1 ,μ x2 ,μx3 ,μ x4 ,μ x5 ,μ x6 ] T , μ y =[μ y1 ,μ y2 ,μ y3 ,0,0,0] T , μ τ =[0,0,0,μ1,μ τ2 ,μ τ3 ] T At this point, the imbalance equation of the three-bar flexible joint robotic arm is:

[0134]

[0135] In the above equation, r1 and r2 represent the imbalance in the dynamic equations of the three-bar flexible joint manipulator, and the remaining symbols refer to the dynamic equations of the three-bar flexible joint manipulator. r3 represents the imbalance between the expected and output of the three-bar flexible joint manipulator (r3 is the sum of squared errors between the expected end position and the simulated output position). In the equation, x 实际 y 实际 To output the displacement of the end trajectory in the x and y directions:

[0136] x 实际 =L1cos(μ x1 )+L2cos(μ x1 +μ x2 )+L3cos(μ x1 +μ x2 +μ x3 ) (twenty four)

[0137] y 实际 =L1sin(μ x1 )+L2sin(μ x1 +μ x2 )+L3sin(μ x1 +μ x2 +μ x3 (25)

[0138] Next, the appropriateness of the DNN output value is determined by the loss function, and the simulation results of the two- and three-link flexible joint robotic arms are obtained.

[0139] Table 4 Simulation results of the two-bar flexible joint robotic arm

[0140]

[0141] Table 4 shows the simulation results of the two-bar flexible joint manipulator. After being controlled by the deep reinforcement learning algorithm, the comparison between the joint angle output at each time step and the expected angle, as well as their errors, are shown in the table. Figure 8 As shown in Table 4 or... Figure 8 As can be seen, after deep reinforcement learning control, the joint angles of the robotic arm basically match the desired angles, and the absolute average error of joint 1 is approximately 2.22 × 10⁻⁶. -4 The average relative error is approximately 0.19% (rad); the average absolute error of joint 2 is approximately 3.27 × 10⁻⁶. -4 The average relative error is approximately 0.024%, measured in rad. The graphs showing the x and y displacements of the deep reinforcement learning output's terminal trajectory at each time step compared to the expected x and y displacements, along with their errors, are shown below. Figure 9 As shown in Table 4 or... Figure 9 As can be seen, the terminal trajectory output by deep reinforcement learning basically matches the expected trajectory in the x and y directions, and the average absolute error in the x direction is approximately 2.61 × 10⁻⁶. -4 The average relative error in the m direction is approximately 0.033%; the average absolute error in the y direction is approximately 2.44 × 10⁻⁶. -4 The average relative error is approximately 0.030%. Calculations show that the average error between the actual and expected trajectories is approximately 3.946 × 10⁻⁶ m. -4 m.

[0142] Table 5 Simulation results of the three-bar flexible joint robotic arm

[0143]

[0144] Table 5 shows the simulation results of the three-bar flexible joint manipulator. Under the control of the deep reinforcement learning algorithm, the output end-effector trajectory displacements in the x and y directions are compared with the expected end-effector trajectory displacements in the x and y directions, and their errors are shown in the table. Figure 10 As shown in Table 5 or... Figure 10 As can be seen, for the three-bar flexible joint robotic arm, after deep reinforcement learning control, the actual end effector trajectory in the x and y directions basically matches the expected trajectory, and the average absolute error in the x direction is approximately 2.61 × 10⁻⁶. -4 The average relative error in the m direction is approximately 0.031%; the average absolute error in the y direction is approximately 2.54 × 10⁻⁶. -4 The average relative error is approximately 0.031%. Calculations show that the average error between the actual and expected trajectories is approximately 4.010 × 10⁻⁶ m. -4 m.

[0145] Based on previous research, this invention uses PyCharm to perform simulation calculations on chain-type flexible joint robotic arms, providing a new approach to combining the mathematical model of chain-type flexible joint robotic arms with deep reinforcement learning methods, and providing researchers in this field with more comprehensive data and image data.

[0146] The simulation method for chain-link flexible joint manipulators described above can be extended to flexible joint manipulators with other numbers of links. For the sake of brevity, simulation calculations were not performed for all chain-link flexible joint manipulators. The above description is only one simulation example of the two types of chain-link flexible joint manipulators in this application, but it can also be used in other simulation examples of other chain-link flexible joint manipulators. Modifications and improvements can be made without departing from this application, and these all fall within the scope of protection of this application. Therefore, the scope of protection of this patent application should be determined by the appended claims.

Claims

1. A dynamic simulation method for a chain-type multi-link flexible joint robotic arm based on deep reinforcement learning, characterized in that, Includes the following steps: Step 1: Establish the physical model of the chain-link flexible joint manipulator and set the parameters of the chain-link flexible joint manipulator. Step 2: In the floating coordinate system, first establish the mathematical model of the flexible joint in the chain-type multi-link flexible joint manipulator, and then use the second kind of Lagrange equation to establish the rigid-flexible coupling dynamic equation of the chain-type multi-link flexible joint manipulator. Step 3: Set the desired end effector trajectory of the chain-link flexible joint robotic arm and divide the trajectory equally according to the total time required to complete the trajectory. Step 4: Construct a deep reinforcement learning system. Learn the output rotation angle, motor rotation angle, output rotation angular velocity, output rotation angular acceleration, and motor output torque at each joint through a deep neural network. Substitute these parameters into the dynamic equation of the chain-type flexible joint manipulator to calculate the imbalance term. Construct a loss function based on the imbalance term and compare it with a set threshold to determine whether the training termination condition is met, thereby completing the trajectory tracking of the chain-type multi-link flexible joint manipulator.

2. The dynamic simulation method for a chain-type multi-link flexible joint robotic arm based on deep reinforcement learning according to claim 1, characterized in that, Step 1: Establish the physical model of the chain-link flexible joint robotic arm and set its parameters. The specific method is as follows: (1) Physical model of chain-type multi-link flexible joint robotic arm A chain-type multi-link flexible joint robotic arm consists of flexible joints and rigid links; The mathematical model of the flexible joint in the chain-type multi-link flexible joint robotic arm is divided into two parts: a rigid reduction device, which is simplified from a harmonic reduction gear; and a flexible torsion bar, which is simplified from a series of torsion springs. The left end of the entire flexible joint is the motor side, the right end is the connecting rod side, and the middle is the harmonic reduction gear and the torsion spring. The center of gravity of the rigid link is at the end of the link. The first link can rotate around one end and the end is connected to a flexible joint. The second link receives the flexible joint at the end of the first link, can rotate around the flexible joint, and the end is connected to another flexible joint. The next link is connected in the same way as before. The entire chain-type multi-link flexible joint robot arm performs planar motion. (2) Physical parameters of chain-type multi-link flexible joint robotic arm Physical parameters of a chain-link flexible articulated robotic arm: length of the link ,quality Stiffness of flexible joints Damping coefficient Moment of inertia of the motor at the joint .

3. The dynamic simulation method for a chain-type multi-link flexible joint robotic arm based on deep reinforcement learning according to claim 2, characterized in that, Step 2: In a floating coordinate system, first establish the mathematical model of the flexible joints in the chain-link flexible joint manipulator, and then use the second kind of Lagrange equation to establish the rigid-flexible coupling dynamic equations of the chain-link flexible joint manipulator. The specific method is as follows: (1) Mathematical model of the flexible joint of the chain-link flexible joint robot arm Based on the above simplified model and assumptions of flexible joints, the flexible deformation of the joint is derived; The torsional force of the spring in the joint is: ; in, —Spring torsional force and motor output torque; —Joint stiffness; —Motor angular displacement; —Connecting rod angular displacement; The spring damping force of the joint is: ; in, —Spring damping force; —Spring damping coefficient; —Connecting rod angular velocity; —Motor angular velocity; (2) Dynamic equations of a chain-type multi-link flexible joint robotic arm Using the second type of Lagrange equations, the rigid-flexible coupling dynamic equations of a chain-link flexible joint manipulator are established; the total kinetic energy of the manipulator... It is divided into the rotational kinetic energy of the motor and the rotational kinetic energy of the connecting rod: ; in, —Motor rotational inertia matrix; —Link rotational inertia matrix; —Motor angular velocity vector; —Link angular velocity vector; Total potential energy of the robotic arm It is divided into the elastic potential energy of the flexible joint and the gravitational potential energy of the connecting rod: ; in, —The mass vector of the connecting rod; —Acceleration due to gravity; —The column vector of distances from the center of gravity of the link to the zero potential energy surface; The second kind of Lagrange equation is generally written as: ; Considering joint damping, and Substituting into equation (5), we obtain the dynamic equations of the chain-type multi-link flexible joint robotic arm: ; in, —Motor output torque vector; —Motor output angular acceleration vector; —Joint stiffness coefficient matrix; —Link angular displacement vector; —Joint output torque matrix; —Link inertia matrix; —Joint output angular acceleration vector; —Spring damping coefficient matrix; — Vectors of centrifugal force and Coriolis force; —Gravity term.

4. The dynamic simulation method for a chain-type multi-link flexible joint robotic arm based on deep reinforcement learning according to claim 3, characterized in that, Step 4: Construct a deep reinforcement learning system. Using a deep neural network, learn the output rotation angle, motor rotation angle, output rotation angular velocity, output rotation angular acceleration of each joint, and the motor output torque at each joint. Substitute these parameters into the dynamic equations of the chain-link flexible joint manipulator to calculate the imbalance term. Based on the imbalance term, construct a loss function and compare it with a set threshold to determine if the training termination condition is met, thereby completing the trajectory tracking of the chain-link multi-link flexible joint manipulator. Where: The deep reinforcement learning model consists of two DNNs, with each time period selecting... Each time point serves as the input to two DNNs. The first DNN has three layers: the first is the input layer with 32 neurons, the second is the hidden layer with 64 neurons, and the third is the output layer with 32 neurons. It outputs the rotation angle of each joint, the rotation angle of the motor, the angular velocity of the output rotation, the angular acceleration of the output rotation, and the output torque of the motor at each joint. The second DNN has one layer, which is the input-output layer with 32 neurons. The output value is the standard deviation. The desired trajectory of the robotic arm's end effector is the simulation target. Based on the output information of the first DNN and the dynamic equations of the chain-type flexible joint robotic arm, the imbalance terms of the dynamic equations, as well as the imbalance terms between the expected and output values ​​of each joint, are calculated. From the imbalance terms and standard deviation, a Gaussian probabilistic strategy and its corresponding loss function are obtained. If the calculated value of the loss function is greater than the error threshold, the Adam optimizer is used to minimize the system's loss function, update the parameters in the two DNNs, and iterate again. If the calculated value of the loss function is less than the error threshold, the simulation value of the current time step is output.

5. The dynamic simulation method for a chain-type multi-link flexible joint robotic arm based on deep reinforcement learning according to claim 4, characterized in that, In the simulation method of a two-bar flexible joint robotic arm, the first DNN outputs the rotation angles of the two joints. The motor rotation angle of the two joints Output rotational angular velocity of the two joints Output rotational angular acceleration of the two joints Motor output torque at both joints ,make , , , , At this point, the imbalance term of the two-bar flexible joint robotic arm is: ; In the above equation, r1 and r2 correspond to the imbalance terms in the dynamic equation of the two-bar flexible joint manipulator, and r3 and r4 correspond to the imbalance terms between the expected value and the output of the two-bar flexible joint manipulator. , This outputs the displacement of the end trajectory in the x and y directions.

6. The dynamic simulation method for a chain-type multi-link flexible joint robotic arm based on deep reinforcement learning according to claim 4, characterized in that, In the simulation method of a three-bar flexible joint robotic arm, the first DNN output is the output rotation angle of the three joints. The motor rotation angle of the three joints Output rotational angular velocity of the three joints Output rotational angular acceleration of the three joints Motor output torque at three joints ,make , , , , At this point, the imbalance term of the three-bar flexible joint robotic arm is: ; In the above equation, r1 and r2 correspond to the imbalance terms in the dynamic equation of the three-bar flexible joint manipulator, and r3 corresponds to the imbalance terms between the expected value and the output of the three-bar flexible joint manipulator. , This outputs the displacement of the end trajectory in the x and y directions.

7. A dynamic simulation system for a chain-type multi-link flexible joint robotic arm based on deep reinforcement learning, characterized in that, Based on the dynamic simulation method of the chain-type multi-link flexible joint manipulator according to any one of claims 1-6, the dynamic simulation of the chain-type multi-link flexible joint manipulator based on deep reinforcement learning is realized.

8. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, it realizes the dynamic simulation of a chain-type multi-link flexible joint manipulator based on deep reinforcement learning, based on the dynamic simulation method of the chain-type multi-link flexible joint manipulator according to any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, it realizes the dynamic simulation of a chain-type multi-link flexible joint manipulator based on deep reinforcement learning, according to the dynamic simulation method of the chain-type multi-link flexible joint manipulator according to any one of claims 1-6.