Flexible mechanical arm reinforcement learning control method based on asymmetric time-varying output constraint

By constructing a distributed parameter model and using reinforcement learning methods, the asymmetric time-varying output constraint problem of the flexible robotic arm was solved, achieving efficient and energy-saving control and improving the stability and accuracy of the system.

CN121132684APending Publication Date: 2025-12-16CHANGCHUN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511520628.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-23
Publication Date
2025-12-16

AI Technical Summary

Technical Problem

Traditional control methods struggle to achieve asymmetric time-varying output constraints for flexible robotic arms, resulting in insufficient accuracy and reliability in complex environments and poor robustness in the face of model uncertainties.

Method used

By constructing a dynamic model of a flexible robotic arm system based on a distributed parameter model, and combining Hamiltonian principle and reinforcement learning methods, Lyapunov function and backstepping method are designed to construct a control law to satisfy asymmetric time-varying output constraints, and the network parameters are updated through gradient descent.

Benefits of technology

It achieves efficient, energy-saving, and safe control of flexible robotic arms in complex environments, avoids controller overflow issues, and improves system stability and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121132684A_ABST
    Figure CN121132684A_ABST
Patent Text Reader

Abstract

The invention relates to a reinforcement learning control method for a flexible mechanical arm with asymmetric time-varying output constraints. The reinforcement learning neural network control problem of a flexible mechanical arm system with asymmetric time-varying output constraints is researched. On the basis of the Hamiltonian principle, a distributed parameter model of the system is established, an execution-judgment structure reinforcement learning controller is designed according to the model, and the system cost function and the model uncertainty are approached while the optimal performance of the system is guaranteed. A time-varying asymmetric obstacle Lyapunov function is utilized to construct output constraints of the system on angle errors and vibration, and trajectory tracking and vibration suppression of the flexible arm are achieved. And finally, verifying the effectiveness of the proposed scheme on MATLAB software through a numerical simulation experiment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a reinforcement learning control method for a flexible robotic arm with asymmetric time-varying output constraints, belonging to the field of flexible robot control systems and control methods. Background Technology

[0002] Flexible robotic arms, due to their unique compliance and flexibility, show broad application prospects in many fields. In the medical field, they can operate precisely in minimally invasive surgery, reducing damage to human tissues; in the aerospace field, they can adapt to complex space environments to perform tasks; and in industrial production, they can flexibly meet diverse production needs. However, flexible robotic arms face a series of severe challenges in actual operation. Furthermore, existing research methods all focus on lumped parameter finite-dimensional models, which, by neglecting higher-order modes, can lead to control overflow problems. Therefore, research on partial differential equation models of distributed parameters for flexible robotic arm systems is of great significance.

[0003] In complex and ever-changing work scenarios, the output state of a robotic arm needs to remain within a limited range due to the limitations of its physical structure and the external environment. For example, in medical surgery, the robotic arm's operational precision and displacement range need to be adjusted in real time according to the physiological structure of human tissue and the progress of the surgery. This requires its output to conform to asymmetric time-varying output constraints to ensure the safety and effectiveness of the surgery. In industrial production, the robotic arm's motion trajectory and force output also need to meet specific asymmetric time-varying output constraints when processing workpieces of different shapes and materials; otherwise, product quality defects will result. Traditional control methods struggle to control these asymmetric time-varying output constraints, thus affecting the accuracy and reliability of the robotic arm when performing complex tasks.

[0004] Furthermore, flexible robotic arms exhibit high nonlinearity and strong coupling, and traditional model-based control methods show poor robustness and adaptability when facing model uncertainties. Reinforcement learning control, as an intelligent control strategy, enables the robotic arm to autonomously learn and optimize its control strategy through continuous interaction with the environment. This effectively addresses the nonlinear characteristics and model uncertainties of flexible robotic arms, demonstrating better control performance in complex environments.

[0005] In summary, when flexible robotic arms face problems such as asymmetric time-varying output constraints and model uncertainties, it is necessary to study an efficient reinforcement learning optimal control method that can solve asymmetric time-varying output constraints. This is of great theoretical significance and urgent practical need for improving the control performance and reliability of flexible robotic arms in complex environments and promoting their widespread application in more fields. Summary of the Invention

[0006] To address the aforementioned issues, this paper presents a reinforcement learning control method for a flexible robotic arm with asymmetric time-varying output constraints. This method includes:

[0007] 1. Obtain the total kinetic energy of the flexible robotic arm system through kinematic analysis. Total potential energy The work done by non-conservative forces ;

[0008] 2. Based on the total kinetic energy The total potential energy and the virtual work done by the non-conservative forces A dynamic model of the flexible robotic arm system is constructed using Hamilton's principle;

[0009] 3. By performing gradient descent on the error term, the update rates of the execution network and the evaluation network are obtained;

[0010] 4. By constructing a barrier Lyapunov function, the system output state constraints are guaranteed to be within the specified range;

[0011] 5. Design a Lyapunov candidate function and combine it with the backstepping method to obtain the control law of the system;

[0012] 6. Vibration control of the flexible robotic arm system based on the aforementioned control signals.

[0013] The total kinetic energy of the system is

[0014]

[0015] in, Let be the moment of inertia of the motor. This represents the rotation angle of the motor. for The first derivative with respect to time, Point on the robotic arm Compared to arc length, for The first derivative with respect to time, The local rotation coordinate system of the robotic arm. Indicates the position of the robotic arm Place The vibration offset at any given moment; For the length of the flexible robotic arm, This refers to the density of the robotic arm.

[0016] The total potential energy of the system is

[0017]

[0018] in, For the bending stiffness of the flexible robotic arm, for right The second-order partial derivatives of .

[0019] The work done by non-conservative forces is

[0020]

[0021] in, This refers to the control torque of the flexible robotic arm.

[0022] Combining kinetic energy, potential energy, and virtual work, the dynamic model of the flexible robotic arm system can be obtained as follows:

[0023]

[0024] in, for The second derivative with respect to time, express The second derivative with respect to time, express for The fourth-order partial derivative.

[0025] The controller is:

[0026]

[0027] in, , , It is a positive constant parameter. , This is the estimation term for the neural network.

[0028] The beneficial effects of this invention are:

[0029] 1. This invention presents a reinforcement learning control method for flexible robotic arms with asymmetric time-varying output constraints. Addressing the time-varying trajectory tracking and vibration suppression problems of flexible robotic arm systems considering asymmetric time-varying output constraints and infinite-dimensional modes, a boundary vibration controller for motor input is constructed. Compared to traditional controllers, the method of this invention is based on a distributed parameter model, effectively avoiding instability phenomena such as controller overflow caused by neglecting some modes in traditional lumped parameter models, thus significantly reducing safety hazards. Furthermore, compared to existing distributed parameter control methods, the control method of this invention only requires the installation of controllers and sensors at boundary locations, thereby greatly reducing costs and energy consumption.

[0030] 2. This invention is a reinforcement learning control method for a flexible robotic arm with asymmetric time-varying output constraints. It fully considers the asymmetric time-varying output constraint problem and the optimal control problem of the system. It uses reinforcement learning control method to estimate unknown physical parameters and compensate for control errors. It can be applied to complex control systems. Therefore, the control method proposed in this invention has the characteristics of energy saving, high efficiency and safety, and is suitable for various complex systems. Attached Figure Description

[0031] To more clearly illustrate the design schemes in the prior art, the following is a brief introduction to the accompanying drawings used in the implementation scheme.

[0032] Figure 1 This is a flowchart of a reinforcement learning controller design method for a flexible robotic arm with asymmetric time-varying output constraints.

[0033] Figure 2 This is a schematic diagram of a flexible manipulator using a reinforcement learning control method based on asymmetric time-varying output constraints.

[0034] Figure 3 This is a vibration 3D diagram of a reinforcement learning control method for a flexible robotic arm with asymmetric time-varying output constraints.

[0035] Figure 4 The trajectory tracking curve is a reinforcement learning control method for a flexible robotic arm with asymmetric time-varying output constraints.

[0036] Figure 5 This is a control input curve diagram of a reinforcement learning control method for a flexible robotic arm with asymmetric time-varying output constraints.

[0037] Figure 6 This is a vibration constraint curve of the end effector of a flexible manipulator, which is a reinforcement learning control method for asymmetric time-varying output constraints.

[0038] Figure 7 This is a curve representing the trajectory tracking error constraint of a flexible robotic arm, which is a reinforcement learning control method for asymmetric time-varying output constraints. Detailed Implementation

[0039] The invention will now be described in more detail with reference to the accompanying drawings and embodiments. For ease of explanation, only the parts relevant to the invention are shown in the drawings. It should be noted that this invention is a reinforcement learning control method for a flexible robotic arm with asymmetric time-varying output constraints. It fully considers the asymmetric time-varying output constraint problem and the optimal control problem of the system. Using reinforcement learning control methods, it estimates unknown physical parameters and compensates for control errors. It can be applied to complex control systems. Therefore, the control method proposed in this invention has the characteristics of energy saving, high efficiency, and safety, and is suitable for various complex systems.

[0040] like Figure 1 As shown, this invention relates to a reinforcement learning control method for a flexible robotic arm with asymmetric time-varying output constraints. The specific implementation method and process are as follows:

[0041] 1. Establishment of the dynamic model

[0042] Total kinetic energy of the system for:

[0043] (1)

[0044] The total potential energy of the system is :

[0045] (2)

[0046] The work done by non-conservative forces is

[0047] (3)

[0048] Hamilton's principle:

[0049] (4)

[0050] in, This is the moment of inertia of the motor. This represents the rotation angle of the motor; for The first derivative with respect to time; Point on the robotic arm Compared to The arc length; for The first derivative with respect to time; The local rotation coordinate system representing the robotic arm; Indicates the position of the robotic arm Place The vibration offset at any given moment; The length of the flexible robotic arm; For robotic arm density; The bending stiffness of the flexible robotic arm; for right The second-order partial derivative; This serves as the control input for the flexible robotic arm. Furthermore, according to... It can be seen that, , , , Furthermore, for ease of calculation, let , , .

[0051] According to formula (4), combined with kinetic energy (1), potential energy (2), and virtual work (3), the system dynamics model of the flexible robotic arm can be obtained as follows:

[0052] (5)

[0053] (6)

[0054] (7)

[0055] in, for The second derivative with respect to time, express The second derivative with respect to time, express for The fourth-order partial derivative.

[0056] 2. Reinforcement Learning Control Method Design

[0057] First, define the total system error: , , , ,in, This represents the actual trajectory of the robotic arm. For the desired trajectory of the motion, For the vibration of the end effector of the flexible robotic arm, and These are virtual control variables.

[0058] To ensure the flexible robotic arm system achieves optimal control under limited energy conditions, a performance index function is designed for the robotic arm system based on a distributed parameter model:

[0059] (8)

[0060] in, , , , These are the designed positive parameters.

[0061] To calculate the performance index function, a evaluation network is used. Approximation:

[0062] (9)

[0063] in Represents input, for The weight. Let f be a basis function, where f is a basis function.

[0064] (10)

[0065] in, It is the central node. It is the width of the Gaussian function.

[0066] definition and ,in, and These are the optimal weights and the estimated weights, respectively. , , These are the basis functions for evaluating the network.

[0067] Therefore, the approximate error of the cost function can be obtained as:

[0068] (11)

[0069] Due to constant You can get :

[0070] (12)

[0071] Among them, the definition To The gradient. The network update rate is used as the evaluation criterion:

[0072] (13)

[0073] in, .

[0074] Therefore, a new evaluation of the network update rate can be obtained:

[0075] (14)

[0076] in It's the learning rate. .

[0077] A reinforcement learning controller can be obtained:

[0078] (15)

[0079] Using neural networks to estimate uncertain physical parameters of a system , and As shown below:

[0080] (16)

[0081] in, This is the approximation error.

[0082] Thus, we can obtain the controller (17):

[0083] (17)

[0084] in, and These are the optimal weights and the estimated weights, respectively. , , These are the basis functions for evaluating the network.

[0085] Define the estimation error as:

[0086] (18)

[0087] Gradient descent method is used, utilizing The designed update rate:

[0088] (19)

[0089] in, This is the learning rate of the neural network, therefore the update rate of the neural network can be obtained as follows:

[0090] (20)

[0091] Introduce a Related functions:

[0092] (twenty one)

[0093] Furthermore, define a coordinate variable:

[0094] (twenty two)

[0095] The design barrier Lyapunov function is as follows:

[0096] (twenty three)

[0097] Rewriting formula (21) yields the following formula:

[0098] (twenty four)

[0099] From formula (23), it can be concluded that when ,have and Therefore, we can obtain ,when , have and , can be obtained .

[0100] 3. Stability Analysis:

[0101] (25)

[0102] (26)

[0103] (27)

[0104] (28)

[0105] (29)

[0106] (30)

[0107] in, , and These are design parameters.

[0108] Prove that formula (25) has upper and lower bounds using formula (31):

[0109] (31)

[0110] (32)

[0111] if ,satisfy Then formula (30) holds true:

[0112] (33)

[0113] According to formulas (32), (33) and Young's inequality, we can... Rewritten as (34):

[0114] (34)

[0115] in .and satisfy We can obtain formula (35):

[0116] (35)

[0117] in , .

[0118] Therefore, it can be proved that the Lyapunov function It has upper and lower boundaries.

[0119] Next, we need to prove that the first derivative of the Lyapunov function (25) with respect to time satisfies Separate pairs Taking the derivative of each part with respect to time, we can obtain... :

[0120] (36)

[0121] Further simplification yields formula (37):

[0122] (37)

[0123] As shown in formula (38):

[0124] (38)

[0125] The derivative with respect to time is expressed as formula (39):

[0126] (39)

[0127] This can be expressed as formula (40):

[0128] (40)

[0129] also, This can be expressed as formula (41):

[0130] (41)

[0131] because

[0132] (42)

[0133] Combining formulas (41) and (42), we can obtain formula (43):

[0134] (43)

[0135] Therefore, we obtain As shown in formula (44):

[0136] (44)

[0137] in,

[0138] Multiply both sides of formula (44) We can obtain:

[0139] (45)

[0140] From formula (45), we can obtain formula (46):

[0141] (46)

[0142] This shows that It is bounded. According to formulas (28), (33) and (35), we can obtain:

[0143] (47)

[0144] Furthermore, we can obtain:

[0145] (48)

[0146] Finally, there is:

[0147] (49)

[0148] Similarly, we can conclude that:

[0149] (50)

[0150] (51)

[0151] It can be seen that formula (25) satisfies .

[0152] Finally, simulation verification was performed using the MATLAB platform. The initial state of the flexible robotic arm was set to 0, and the expected target trajectory was... The simulation time is 10 seconds, and the parameters are selected as follows: , , , , , , , The learning rates for the evaluation neural network and the execution neural network are 0.01 and 20, respectively. The constraint functions are selected as follows:

[0153] 1. The constraint function for the trajectory tracking error of the flexible robotic arm is chosen as follows:

[0154] (52)

[0155] 2. The constraint function for the vibration error at the end of the flexible robotic arm is selected as follows:

[0156] (53)

[0157] Figure 2 This is a modeling principle diagram of a flexible robotic arm in an example of the present invention. Figure 3 The image shows a 3D vibration diagram of the flexible robotic arm. The image clearly shows that the proposed controller can effectively suppress the vibration of the flexible robotic arm. Figure 4 The figure shows the trajectory tracking curve, and it can be seen from the figure that the system can track the desired signal in a relatively short time. Figure 5 To control the input curve, in Figure 5 It can be seen that the control torque gradually stabilizes after 2 seconds and converges to around 0. Figure 6 The figure shows the vibration constraint curve at the end of the flexible robotic arm. As can be seen from the figure, the end vibration is completely constrained within the set range. Figure 7 The figure shows the system trajectory tracking error constraint curve. As can be seen from the figure, the actual trajectory can quickly track the expected trajectory.

[0158] This invention relates to a reinforcement learning control method for a flexible robotic arm with asymmetric time-varying output constraints. The technical solution of this invention has been described in detail with reference to the preferred embodiments shown in the accompanying drawings. However, those skilled in the art will readily recognize that the scope of protection of this invention is not limited to these specific embodiments. Without departing from the principles of this invention, those skilled in the art can make equivalent modifications or substitutions to the relevant technical features, and the technical solutions formed after these modifications or substitutions will all be included within the scope of protection of this invention.

Claims

1. A reinforcement learning control method for a flexible robotic arm with asymmetric time-varying output constraints, characterized in that, Includes the following steps: Step 1: First, establish the dynamic equations of the system based on Hamilton's principle; second, use the evaluation network to approximate the cost function and utilize the execution network to approximate the uncertainty of the system model parameters. Step 2: Design an obstacle Lyapunov function to ensure that the constrained object is within the specified range. At the same time, use the backstepping method to design a virtual controller for correction, ensuring that the system can suppress the vibration of the flexible robotic arm while completing trajectory tracking.

2. The reinforcement learning control method for a flexible robotic arm with asymmetric time-varying output constraints according to claim 1, characterized in that, Includes the following steps: Step 1: Based on Hamilton's principle, establish the dynamic model of the flexible robotic arm system: The total kinetic energy of the system is : (1) The total potential energy of the system is : (2) The work done by non-conservative forces is : (3) Hamilton's principle: (4) in, Let be the moment of inertia of the motor. This represents the rotation angle of the motor. for The first derivative with respect to time, Point on the robotic arm Compared to arc length, for The first derivative with respect to time, The local rotation coordinate system of the robotic arm. Indicates the position of the robotic arm Place The vibration offset at time t, For the length of the flexible robotic arm, For the density of robotic arms, For the bending stiffness of the flexible robotic arm, for right The second-order partial derivative, This refers to the control torque of the flexible robotic arm; furthermore, according to... It can be seen that, , , , Furthermore, for ease of calculation, let , , ; According to formula (4), combined with kinetic energy (1), potential energy (2), and virtual work (3), the dynamic model of the flexible robotic arm system can be obtained as follows: (5) (6) (7) in, for The second derivative with respect to time, express The second derivative with respect to time, express for The fourth-order partial derivative; Step two: To achieve the optimal control objective of the flexible robotic arm system, an execution-evaluation structured reinforcement learning control method is introduced. First, the total system error is defined as follows: , , , ,in, This represents the actual trajectory of the robotic arm. For the desired trajectory of the motion, For the vibration of the end effector of the flexible robotic arm, and These are virtual control variables; The evaluation neural network is designed with the following cost function: (8) in, , , , The positive parameters are those designed. To calculate the performance index function, an approximation is made using a judging network. Furthermore, the neural network can be compact set. The approximate value of any continuous function can achieve the desired accuracy, as shown in formula (9): (9) in, To approximate the error; definition and ,in, and These are the optimal weights and the estimated weights, respectively. , , These are the basis functions for evaluating the network; Therefore, the approximate error of the cost function can be obtained as: (11) Due to constant You can get : (12) In the formula, it is defined To The gradient is used to evaluate the network update rate: (13) in, ; The new evaluation network update rate can be obtained as follows: (14) in It is used to evaluate the learning rate of a neural network. ; Design a neural network to perform the estimation error, and define the estimation error as: (15) according to Using gradient descent, the update rate of the neural network is obtained as follows: (16) in, This is the learning rate of the neural network, therefore the new update rate of the neural network can be obtained as follows: (17) Step 3, design Lyapunov candidate functions: (18) (19) (20) (21) (22) (23) in, , and These are design parameters; The system controller can be obtained through the backstepping method: (24) in, , , It is a positive constant parameter. , For neural network estimation terms; Using neural networks to estimate uncertain physical parameters of a system , and As shown below: (25) in, To approximate the error; Thus, we can obtain the controller (26): (26) Finally, the proposed controller is used to perform trajectory tracking control and vibration suppression on the flexible robotic arm.