Hydraulic mechanical arm PID control method based on reinforcement learning

Through the reinforcement learning-based PID control method of the hydraulic robotic arm, the evaluation neural network and the execution neural network are used to adjust the PID gain parameters in real time, which solves the problem of high-precision motion control of the hydraulic robotic arm under high load and harsh environment, and realizes high-precision tracking control and system stability.

CN120663327APending Publication Date: 2025-09-19YANSHAN UNIV

Patent Information

Application Number
CN202511083892.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-04
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

How to achieve high-precision motion control of hydraulic manipulators under high loads and harsh environments, especially to solve the impact of strong nonlinearity, parameter uncertainty and external disturbances on tracking control performance.

Method used

A PID control method for a hydraulic robotic arm based on reinforcement learning is adopted. The evaluation neural network and execution neural network architecture are used to adjust the gain parameters of the PID controller in real time. The nonlinearity and unknown disturbances in the compensation system are compensated through real-time learning and optimization to achieve high-precision tracking control.

Benefits of technology

The tracking accuracy of the hydraulic robotic arm is significantly improved, and the displacement tracking error is controlled within 0.5mm. It is suitable for heavy-load operation scenarios with high precision requirements and maintains system stability under complex working conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120663327A_ABST
    Figure CN120663327A_ABST
Patent Text Reader

Abstract

The invention discloses a reinforcement learning-based PID (Proportion Integration Differentiation) control method for a hydraulic mechanical arm, which belongs to the technical field of heavy-load hydraulic mechanical arm control and comprises the following steps of: S1, acquiring hydraulic cylinder displacement data of the hydraulic mechanical arm in an experiment process by utilizing a displacement sensor; s2, designing a PID-like controller based on the obtained hydraulic cylinder displacement signal and the hydraulic cylinder instruction displacement signal; s3, automatically adjusting a time-varying gain by using an evaluation neural network-execution neural network architecture, and correcting a gain parameter of the PID controller in real time to obtain an optimal gain parameter of the PID controller; and S4, realizing high-precision tracking control of the hydraulic cylinder based on the optimal control parameters of the PID controller. The method is used for compensating for strong nonlinearity, unknown disturbance, multi-joint coupling and the like existing in mechanical and hydraulic systems, the tracking precision of a hydraulic mechanical arm system is improved, the complex and time-consuming manual parameter adjusting process in traditional PID control is avoided, the structure is simple, and meanwhile the strong self-adaption and learning capacity is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of heavy-duty hydraulic mechanical arm control, and in particular to a hydraulic mechanical arm PID control method based on reinforcement learning. Background Art

[0002] With the rapid development of modern industrial automation and robotics, the application of industrial robots is becoming increasingly widespread, gradually penetrating a wider range of industries beyond traditional sectors such as automotive manufacturing and electronics assembly. In particular, industrial robots can replace manual labor in heavy, precise, or high-risk tasks in hazardous environments such as high temperatures, high pressures, and toxic and hazardous materials, significantly improving production efficiency and safety. Hydraulic heavy-duty robotic arms, with their high power density, strong load capacity, and excellent impact resistance, have demonstrated unique advantages in specialized fields such as the nuclear industry, ocean exploration, and emergency rescue, becoming indispensable key equipment.

[0003] Hydraulic manipulators face numerous challenges in their control accuracy when used in complex environments. Furthermore, factors such as the strong nonlinearity of the hydraulic manipulator system, parameter uncertainty, and external disturbances significantly impact the manipulator's tracking control performance. Therefore, achieving high-precision motion control of hydraulic manipulators under high loads and harsh environments has become a hot and challenging research topic. Summary of the Invention

[0004] To address the difficulty of high-precision motion control for heavy-duty, multi-degree-of-freedom hydraulic manipulators, this paper provides a reinforcement learning-based PID control method for hydraulic manipulators. This method generates an optimal control law based on an execution neural network-evaluation neural network architecture, learns and adjusts the gain parameters of the PID controller in real time, and effectively compensates for strong nonlinearities, unknown disturbances, and multi-joint coupling in the mechanical and hydraulic systems, thereby improving system tracking accuracy.

[0005] The technical solution adopted by the PID control method of a hydraulic mechanical arm based on reinforcement learning of the present invention is:

[0006] A PID control method for a hydraulic manipulator based on reinforcement learning comprises the following steps:

[0007] S1. Use the displacement sensor to obtain the displacement data of the hydraulic cylinder of the hydraulic manipulator during the experiment;

[0008] S2. designing a PID-like controller based on the acquired hydraulic cylinder displacement signal and the hydraulic cylinder command displacement signal;

[0009] S3, using the evaluation neural network-executor neural network architecture to automatically adjust the time-varying gain, correct the gain parameters of the PID controller in real time, and obtain the optimal gain parameters of the PID controller;

[0010] S4. Based on the optimal control parameters of the PID controller, high-precision tracking control of the hydraulic cylinder is achieved.

[0011] A further improvement of the technical solution of the present invention is that: step S1 specifically includes using a displacement sensor to collect the displacement data of each hydraulic cylinder of the hydraulic robotic arm in real time, feeding the collected displacement data back to the controller in real time, and storing the data through the storage module of the controller, and the displacement sensor is a high-precision wire-type displacement sensor.

[0012] A further improvement of the technical solution of the present invention is that the expression of the PID controller in step S2 is:

[0013]

[0014] Among them, k p 、k i and k d is a positive constant, k P (·), k I (·) and k D (·) is the time-varying gain; the proportional, integral, and differential gain parameters are connected using a constant coefficient ω>0 and rewritten in terms of differential gain as k d =k p / 2ω=k i / ω 2 , similarly k D (·)=k P (·) / 2ω=k I (·) / ω 2 ;

[0015]

[0016] A further improvement of the technical solution of the present invention is that in step S3, the working process of the evaluation neural network includes constructing a long-term cost function to approximate the cumulative amount of the generalized error of the hydraulic cylinder displacement and the control signal, and the long-term cost function is defined as:

[0017]

[0018] Where ψ is the time constant for discounting future costs, t≤m<∞, and r(t) is the instantaneous cost function expressed as,

[0019] r(t)=Z(t) T R1Z(t)+U(t) T R2U(t)(4)

[0020] Among them, R1 and R2 are semi-positive definite matrices that are not zero at the same time, Z(t) is the PID generalized error signal, and U(t) is the control signal.

[0021] A further improvement of the technical solution of the present invention is to define J=W c2 φ(W c1 Z)+ε c It is used to approximate the evaluation neural network and the output of the evaluation neural network is designed to be

[0022]

[0023] in is the estimated weight between the input layer and the hidden layer, is the estimated weight between the hidden layer and the output layer, φ(x) is the hyperbolic tangent activation function;

[0024] By derivation of formula (3), we can get

[0025]

[0026] If the current estimate is perfect, satisfying Equation (6); if this condition is not met, the prediction is adjusted to reduce the inconsistency, then the following function is defined and made close to 0;

[0027]

[0028] In order to make δ(t) approach zero, the objective function is defined as:

[0029] E c =δ T δ / 2(8)

[0030] Using the gradient descent method and a deep neural network, the control rate of the evaluation neural network is designed to be,

[0031]

[0032] Among them, σ c >0 is the learning rate of the evaluation neural network;

[0033] right The derivative is,

[0034]

[0035] Substituting equations (5) and (10) into equation (7), we obtain:

[0036]

[0037] Further seek,

[0038]

[0039] Substituting formula (12) into formula (9), we can obtain

[0040]

[0041] in,

[0042] In order to enhance the robustness of the closed-loop system, a correction term is added to the control rate (14) of the evaluation neural network and rewritten as,

[0043]

[0044] Among them, η c Is a positive number.

[0045] The further improvement of the technical solution of the present invention is that the control signal U is as shown in formula (2), where k D (·) Automatically adjust the adaptive rate of neural network updates using execution,

[0046]

[0047] The input is x is the displacement of the hydraulic cylinder, x d is the hydraulic cylinder command displacement, e is the hydraulic cylinder displacement tracking error, Then define the estimation error

[0048] The neural network ensemble error is defined as where k a >0 is the control gain, and the minimum action network error function is defined as,

[0049]

[0050] Based on the gradient descent method, the update law of the execution network can be obtained by using a deep neural network:

[0051]

[0052] Since it is not possible to obtain a The actual value of And further add correction items where η a > 0, and finally modify the update law to,

[0053]

[0054] Due to the adoption of the above technical solution, the technical advancements achieved by the present invention are:

[0055] The present invention uses an evaluation neural network-execution neural network architecture to automatically adjust the PID time-varying gain in real time without manual intervention, significantly simplifying the parameter adjustment process of a multi-degree-of-freedom heavy-duty hydraulic manipulator and reducing dependence on the operator's professional experience.

[0056] The present invention utilizes the online learning capability of reinforcement learning and provides an optimization direction for executing the neural network by evaluating the real-time approximation error accumulation of the neural network, so that the PID gain can dynamically track system changes, effectively compensate for nonlinear characteristics and unknown disturbances, and improve system adaptability.

[0057] The evaluation neural network of this invention accurately assesses tracking performance by constructing a long-term cost function that comprehensively considers the generalized displacement error and the cumulative amount of the control signal. The execution neural network then adjusts the gain in real time based on the evaluation results and system status to ensure optimal control signals. In practical applications, the hydraulic cylinder displacement tracking error can be controlled to within 0.5mm, significantly outperforming traditional PID control. This makes it particularly suitable for heavy-duty applications such as the nuclear industry and ocean exploration, where motion precision is extremely critical.

[0058] The present invention introduces robustness correction terms in the update laws of both the evaluation neural network and the execution neural network, effectively suppressing the impact of parameter fluctuations and external noise on the system, avoiding oscillation or instability in the closed-loop system, and ensuring long-term stable operation of the hydraulic manipulator under complex working conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 This is a schematic diagram of the hydraulic system principle of the heavy-duty robotic arm system;

[0060] Figure 2 This is a schematic diagram of the hydraulic manipulator PID control principle of a hydraulic manipulator PID control method based on reinforcement learning in the present invention. DETAILED DESCRIPTION

[0061] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings. In the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessary confusion of the concept of the present invention.

[0062] like Figure 1As shown, the controller inputs electrical signals to control each component, the electric motor drives the variable pump to supply energy to the hydraulic system, and the overflow valve acts as a safety valve to provide safety protection for the entire hydraulic system; on the one hand, the unloading valve returns the minimum flow of the variable pump in the standby state of the heavy-loaded robotic arm to the oil tank, and on the other hand, it acts as a speed regulating element to perform bypass throttling speed regulation; the boom, dipper arm, and bucket hydraulic cylinders are respectively controlled by three-position four-way servo valves, and at the same time, the hydraulic cylinder inlet and outlet are equipped with oil replenishment overflow valves, which play a safety protection role when the hydraulic cylinder moves and also have oil return regeneration function; pressure sensors and displacement sensors are installed in the circuit to detect system pressure and hydraulic cylinder displacement, and the output electrical signals are fed back to the controller and stored.

[0063] This embodiment provides a PID control method for a hydraulic manipulator based on reinforcement learning. Figure 2 As shown, a PID controller is designed based on the acquired hydraulic cylinder displacement signal and hydraulic cylinder displacement command signal, and the optimal control signal is generated by the evaluation neural network-execution neural network, which is then input into the servo valve to control the movement of the hydraulic cylinder.

[0064] The specific steps include:

[0065] S1. Use the displacement sensor to obtain the displacement of the hydraulic cylinder of the hydraulic manipulator during the experiment, and feed it back to the controller in real time and store it.

[0066] S2. Design a PID controller based on the acquired hydraulic cylinder displacement signal and hydraulic cylinder command displacement signal.

[0067]

[0068] Among them, k p 、k i and k d is a positive constant, k P (·), k I (·) and k D (·) is the time-varying gain. To reduce the complexity of the gain adjustment parameters, a constant coefficient ω>0 is used to connect the proportional, integral and differential gain parameters and rewrite them in the form of differential gain as k d =k p / 2ω=k i / ω 2 , similarly k D (·)=k P (·) / 2ω=k I (·) / ω 2 .

[0069]

[0070] Therefore, the complex task of determining the PID gains is reduced to choosing two constants k dand ω, the time-varying gain parameter k D (·) Automatically adjust the neural network by implementing its update law.

[0071] S3. Automatically adjust the time-varying gain k using the evaluation neural network and the execution neural network D (·) to modify the PID control gain parameters in real time.

[0072] The evaluation neural network is used to approximate the cumulative amount of the generalized error of the hydraulic cylinder displacement and the control signal to evaluate the performance of the hydraulic cylinder displacement tracking, provide guidance for the execution of neural network update iteration, and improve the tracking performance of the manipulator hydraulic cylinder displacement. The long-term cost function is defined as:

[0073]

[0074] Where ψ is the time constant for discounting future costs, t≤m<∞, and r(t) is the instantaneous cost function expressed as:

[0075] r(t)=Z(t) T R1Z(t)+U(t) T R2U(t) (23)

[0076] Where R1 and R2 are semi-positive matrices that are not zero at the same time. Z(t) is a PID generalized error signal and is defined as Where ω is a constant and e(t) is the tracking error between the actual displacement of the hydraulic cylinder and the command displacement.

[0077] Define J = W c2 φ(W c1 Z)+ε c It is used to approximate the evaluation neural network and the output of the evaluation neural network is designed to be

[0078]

[0079] in, is the estimated weight between the input layer and the hidden layer, is the estimated weight between the hidden layer and the output layer, and φ(x) is the hyperbolic tangent activation function.

[0080] From formula (22), we can get

[0081]

[0082] If the current estimate is perfect, it should satisfy Equation (25). If this condition is not met, the prediction should be adjusted to reduce the inconsistency, and the following function is defined and made close to 0:

[0083]

[0084] In order to make δ(t) approach zero, the objective function is defined as:

[0085]

[0086] Using the gradient descent method and a deep neural network, the control rate of the evaluation neural network is designed as:

[0087]

[0088] Among them, σ c >0 is the learning rate for evaluating the neural network.

[0089] right Taking the derivative we get:

[0090]

[0091] Substituting equations (24) and (29) into equation (26), we can obtain

[0092]

[0093] Further obtain:

[0094]

[0095] Substituting formula (31) into formula (28) yields:

[0096]

[0097] in,

[0098] In order to enhance the robustness of the closed-loop system, a correction term is added to the control rate (33) of the evaluation neural network and rewritten as:

[0099]

[0100] Among them, η c Is a positive number.

[0101] The execution neural network adjusts the PID gain to generate the control signal and realizes good motion control by correcting the PID parameters in real time.

[0102] The control signal U is shown in formula (21), where k D (·) Automatically adjust the adaptive rate of neural network updates using:

[0103]

[0104] The input is x is the displacement of the hydraulic cylinder, xd is the hydraulic cylinder command displacement, e is the hydraulic cylinder displacement tracking error, Then define the estimation error The goal of executing the neural network update law is to make the estimated error ξ a and estimated cost function Therefore, by defining the neural network integration error as where k a >0 is the control gain, and the minimization action network error function is defined as:

[0105]

[0106] Based on the gradient descent method, the update law of the execution network can be obtained by using a deep neural network:

[0107]

[0108] In addition, due to the inability to obtain a The actual value of , we use its estimated value And further add correction items where η a > 0, and finally the update law is modified to:

[0109]

[0110] Using the execution neural network to automatically adjust the time-varying gain k in real time D (·), and then the optimal control parameters of the PID controller are obtained.

[0111] S4. Based on the optimal control parameters of the PID controller obtained by executing the network update law, high-precision tracking control of the hydraulic cylinder is achieved.

[0112] In the above embodiment, the present invention provides a hydraulic robotic arm PID control method based on reinforcement learning. The present invention uses an evaluation neural network-execution neural network architecture to automatically adjust the PID time-varying gain in real time without manual intervention, significantly simplifying the parameter adjustment process of the multi-degree-of-freedom heavy-duty hydraulic robotic arm and reducing the dependence on the operator's professional experience; the present invention uses the online learning ability of reinforcement learning to approximate the error accumulation in real time through the evaluation neural network, and provides an optimization direction for the execution neural network, so that the PID gain can dynamically track system changes, effectively compensate for nonlinear characteristics and unknown disturbances, and improve system adaptability; the evaluation neural network of the present invention constructs a long-term cost function, comprehensively considers the displacement generalized error and the control signal accumulation, and accurately evaluates the tracking performance; the execution neural network corrects the gain in real time based on the evaluation results and system status to ensure the optimal control signal. In practical applications, the displacement tracking error of the hydraulic cylinder can be controlled within 0.5mm, which is significantly better than traditional PID control. It is especially suitable for heavy-load operation scenarios such as the nuclear industry and ocean exploration that require extremely high motion accuracy. At the same time, the present invention introduces robustness correction terms in the update laws of both the evaluation neural network and the execution neural network to effectively suppress the influence of parameter fluctuations and external noise on the system, avoid oscillation or instability in the closed-loop system, and ensure long-term stable operation of the hydraulic robotic arm under complex working conditions.

[0113] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the concept and scope of the present invention. Any modifications and improvements made to the technical solution of the present invention by a person of ordinary skill in the art without departing from the design concept of the present invention shall fall within the scope of protection of the present invention. The technical content for which protection is sought in the present invention is fully set forth in the claims.

Claims

1. A PID control method for a hydraulic manipulator based on reinforcement learning, characterized by: The following steps are included: S1. Use the displacement sensor to obtain the displacement data of the hydraulic cylinder of the hydraulic manipulator during the experiment; S2. designing a PID-like controller based on the acquired hydraulic cylinder displacement signal and the hydraulic cylinder command displacement signal; S3, using the evaluation neural network-executor neural network architecture to automatically adjust the time-varying gain, correct the gain parameters of the PID controller in real time, and obtain the optimal gain parameters of the PID controller; S4. Based on the optimal control parameters of the PID controller, high-precision tracking control of the hydraulic cylinder is achieved.

2. The PID control method for a hydraulic manipulator based on reinforcement learning according to claim 1, characterized in that: The step S1 specifically includes using a displacement sensor to collect displacement data of each hydraulic cylinder of the hydraulic mechanical arm in real time, feeding the collected displacement data back to the controller in real time, and storing the data through a storage module of the controller. The displacement sensor is a high-precision wire-type displacement sensor.

3. The PID control method for a hydraulic manipulator based on reinforcement learning according to claim 1, characterized in that: The expression of the PID controller in step S2 is: Among them, k p , k i and k d is a positive constant, k P (·), k I (·) and k D (·) is the time-varying gain; the proportional, integral, and differential gain parameters are connected using a constant coefficient ω>0 and rewritten in terms of differential gain as k d =k p / 2ω=k i / ω 2 , similarly k D (·)=k P (·) / 2ω=k I (·) / ω 2 ; 4. The PID control method for a hydraulic manipulator based on reinforcement learning according to claim 1, characterized in that: In step S3, the working process of the evaluation neural network includes constructing a long-term cost function to approximate the cumulative amount of the generalized error of the hydraulic cylinder displacement and the control signal. The long-term cost function is defined as: Where ψ is the time constant for discounting future costs, t≤m<∞, and r(t) is the instantaneous cost function expressed as, r(t)=Z(t) T R1Z(t)+U(t) T R2U(t)(4) Among them, R1 and R2 are semi-positive definite matrices that are not zero at the same time, Z(t) is the PID generalized error signal, and U(t) is the control signal.

5. The PID control method for a hydraulic manipulator arm based on reinforcement learning according to claim 4, characterized in that: Define J = W c2 φ(W c1 Z)+ε c It is used to approximate the evaluation neural network and the output of the evaluation neural network is designed to be in is the estimated weight between the input layer and the hidden layer, is the estimated weight between the hidden layer and the output layer, φ(x) is the hyperbolic tangent activation function; By derivation of formula (3), we can get If the current estimate is perfect, satisfying Equation (6); if this condition is not met, the prediction is adjusted to reduce the inconsistency, then the following function is defined and made close to 0; In order to make δ(t) approach zero, the objective function is defined as: E c =d T d / 2(8) Using the gradient descent method and a deep neural network, the control rate of the evaluation neural network is designed to be, Among them, σ c >0 is the learning rate of the evaluation neural network; right The derivative is, Substituting equations (5) and (10) into equation (7), we obtain: Further seek, Substituting formula (12) into formula (9), we can obtain in, In order to enhance the robustness of the closed-loop system, a correction term is added to the control rate (14) of the evaluation neural network and rewritten as, Among them, η c Is a positive number.

6. The PID control method for a hydraulic manipulator arm based on reinforcement learning according to claim 5, characterized in that: The control signal U is shown in formula (2), where k D (·) Automatically adjust the adaptive rate of neural network updates using execution, The input is x is the displacement of the hydraulic cylinder, x d is the hydraulic cylinder command displacement, e is the hydraulic cylinder displacement tracking error, Then define the estimation error The neural network ensemble error is defined as where k a >0 is the control gain, and the minimization action network error function is defined as, Based on the gradient descent method, the update law of the execution network can be obtained by using a deep neural network: Since it is not possible to obtain a The actual value of And further add correction items where η a > 0, and finally modify the update law to,

Citation Information

Patent Citations

  • Wavelet neural network PID online control method and system of hydraulic actuator

    CN111796511A

  • Multi-degree-of-freedom hydraulic mechanical arm real-time control system and method

    CN114800517A

  • Hydraulic actuator driving force control method based on PID (Proportion Integration Differentiation)

    CN117231599A

  • Neural network adaptive tracking control method for joint robots

    US20220152817A1

Cited By

  • A mechanical arm trajectory tracking control method and system based on a double-layer controller

    CN122683733A