Hydraulic multi-way valve intelligent shaft control method based on reinforcement learning

By adopting an intelligent axis control method based on reinforcement learning in the hydraulic multi-channel valve shaft control system, the control accuracy and stability problems caused by the system's nonlinear characteristics and modeling uncertainty are solved, and the high-precision tracking performance and strong anti-interference ability are improved.

CN120029055AActive Publication Date: 2025-05-23NANJING UNIV OF SCI & TECH +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510094461.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-23
Estimated Expiration
2045-01-21

AI Technical Summary

Technical Problem

Due to its nonlinear characteristics and modeling uncertainty, traditional control methods are difficult to achieve high-precision and high-frequency tracking performance, and are prone to system instability and differential explosion problems.

Method used

Using the intelligent shaft control method of hydraulic multi-way valve based on reinforcement learning, a position shaft control controller based on reinforcement learning is designed by establishing a mathematical model of the hydraulic multi-way valve shaft control system, and using the Liyapunov stability theory to prove stability, the active compensation of unknown disturbances of the system and the improvement of anti-interference ability.

Benefits of technology

The high-precision tracking performance of the hydraulic multi-channel valve shaft control system is realized, which avoids the differential explosion problem in traditional reverse step control, reduces the impact of measurement noise on control accuracy, and improves the anti-interference ability and intelligence level of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029055A_ABST
    Figure CN120029055A_ABST
Patent Text Reader

Abstract

The invention discloses a hydraulic multi-way valve intelligent shaft control method based on reinforcement learning, and the shaft control method is based on an execution-evaluation neural network reinforcement learning framework, combines a dynamic surface control thought, and designs a position shaft control intelligent controller considering unknown friction dynamic active compensation. Aiming at the problem of hydraulic multi-way valve position shaft control, the method not only can ensure the unknown friction dynamic active learning compensation of the system and improve the anti-interference capability and the intelligent level of the system, but also can avoid the differential explosion problem in the traditional backstepping control of the electro-hydraulic system, reduce the influence of measurement noise on the control precision, and realize the asymptotic tracking performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electromechanical servo control, and in particular to a hydraulic multi-way valve intelligent axis control method (RLAC) based on reinforcement learning. Background Art

[0002] The hydraulic multi-way valve axis control system has a pivotal position in the fields of robots, heavy machinery, high-performance loading and testing equipment, etc., due to its high power density, large force / torque output, and fast dynamic response. The hydraulic multi-way valve axis control system is a typical nonlinear system, which contains many nonlinear characteristics and modeling uncertainties. The nonlinear characteristics include input nonlinearities such as hysteresis and saturation, multi-way valve flow pressure nonlinearity, friction nonlinearity, etc. The modeling uncertainty includes parameter uncertainty and uncertainty nonlinearity, among which the parameter uncertainty mainly includes load mass, actuator viscous friction coefficient, leakage coefficient, multi-way valve flow gain, hydraulic oil elastic modulus, etc. The uncertainty nonlinearity mainly includes unmodeled friction dynamics, system high-order dynamics, external interference and unmodeled leakage. When the hydraulic multi-way valve axis control system develops towards high precision and high frequency response, the nonlinear characteristics of the system will have a more significant impact on the system performance, and the existence of modeling uncertainty will make the controller designed with the nominal model of the system unstable or downgraded. Therefore, the nonlinear characteristics and modeling uncertainty of the hydraulic multi-way valve axis control system are important factors that limit the improvement of system performance. With the continuous advancement of technology in the industrial and defense fields, the controllers previously designed based on traditional linear theory have gradually failed to meet the high performance requirements of the system. Therefore, it is necessary to study more advanced nonlinear control strategies based on the nonlinear characteristics of the hydraulic multi-way valve axis control system.

[0003] Many methods have been proposed to solve the nonlinear control problem of hydraulic multi-way valve axis control system. Among them, the adaptive control method is a very effective method for dealing with parameter uncertainty problems, and can obtain steady-state performance of asymptotic tracking. However, it is powerless for uncertainty nonlinearity such as external load interference. When the uncertainty nonlinearity is too large, the system may become unstable. The actual hydraulic multi-way valve axis control system has uncertainty nonlinearity, so the adaptive control method cannot obtain high-precision control performance in practical applications; as a robust control method, classical sliding mode control can effectively deal with any bounded modeling uncertainty and obtain steady-state performance of asymptotic tracking, but the discontinuous controller designed by classical sliding mode control is prone to cause the chatter problem of the sliding surface, thereby deteriorating the tracking performance of the system; in order to solve the problems of parameter uncertainty and uncertainty nonlinearity at the same time, the adaptive robust control method is proposed. This control method can enable the system to obtain certain transient and steady-state performance when the two modeling uncertainties exist at the same time. If high-precision tracking performance is to be obtained, the feedback gain must be increased to reduce the tracking error. Due to the existence of measurement noise, the gain is too large, which often leads to high-gain feedback, thereby causing chattering of the control input, thereby deteriorating the control performance and even causing system instability. Summary of the invention

[0004] The purpose of the present invention is to provide a hydraulic multi-way valve intelligent axis control method with intelligent learning ability, strong anti-interference ability and high tracking performance, which can not only ensure active learning compensation for unknown friction dynamics of the system, improve the anti-interference ability and intelligence level of the system, but also avoid the differential explosion problem in traditional backstepping control of electro-hydraulic systems, reduce the influence of measurement noise on control accuracy, and achieve asymptotic tracking performance.

[0005] The technical solution to achieve the purpose of the present invention is: a hydraulic multi-way valve intelligent axis control method based on reinforcement learning, comprising the following steps:

[0006] Step 1: Establish a mathematical model of the hydraulic multi-way valve axis control system and proceed to step 2.

[0007] Step 2: Based on the mathematical model of the hydraulic multi-way valve axis control system, design a position axis control controller based on reinforcement learning, and then proceed to step 3.

[0008] Step 3: Use Lyapunov stability theory to prove the stability of the position axis control controller and obtain the result that the system tracking error is asymptotically stable.

[0009] Compared with the prior art, the present invention has the following significant advantages: (1) it has intelligent learning capability and a high level of intelligence; (2) it can realize active compensation of unknown disturbances in the system and has strong anti-interference capability; (3) it can avoid the differential explosion problem in the traditional backstepping control of the electro-hydraulic system, reduce the influence of measurement noise on control accuracy, and achieve high-precision tracking performance. The simulation results verify its effectiveness. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 It is a schematic diagram of the principle of the hydraulic multi-way valve intelligent axis control method based on reinforcement learning in the present invention.

[0011] Figure 2 It is a schematic diagram of the principle of the hydraulic multi-way valve axis control system of the present invention.

[0012] Figure 3 It is a curve diagram of the tracking process of the system output to the expected instruction under the action of the RLAC controller designed by the present invention.

[0013] Figure 4 It is a curve diagram showing the tracking error of the system changing with time under the action of the RLAC controller designed by the present invention.

[0014] Figure 5 It is a tracking error comparison curve diagram of the system under the action of the RLAC controller designed by the present invention and the traditional PID controller.

[0015] Figure 6 It is a control input curve diagram of the system under the action of the RLAC controller designed by the present invention. DETAILED DESCRIPTION

[0016] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0017] Combination Figure 1 and Figure 2 The present invention provides a hydraulic multi-way valve intelligent axis control method based on reinforcement learning, comprising the following steps:

[0018] Step 1: Establish a mathematical model of the hydraulic multi-way valve axis control system, as follows:

[0019] Step 1-1, the hydraulic multi-way valve axis control system is applied to the linear motion of large industrial heavy-load mechanical equipment, wherein the load is fixedly connected to the piston rod on the hydraulic cylinder, and the hydraulic multi-way valve controls the movement of the piston rod on the hydraulic cylinder, thereby driving the load to move.

[0020] According to Newton's second law, the force balance equation of the hydraulic multi-way valve axis control system is:

[0021]

[0022] In formula (1), m represents the mass of the load, y represents the displacement of the hydraulic cylinder piston rod, Indicates the speed of the hydraulic cylinder piston rod. represents the acceleration of the hydraulic cylinder piston rod, A represents the effective area of ​​the hydraulic cylinder piston, P 1 Indicates the oil pressure in the hydraulic cylinder oil inlet chamber, P 2 Indicates the oil pressure of the hydraulic cylinder outlet chamber. Indicates the friction force on the load, d 1 (t) represents the unmodeled mechanical disturbance of the system, and t represents time.

[0023] Then formula (1) can be rewritten as:

[0024]

[0025] In the hydraulic multi-way valve axis control system, ignoring the leakage of the cylinder oil, the pressure dynamic equation is:

[0026]

[0027] Formula (3), β e Indicates the effective elastic modulus of the oil, C t Indicates the leakage coefficient of the hydraulic cylinder, the oil pressure difference P in the oil chamber and out of the oil chamber L =P 1 -P 2 , oil inlet chamber control volume V 1 =V 01 +Ay, oil outlet chamber control volume V 2 =V 02 -Ay, V 01 Represents the initial volume of the oil inlet chamber, V 02 Indicates the initial volume of the oil chamber, Q 1 Indicates the flow rate of the oil inlet cavity, Q 2 Indicates the flow rate of the oil chamber, q 1 Indicates P 1 The unmodeled disturbance, q 2 Indicates P 2 The unmodeled disturbance Indicates P 1 The first derivative of Indicates P 2 The first derivative of .

[0028] Q 1 , Q 2 Respectively with the displacement x of the hydraulic multi-way valve core v There are the following relationships:

[0029]

[0030] Among them, the hydraulic multi-way valve coefficient C d Indicates the flow coefficient of the hydraulic multi-way valve, w 0 It represents the valve core area gradient of the hydraulic multi-way valve, ρ represents the oil density, P s Indicates the oil supply pressure, P r represents the return oil pressure, sg(·) represents the function of the intermediate variable ·, and is defined as:

[0031]

[0032] Ignoring the dynamics of the hydraulic multi-way valve core, assume that the control input u acting on the valve core and the valve core displacement x v Proportional relationship, that is, satisfying x v =k i u, where k i represents the voltage-valve core displacement gain coefficient, so equation (4) is rewritten as:

[0033]

[0034] Formula (6), intermediate variable k u =k q k i , intermediate variable Intermediate variables

[0035] Step 1-2, define state variables: Among them, the intermediate variable x 1 =y, intermediate variable Intermediate variable x 3 =(AP 1 -AP 2 ) / m, then transform equation (2) into the state equation:

[0036]

[0037] Formula (7), Represents x 1 The first derivative of Represents x 2 The first derivative of Represents x 3 The first-order derivative of the unknown dynamics of the system D 1 =d 1 (t) / m, intermediate variable F(x 2 )=F f (x 2 ) / m, intermediate variable Intermediate variables Intermediate variables System unknown dynamics

[0038] To facilitate the design of the controller, the following assumptions are made:

[0039] Assumption 1: The system is expected to track the position command x d It is second-order continuous, and the system expects that the position command, velocity command and acceleration command are all bounded;

[0040] Assumption 2: The system has unknown dynamics D 1 With D 2 satisfy:

[0041] |D 1 |≤δ 1 ,|D 2 |≤δ 2 (8)

[0042] Formula (8), δ 1 and δ 2 are all unknown positive constants.

[0043] Go to step 2.

[0044] Step 2: Based on the mathematical model of the hydraulic multi-way valve axis control system, a position axis control controller based on reinforcement learning is designed, as follows:

[0045] Step 2-1: To facilitate controller design, define the tracking error z of the system 1 =x 1 -x d , x d The system expects to track the position command and designs the following nonlinear filter:

[0046]

[0047] Formula (9), filter gain τ 1 >0,s 1 Represents x 2 Virtual control, s 1f Indicates 1 The filtered signal, s 1f With x 2 The error z 2 =x 2 -s 1f ,s 1 Filter error ε 1 =s 1f -s 1 , gain l 1 >0 means The upper bound of , σ(t) represents a function that is always positive and satisfies Where ν represents the integral variable, represents a positive constant, Indicates1 The first derivative of Indicates 1f The first derivative of .

[0048] Right 1 Taking the derivative, we get:

[0049]

[0050] Design virtual controllers 1 for:

[0051]

[0052] Formula (11), gain k 1 >0, then

[0053]

[0054] Step 2-2, z 2 The derivative is:

[0055]

[0056] Design the following nonlinear filter:

[0057]

[0058] Formula (14), filter gain τ 2 >0,s 2 Represents x 3 Virtual control, s 2f Indicates 2 The filtered signal, s 2f With x 3 The error z 3 =x 3 -s 2f ,s 2 Filter error ε 2 =s 2f -s 2 , gain l 2 >0 means The upper bound of Indicates 2 The first derivative of Indicates 2f The first derivative of .

[0059] Design virtual controllers 2 for:

[0060]

[0061] Formula (15), gain k 2 >0,s2s represents the intermediate variable, δ 2 represents a positive constant, It means F(x 2 ), The specific form is:

[0062]

[0063] Formula (16), W a The estimated value of W a represents the weights of executing the neural network, represents the activation function of executing the neural network, X a Represents the input to execute the neural network.

[0064] Accordingly, the reinforcement learning signal R(t) can be designed as

[0065]

[0066] Formula (17), Represents the evaluation neural network weight W c The estimated value of represents the activation function of the evaluation neural network, χ represents the intermediate variable, represents the first-order derivative of χ, B s represents the intermediate variable, B s1 represents the intermediate variable, B s2 represents the intermediate variable, express The first derivative of 0 Represents a positive constant.

[0067] Then the weight update law of the execution-evaluation neural network can be designed as:

[0068]

[0069] Formula (18), W a The first derivative of , Γ a represents the weight gain matrix of the executed neural network, W c The first derivative of , Γ c represents the weight gain matrix of the evaluation neural network, and Proj(·) represents the discontinuous mapping function.

[0070] Substituting equations (15) to (18) into equation (13), we obtain:

[0071]

[0072] Formula (19), Represents the execution of the neural network weight W a The estimation error, ε a Represents the approximation error of executing the neural network.

[0073] Step 2-3, z 3 The derivative is:

[0074]

[0075] According to formula (20), the control input of the valve core, that is, the position axis control controller u based on reinforcement learning is:

[0076]

[0077] Formula (21), gain k 3 >0,u s represents the intermediate variable, δ 3 Represents a positive constant.

[0078] Substituting formula (21) into formula (20), we get:

[0079]

[0080] Go to step 3.

[0081] Step 3, using Lyapunov stability theory to prove the stability of the position axis control controller, the result of asymptotic stability of the system tracking error is obtained, as follows:

[0082] The Lyapunov function V is defined as follows:

[0083]

[0084] in, Represents the evaluation neural network weight W c The estimation error.

[0085] Deriving equation (23) and substituting equations (9), (12), (14), (18), (19) and (22) into equation (23) yields:

[0086]

[0087] Considering and The expression can be obtained:

[0088]

[0089] Notice

[0090]

[0091] Substituting equation (26) into equation (25), we can obtain

[0092]

[0093] Define the intermediate variables z and Λ as:

[0094] z=[z 1 ; z 2 ; z 3 ; ε 1 ; ε 2 ](28)

[0095]

[0096] Formula (29), the intermediate variable Λ 1 and Λ 2 They are

[0097]

[0098] By adjusting the gain k 1 , k 2 , k 3 and filter gain τ 1 , τ 2 , the symmetric matrix Λ can be made a positive definite matrix, then we can get:

[0099]

[0100] Formula (31), intermediate variable Φ = z T Λz,T represents transpose.

[0101] Integrating both sides of equation (31) we can obtain:

[0102]

[0103] From equation (32), we can see that V is bounded and the integral of Φ is bounded. It can be concluded that all signals in the system are bounded. Therefore, Φ is uniformly continuous. According to Barbalat's lemma, when time tends to positive infinity, the tracking error z 1 Tends toward 0.

[0104] Therefore, it is concluded that by adjusting the gain k 1 , k 2 , k 3 and filter gain τ 1 , τ 2 The position axis control controller based on reinforcement learning designed for the hydraulic multi-way valve axis control system can make the system obtain the result that the tracking error converges to 0 asymptotically. The principle diagram of the position axis control controller based on reinforcement learning for the hydraulic multi-way valve axis control system is shown in Figure 1 shown.

[0105] Example

[0106] In order to evaluate the performance of the designed controller, the physical parameters of the hydraulic multi-way valve axis control system in the simulation are shown in Table 1:

[0107] Table 1 System physical parameters

[0108] Physical parameters Numeric Physical parameters Numeric <![CDATA[A(m 2 )]]> <![CDATA[2×10 -4 ]]> <![CDATA[β e (Well)]]> <![CDATA[2×10 8 ]]> m(kg) 40 B(N·s / m) 80 <![CDATA[C t (m 5 / (N·s))]]> <![CDATA[7×10 -12 ]]> <![CDATA[k u (m / V)]]> <![CDATA[4×10 -8 ]]> <![CDATA[V 01 (m 3 )]]> <![CDATA[1×10 -3 ]]> <![CDATA[V 02 (m 3 )]]> <![CDATA[1×10 -3 ]]> <![CDATA[P s (MPa)]]> 7 <![CDATA[P r (MPa)]]> 0

[0109] Given a system with the expected instruction x d =0.1sin(πt)×(1-e -0.1t )m.

[0110] The following controllers are used for comparison in the simulation:

[0111] Hydraulic multi-way valve intelligent axis control method based on reinforcement learning (RLAC): taking gain k 1 =10, k 2 =1, k 3 =1,τ 1 =2000,τ 2 =2000,l 1 = l 2 =1.

[0112] PID controller: The steps for selecting PID controller parameters are: first, ignoring the nonlinear dynamics of the hydraulic multi-way valve axis control system, a set of controller parameters are obtained through the PID parameter self-tuning function in Matlab, and then the self-tuning parameters are fine-tuned after adding the nonlinear dynamics of the system to obtain the best tracking performance. The selected controller parameters are k P =10, k I =1, k D =1.

[0113] The expected command of the system, the tracking error of the RLAC controller, and the tracking error comparison between the RLAC controller and the PID controller are shown as follows: Figure 3 , Figure 4 and Figure 5 As shown. Figure 4 It can be seen that under the action of the RLAC controller, the position output of the hydraulic multi-way valve axis control system has a high tracking accuracy for the command, and the amplitude of the steady-state tracking error is about 6×10 -5 m. From Figure 5 The comparison of the tracking errors of the two controllers shows that the tracking error of the RLAC controller proposed in the present invention is much smaller than that of the PID controller, and the tracking performance is more superior.

[0114] Figure 6It is a curve diagram of the control input of the hydraulic multi-way valve axis control system under the action of the RLAC controller changing with time. It can be seen from the figure that the obtained control input is a low-frequency continuous signal, which is more conducive to execution in practical applications.

Claims

1. A hydraulic multi-way valve intelligent axis control method based on reinforcement learning, characterized in that: The following steps are involved: Step 1: Establish a mathematical model of the hydraulic multi-way valve axis control system, and then proceed to step 2; Step 2: Based on the mathematical model of the hydraulic multi-way valve axis control system, design a position axis control controller based on reinforcement learning, and then proceed to step 3; Step 3: Use Lyapunov stability theory to prove the stability of the position axis control controller and obtain the result that the system tracking error is asymptotically stable.

2. The intelligent axis control method of a hydraulic multi-way valve based on reinforcement learning according to claim 1 is characterized in that: In step 1, a mathematical model of the hydraulic multi-way valve axis control system is established, as follows: Step 1-1, the hydraulic multi-way valve axis control system is applied to the linear motion of large industrial heavy-load mechanical equipment, wherein the load is fixedly connected to the piston rod on the hydraulic cylinder, and the hydraulic multi-way valve controls the movement of the piston rod on the hydraulic cylinder, thereby driving the load to move; Step 1-2: Define state variables and obtain the state equation of the hydraulic multi-way valve axis control system.

3. The intelligent axis control method of a hydraulic multi-way valve based on reinforcement learning according to claim 2 is characterized in that: In step 1-1, the hydraulic multi-way valve axis control system is applied to the linear motion of large industrial heavy-duty mechanical equipment, in which the load is fixedly connected to the piston rod on the hydraulic cylinder, and the hydraulic multi-way valve controls the movement of the piston rod on the hydraulic cylinder, thereby driving the load to move, as follows: According to Newton's second law, the force balance equation of the hydraulic multi-way valve axis control system is: In formula (1), m represents the mass of the load, y represents the displacement of the hydraulic cylinder piston rod, Indicates the speed of the hydraulic cylinder piston rod. represents the acceleration of the hydraulic cylinder piston rod, A represents the effective action area of ​​the hydraulic cylinder piston, P1 represents the oil pressure in the hydraulic cylinder oil inlet chamber, P2 represents the oil pressure in the hydraulic cylinder oil outlet chamber, represents the friction force on the load, d1(t) represents the unmodeled mechanical disturbance of the system, and t represents time; Then formula (1) can be rewritten as: In the hydraulic multi-way valve axis control system, ignoring the leakage of the cylinder oil, the pressure dynamic equation is: Formula (3), β e Indicates the effective elastic modulus of the oil, C t Indicates the leakage coefficient of the hydraulic cylinder, the oil pressure difference P between the oil inlet chamber and the oil outlet chamber of the cylinder L =P1-P2, oil inlet chamber control volume V1=V 01 +Ay, the oil outlet chamber control volume V2 = V 02 -Ay, V 01 Represents the initial volume of the oil inlet chamber, V 02 represents the initial volume of the oil outlet chamber, Q1 represents the flow rate of the oil inlet chamber, Q2 represents the flow rate of the oil outlet chamber, q1 represents the unmodeled disturbance of P1, q2 represents the unmodeled disturbance of P2, represents the first-order derivative of P1, represents the first derivative of P2; Q1 and Q2 are respectively related to the displacement x of the hydraulic multi-way valve core v There are the following relationships: Among them, the hydraulic multi-way valve coefficient C d represents the flow coefficient of the hydraulic multi-way valve, w0 represents the valve core area gradient of the hydraulic multi-way valve, ρ represents the oil density, P s Indicates the oil supply pressure, P r represents the return oil pressure, sg(·) represents the function of the intermediate variable ·, and is defined as: Ignoring the dynamics of the hydraulic multi-way valve core, assume that the control input u acting on the valve core and the valve core displacement x v Proportional relationship, that is, satisfying x v =k i u, where k i represents the voltage-valve core displacement gain coefficient, so equation (4) is rewritten as: Formula (6), intermediate variable k u =k q k i , intermediate variable Intermediate variables 4. The method for intelligent axis control of a hydraulic multi-way valve based on reinforcement learning according to claim 3 is characterized in that: In step 1-2, define the state variables: Among them, the intermediate variable x1 = y, the intermediate variable The intermediate variable x3 = (AP1-AP2) / m, then transform equation (2) into the state equation: Formula (7), represents the first-order derivative of x1, represents the first-order derivative of x2, represents the first-order derivative of x3; The unknown dynamics of the system D1 = d1(t) / m, the intermediate variable F(x2) = F f (x2) / m, intermediate variable Intermediate variables Intermediate variables System unknown dynamics Go to step 2.

5. The method for intelligent axis control of a hydraulic multi-way valve based on reinforcement learning according to claim 4 is characterized in that: To facilitate the design of the controller, the following assumptions are made: Assumption 1: The system is expected to track the position command x d It is second-order continuous, and the system expects that the position command, velocity command and acceleration command are all bounded; Assumption 2: The unknown dynamics D1 and D2 of the system satisfy: |D1|≤δ1,|D2|≤δ2 (8) In formula (8), δ1 and δ2 are both unknown positive constants.

6. The intelligent axis control method of a hydraulic multi-way valve based on reinforcement learning according to claim 5 is characterized in that: In step 2, based on the mathematical model of the hydraulic multi-way valve axis control system, a position axis control controller based on reinforcement learning is designed as follows: Step 2-1: To facilitate controller design, define the tracking error of the system as z1 = x1-x d , x d The system expects to track the position command and designs the following nonlinear filter: Formula (9), filter gain τ1>0, s1 represents the virtual control of x2, s 1f represents the filtered signal of s1, s 1f The error with x2 is z2 = x2-s 1f , the filtering error of s1 ε1 = s 1f -s1; represents the first-order derivative of s1, Indicates 1f The first-order derivative of; gain l1>0, indicating The upper bound of ; σ(t) represents a function that is always positive and satisfies Where ν represents the integral variable, represents a constant that is always positive; Taking the derivative of z1, we get in, Represents x d First-order derivative; Design virtual control s1 as: Formula (11), gain k1>0, then Step 2-2, take the derivative of z2 and get Design the following nonlinear filter: Formula (14), filter gain τ2>0, s2 represents the virtual control of x3, s 2f represents the filtered signal of s2, s 2f The error with x3 is z3 = x3-s 2f , s2 filtering error ε2 = s 2f -s2, gain l2>0, means The upper bound of represents the first-order derivative of s2, Indicates 2f The first derivative of ; Design virtual control s2 as: Formula (15), gain k2>0, s 2s represents an intermediate variable, δ2 represents a positive constant, represents the estimated value of F(x2), The specific form is: Formula (16), W a The estimated value of W a represents the weights of executing the neural network, represents the activation function of executing the neural network, X a Represents the input to execute the neural network; Accordingly, the reinforcement learning signal R(t) is designed as Formula (17), Represents the evaluation neural network weight W c The estimated value of represents the activation function of the evaluation neural network, χ represents the intermediate variable, represents the first-order derivative of χ, B s represents the intermediate variable, B s1 represents the intermediate variable, B s2 represents the intermediate variable, express The first derivative of , δ0 represents a constant that is always positive; The weight update law of the execution-evaluation neural network is designed as: Formula (18), W a The first derivative of , Γ a represents the weight gain matrix of the executed neural network, W c The first derivative of , Γ c represents the weight gain matrix of the evaluation neural network, Proj(·) represents the discontinuous mapping function; Substituting equations (15) to (18) into equation (13), we get: Formula (19) executes the neural network weight W a The estimated error ε a represents the approximation error of executing the neural network; Step 2-3, take the derivative of z3 and get According to formula (20), the control input of the valve core, that is, the position axis control controller u based on reinforcement learning is: Formula (21), gain k3>0, u s represents an intermediate variable, δ3 represents a positive constant; Substituting formula (21) into formula (20), we get: Go to step 3.

7. The intelligent axis control method of a hydraulic multi-way valve based on reinforcement learning according to claim 6 is characterized in that: In step 3, the stability of the position axis control controller is proved by using Lyapunov stability theory, and the result that the system tracking error is asymptotically stable is obtained, as follows: The Lyapunov function V is defined as follows: Among them, the evaluation neural network weight W c The estimated error The stability is proved by using Lyapunov stability theory, and the result that the system tracking error is asymptotically stable is obtained.

Citation Information

Patent Citations

  • Electro-hydraulic proportional servo valve position axial control method considering input time lag

    CN114943146A

  • Self-adaptive control method for multi-degree-of-freedom valve-controlled hydraulic mechanical arm based on reinforcement learning

    CN118046388A

  • Self-learning gain position axial control method for electro-hydraulic proportional servo valve

    CN118192225A

  • System and method for calibrating a control and regulating device for gas exchange valves of an internal combustion engine

    DE102019126246A1

  • Adaptive control device and adaptive control method, as well as control device and control method for injection molding machine

    WO2013031082A1