Fault-tolerant control method for industrial process min-max optimization based on reinforcement learning

By employing a reinforcement learning-based minimum-maximum optimization method for industrial processes and designing fault-tolerant control strategies using real data, this approach overcomes the control difficulties of traditional methods under external disturbances and actuator failures, achieving a wider range of applicability and better control performance, making it suitable for fault-tolerant control of industrial processes.

CN114706356BActive Publication Date: 2025-12-09HAINAN NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210358730.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-06
Publication Date
2025-12-09
Estimated Expiration
2042-04-06

AI Technical Summary

Technical Problem

Existing industrial process control methods mostly rely on system models, which makes it difficult to effectively control under external disturbances and actuator failures, and thus fails to achieve the desired control effect.

Method used

A fault-tolerant control method based on reinforcement learning for industrial processes, namely mini-maximum optimization, is adopted. By learning from real data in the industrial process, the optimal control strategy and worst external disturbance are designed to achieve fault-tolerant control that does not depend on the system model.

Benefits of technology

In the event of external disturbances and actuator failures, it achieves a wider range of applicability, better tracking performance, and superior control effect. It can maintain good control performance under both normal and fault conditions, replacing traditional model-based fault-tolerant control methods and ensuring the safety and efficiency of industrial production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114706356B_ABST
    Figure CN114706356B_ABST
Patent Text Reader

Abstract

This invention relates to the field of industrial control technology, specifically a fault-tolerant control method for minimum-maximum optimization of industrial processes based on reinforcement learning. It includes: (1) establishing an augmented state-space model containing tracking error and state increments based on the original system state-space model with actuator failures and external disturbances, and proposing a performance index function based on the augmented state-space model; (2) proposing a value function and a Q function based on the performance index function, and constructing corresponding expressions for the optimal control input, worst-case external disturbance, optimal control gain, and worst-case external disturbance gain; (3) collecting data θ based on the initial control gain and external disturbance gain that enable system stability. j (k) and ρ k j θ j (k) and ρ k j It is the data containing system production information generated in the j-th iteration; (4) Update the control gain K1 through reinforcement learning. F External disturbance gain K2 F (5) If the iteration termination condition is met, the iteration ends; otherwise, return to step (4) to continue the iteration. This invention has a wide range of applications, good tracking performance, and good control effect.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of industrial control, and in particular to a fault-tolerant control method for minimum-maximum optimization of an industrial process based on reinforcement learning. BACKGROUND

[0002] Modern industrial processes have undergone many changes with the improvement of scientific and technological level. The production process is more intelligent and efficient, the production scale is increasingly large, and the production equipment is more precise and complex. This means that industrial processes are more susceptible to faults or external disturbances during production, which weakens the control effect of control methods designed for ideal conditions or even completely destroys the control effect. In this context, people no longer limit their vision to studying control methods designed for ideal conditions, which has led to the development of robust control aimed at weakening the negative impact of external disturbances on system performance indicators and fault-tolerant control for actuator faults and other fault conditions. However, reviewing their development achievements reveals that past control methods were mostly model-based control methods that relied heavily on system models, so once they deviated from the model, they would be in trouble and unable to achieve the control goal. In this context, people began to seek new control methods, especially new fault-tolerant control methods for industrial processes with both external disturbances and actuator faults.

[0003] At present, industrial processes can generate a large amount of data reflecting the real dynamics of the system during production. These real data have potential value, and how to fully utilize these data to design corresponding control methods in combination with reinforcement learning has attracted the attention of many scholars. Therefore, there is an urgent need to develop a fault-tolerant control method for minimum-maximum optimization of an industrial process based on reinforcement learning. SUMMARY

[0004] The present application provides a fault-tolerant control method for minimum-maximum optimization of an industrial process based on reinforcement learning, which is based on reinforcement learning method and continuously learns by using data generated by actual production of industrial processes, thereby obtaining optimal control strategy and worst external disturbance, and finally achieving good fault-tolerant control effect and tracking performance.

[0005] To achieve the above purpose, the main technical solutions adopted by the present application include:

[0006] The present application provides a fault-tolerant control method for minimum-maximum optimization of an industrial process based on reinforcement learning, which comprises the following steps:

[0007] The beneficial effects of the present application are: the present application proposes a fault-tolerant control method for industrial process minimum-maximum optimization based on reinforcement learning, which can achieve independence from system model in the case of external disturbance and actuator failure, and only with the large amount of real data generated by the industrial process itself in production to continuously learn and design a control method, and finally achieve the ideal control target, effectively solve the fault-tolerant control problem of industrial process. Not only that, the optimal control input, the worst external disturbance, and the control effect achieved by the optimal control gain and the worst external disturbance gain obtained by relying on real production information for learning are also better than the past model-based fault-tolerant control method. Compared with the model-based fault-tolerant control method for systems with actuator failure and external disturbance, the fault-tolerant control method proposed in the present application has a wider range of application, better tracking performance, and better control effect. Whether the system is in normal or fault condition, the fault-tolerant control method based on reinforcement learning for industrial process minimum-maximum optimization proposed in the present application can achieve good control effect. In the current and future industrial process control problem, the fault-tolerant control method proposed in the present application can well replace the traditional model-based fault-tolerant control method, widen the range of solvable actuator failure, and have more use value. It is an excellent control method to ensure the safe and efficient operation of industrial production process and produce quality and quantity products, which is beneficial to maintain the life and property safety involved in the current industrial process. BRIEF DESCRIPTION OF DRAWINGS

[0008] Figure 1 The figure shows the norm convergence of the difference between the matrix K1 F , K2 F and the optimal K1 F* , K2 F* during the learning process;

[0009] Figure 2 The figure shows the norm convergence of the difference between the matrix H and the optimal H * during the learning process;

[0010] Figure 3 The figure shows the output, tracking error, input and external disturbance comparison of the method of the present application and the traditional model-based fault-tolerant control method under normal condition α=1;

[0011] Figure 4 The figure shows the output, tracking error, input and external disturbance comparison of the method of the present application and the traditional model-based fault-tolerant control method under fault coefficient α=0.6;

[0012] Figure 5 The figure shows the output, tracking error, input and external disturbance comparison of the method of the present application and the traditional model-based fault-tolerant control method under fault coefficient α=1.5. DETAILED DESCRIPTION

[0013] For better explaining the present application, in order to facilitate understanding, the present application is described in detail below through specific embodiments combined with the drawings.

[0014] The present application provides a fault-tolerant control method for minimum-maximum optimization of industrial processes based on reinforcement learning. The method comprises the following steps:

[0015] Step (1), constructing a new system model and proposing a performance index function: based on the original system state space model with actuator faults and external disturbances, an augmented state space model containing tracking errors and state increments is established, and a performance index function is proposed according to the augmented state space model.

[0016] Step (1) specifically comprises the following steps:

[0017] u ik F , i = 1, 2, …, m represents the output signal of the faulty actuator, A fault model can be established: u k F = αu k , α > 0. Wherein, α = diag[ α 1 α 2,…, α m ], α = diag[α1 α2,…,α m ], α i ≤ 1, are known constants, and α i is unknown, but is assumed to vary within a known range. Therefore, there will be u k F = αu k . If corresponds to the actual model u ik F = u ik , which means that the system is in an ideal condition where the actuator does not fail; if α i = 0, which corresponds to the system in a complete failure of the actuator; if α i > 0, which corresponds to the system in a partial failure of the actuator.

[0018] Therefore, the original system model of the industrial process with actuator faults is:

[0019]

[0020] where k represents the running time of the original system, are the state of the original system at k time, the actual input signal of the actuator, the external disturbance and the output, respectively, is the state of the original system at k+1 time, and {A, B, C, D} are system matrices with dimensions matching the state, input and external disturbance.

[0021] Consider designing iterative learning control law u k = u k-1 + u Δk , where u k , u k-1 are the input of the original system at k time and k-1 time, respectively, is the iterative update rate at k time.

[0022] For a given desired output trajectory If the tracking error, the increment of disturbance and the difference equation of the original system are denoted as y Δk = y r - y k , w Δk = w k - w k-1 and x Δk = x k - x k-1 , is the output of the original system at k time, w k and w k-1 are the external disturbance at k time and the external disturbance at k-1 time, respectively, and x k , x k-1 are the state of the original system at k time and k-1 time, respectively. The augmented state space model can be obtained as

[0023]

[0024] where, x Δk+1 is the difference between the state of the original system at k+1 time and k time, y Δk+1 is the tracking error of the original system at k+1 time; x Δk is the difference between the state of the original system at k time and k-1 time, y Δk is the tracking error of the original system at k time; u Δk is the iterative update rate of the original system at k time; w Δk is the difference between the external disturbance of the original system at k time and the external disturbance at k-1 time; is the difference between the state of the original system at k time and the state of the original system at k-1 time, and {Z k , u Δk , w ΔkThe system matrix of the original system, which is matched in dimension, is composed of {A, B, C, D} is the system matrix of the original system, and I is the unit matrix; α is the fault coefficient; Z k is the state of the augmented state space model at time k, u Δk is the input of the augmented state space model at time k, and w Δk is the external disturbance of the augmented state space model at time k.

[0025] A performance index function based on the augmented state space model is proposed: wherein Z i is the state of the augmented state space model at time i, u Δi is the input of the augmented state space model at time i, i = k, k + 1,..., ∞, and w Δi is the external disturbance of the augmented state space model at time i; Q and R are positive definite matrices matched in dimension, respectively, with the state Z i and the input u Δi ; γ ≥ 0, and γ represents the sustained disturbance attenuation level.

[0026] The goal of the study is to find an optimal control input and a worst external disturbance u Δk = K1 F Z k , w Δk = K2 F Z k that make the performance index satisfy the minimum-maximum optimization, or an optimal control gain and a worst external disturbance gain

[0027] Step (2), optimal control input and worst external disturbance design: according to the performance index function, a value function and a Q function are proposed, and expressions of the corresponding optimal control input, worst external disturbance, optimal control gain, and worst external disturbance gain are constructed.

[0028] For , the value function is defined as is a symmetric positive definite matrix, and the value function satisfies the condition

[0029]

[0030] J* is the performance index J to be implemented. The Q function is defined as:

[0031] Q*(Z k ,u Δk ,w Δk ) = Z k T QZ k + u Δk T RuΔk -γ 2 w Δk T w Δk +V*(Z k+1 ,u Δk+1 ,w Δk+1 )

[0032] Under certain conditions, there will be J = V k , i.e. There will also be Where, Further, we can get The optimal control input, the worst external disturbance and the control gain, the external disturbance gain u Δk = (H uu -H uw (H ww ) -1 H wu ) -1 (H uw (H ww ) -1 H wZ -H uZ )Z k , K1 F = (H uu -H uw (H ww ) -1 H wu ) -1 (H uw (H ww ) -1 H wZ -H uZ ), w Δk = (H ww -H wu (H uu ) -1 H uw ) -1 (H wu (H uu ) -1 H uZ -H wZ )Z k , K2 F = (H ww -H wu (H uu ) -1 H uw ) -1 (H wu (H uu ) -1 H uZ -H wZ )uu ,H uw ,H ww ,H wu ,H wZ ,H uZ is a component of matrix H derived from Q function.

[0033] where Z k is the state of the augmented state space model at time k, u Δk is the input of the augmented state space model at time k, w Δk is the external disturbance of the augmented state space at time k, and α is the fault coefficient matrix, Q and R are positive definite matrices with dimensions matching the state Z k and the input u Δk , and γ ≥ 0, where γ represents the level of sustained disturbance attenuation, V*(Z k+1 , u Δk+1 , w Δk+1 ) is the value function at time k+1, is composed of the system matrix {A, B, C, D} of the original system.

[0034] Step (3), initialization and data collection: given the initial control gain (K1 F ) 0 and the external disturbance gain (K2 F ) 0 that can stabilize the system, and collect data θ j (k) and ρ k j .

[0035] Let j = 0, j is the iteration index, given the initial control gain (K1 F ) 0 and the external disturbance gain (K2 F ) 0 that can stabilize the system, (K1 F ) 0 and (K2 F ) 0 are the initial control gain and the external disturbance gain, respectively. Collect data θ j (k) and ρ k j , θ j (k) and ρ k j are the data containing system production information generated by the jth iteration.

[0036] Step (4), policy update to solve the optimal control gain and the worst external disturbance gain: update the control gain K1 F and the external disturbance gain K2F :

[0037] The update formula is:

[0038] θ j (k)L j+1 = p j (k)

[0039] = θ j (k) x [(vec(L1 j+1 )) T (vec(L2 j+1 )) T (vec(L3 j+1 )) T (vec(L4 j+1 )) T (vec(L5 j +1 )) T (vec(L6 j+1 )) T ] T

[0040] Learn L1 j+1 to L6 j+1 using the least squares method, and then update the control gain and external disturbance gain:

[0041]

[0042] where (K1 F ) j+1 is the control gain obtained in the j+1th iteration, and (K2 F ) j+1 is the external disturbance gain obtained in the j+1th iteration.

[0043] The target strategy u Δk j , w Δk j will be obtained by step (2):

[0044]

[0045] Combining will have Simplify and combine the Kronecker product to have θ j (k)L j+1 = p k j , where: L j+1 = [(vec(L1 j+1 )) T(vec(L2 j+1 )) T … (vec(L5 j+1 )) T (vec(L6 j+1 )) T ] T , L1 j+1 = P j+1 , L2 j+1 = H Zu j+1 , L3 j+1 = H Zw j+1 , L4 j+1 = H uu j+1 - R, L5 j+1 = (H uw j+1 ) T , L6 j+1 = H ww j+1 + γ 2 I, θ j (k) = [θ1 j (k) θ2 j (k) θ3 j (k) θ4 j (k) θ5 j (k) θ6 j (k)],

[0046] Z k is the state of the augmented state space model at time k, u Δk is the input of the augmented state space model at time k, w Δk is the disturbance of the augmented state space model at time k, Q, R are positive definite matrices with dimensions matching the state Z k , the input u Δk , respectively, γ ≥ 0, and γ represents the level of persistent disturbance attenuation, I is the identity matrix, is composed of the system matrix {A, B, C, D} of the original system, α is the fault coefficient, P j+1 is the P obtained at the j+1th iteration, the matrix H uu j+1 , H uw j+1 , H ww j+1 , H Zw j+1 , H Zu j+1 is the matrix H j+1 derived from the Q function.(K1 F ) j is the control gain obtained in the jth iteration, F ) j is the external disturbance gain obtained in the jth iteration.

[0047] Step (5), if the iteration end condition is reached, then the iteration ends, otherwise go to step (4) to continue the iteration.

[0048] When l > 0 (l is a very small positive integer), the iteration can be stopped, otherwise let j = j + 1 go back to step (4) to continue the algorithm, F ) j and (K1 F ) j+1 are the control gains obtained in the jth and (j+1)th iterations, respectively, F ) j and (K2 F ) j+1 are the external disturbance gains obtained in the jth and (j+1)th iterations, respectively.

[0049] Example 1:

[0050] This example uses an injection molding process to specifically illustrate the method of the present application. The injection molding process is a process of converting plastic particles into various products, which mainly includes three stages: injection, pressure holding and cooling forming. In order to ensure the quality of the products produced by the injection molding process and the production efficiency of the products, the relevant process variables should be required to change as much as possible according to the expected set value at each production stage. In the production process, since the injection speed has a great influence on the quality of the final product, high-precision control of the injection speed is required in the injection stage, and the corresponding variable should be controlled to a given set value. The injection speed response of the proportional valve can be identified as an autoregressive model:

[0051]

[0052] The specific steps of the algorithm are as follows:

[0053] Step 1: Select the appropriate state variable x k , the following discrete state space model is obtained to represent the injection molding process:

[0054]

[0055] where x k+1 , x k , u k , w k , yk are the input, disturbance and output of the system at time k, respectively.

[0056] Consider the augmented state space model containing tracking error and state increment when the system has actuator faults and select the controller parameters as R = 0.01, γ = 1.5, determine the performance index function.

[0057] Step 2: Optimal control input and worst external disturbance design: According to the performance index function, the value function and Q function are proposed, and the expressions of the corresponding optimal control input, worst external disturbance, optimal control gain and worst external disturbance gain are constructed;

[0058] u Δk =(H uu -H uw (H ww ) -1 H wu ) -1 (H uw (H ww ) -1 H wZ -H uZ )Z k

[0059] K1 F =(H uu -H uw (H ww ) -1 H wu ) -1 (H uw (H ww ) -1 H wZ -H uZ )

[0060] w Δk =(H ww -H wu (H uu ) -1 H uw ) -1 (H wu (H uu ) -1 H uZ -H wZ )Z k

[0061] K2 F =(H ww -H wu (H uu ) -1 H uw )-1 (H wu (H uu ) -1 H uZ -H wZ )

[0062] where matrix H uu ,H uw ,H ww ,H wu ,H wZ ,H uZ is the component of matrix H derived from Q function.

[0063] Step3: Initialization and data collection: Given the initial control gain (K1 F ) 0 and external disturbance gain (K2 F ) 0 , collect data θ j (k) and ρ k j (K1 F ) 0 , (K2 F ) 0 are the initial control gain and external disturbance gain, respectively, θ j (k) and ρ k j are the data containing system production information generated by the jth iteration.

[0064] Step4: Strategy update to solve optimal control input and worst external disturbance: Update control gain K1 F and external disturbance gain K2 F by reinforcement learning algorithm.

[0065] The update formula is:

[0066]

[0067] L j+1 = [(vec(L1 j+1 )) T (vec(L2 j+1 )) T … (vec(L5 j+1 )) T (vec(L6 j+1 )) T ] T

[0068] where L1 j+1 = P j+1 , L2 j+1 = HZu j+1 , L3 j+1 = H Zw j+1 , L4 j+1 = H uu j+1 - R, L5 j+1 = (H uw j+1 ) T , L6 j+1 = H ww j+1 + γ 2 I.

[0069] θ j (k) = [θ1 j (k) θ2 j (k) θ3 j (k) θ4 j (k) θ5 j (k) θ6 j (k)]

[0070] wherein,

[0071] wherein, Z k is the state of the augmented state space model at time k, u Δk is the input of the augmented state space model at time k, w Δk is the disturbance of the augmented state space model at time k, Q, R are positive definite matrices with dimensions matching the state Z k , the input u Δk , γ ≥ 0, and I is the identity matrix, is composed of the system matrix {A, B, C, D} of the original system, α is the fault coefficient, P j+1 is the P obtained at the j+1th iteration, the matrix H uu j+1 , H uw j+1 , H ww j +1 , H Zw j+1 , H Zu j+1 is composed of the matrix H j+1 obtained by the Q function, j+1 refers to the j+1th iteration, (K1 F ) j is the control gain obtained at the jth iteration, (K2 F) j is the disturbance gain obtained in the jth iteration.

[0072] L1 is learned by least square method j+1 to L6 j+1 , and the control gain and the external disturbance gain are updated as follows:

[0073]

[0074] wherein (K1 F ) j+1 is the control gain obtained in the j+1th iteration, (K2 F ) j+1 is the external disturbance gain obtained in the j+1th iteration.

[0075] Step 5: if the iteration end condition is reached, the iteration is ended, otherwise, go to step (4) to continue the iteration.

[0076] When l>0 (l is a very small positive integer), the iteration is stopped, otherwise, let j=j+1 return to step (4) to continue the execution of the algorithm, (K1 F ) j and (K1 F ) j+1 are the control gains generated in the jth and j+1th iterations respectively, (K2 F ) j and (K2 F ) j+1 are the external disturbance gains generated in the jth and j+1th iterations respectively.

[0077] Given the initial H, the initial (K1 F ) 0 , (K2 F ) 0 can be obtained by learning, and the optimal H * in the Q function and the optimal (when the fault coefficient α=0.6) can be obtained respectively.

[0078]

[0079] Then the reinforcement learning algorithm is implemented, and after multiple learning, the matrix H and K1 F , K2 F obtained by the fault-tolerant control method gradually converge to the optimal H * and the optimal K1 F* , K2 F* .

[0080] Figures 1 to 5The control effects obtained in the experiment are shown respectively. It can be seen from Figure 1 and Figure 2 that, in the learning process, the gain K i F increases gradually to the optimal gain K i F* , i = 1, 2. In addition, with the increase of the iteration number, the matrix H * increases gradually to the optimal H * . In order to highlight the superiority of the control effect of the present application, the present application also compares the control effects of the present application and the traditional model-based fault-tolerant control method under different fault conditions, and the experimental results correspond to Figures 3 to 5 . Specifically, for the model u F (k) = αu(k), the normal condition and the fault condition can be discussed respectively.

[0081] Case 1: normal condition; α = 1, at this time the system is in a normal condition, and the control method proposed in the present application and the model-based fault-tolerant control method as a comparison can achieve good control effect.

[0082] Case 2: fault condition; this condition can be divided into two conditions: one is 0 < α < 1, and the other is α > 1. In terms of the experimental object, the actual injection speed will be less than the planned injection speed under the fault condition of 0 < α < 1. The actual injection speed will be greater than the planned injection speed under the fault condition of α > 1.

[0083] Under the fault condition, no matter which fault above, the present application can cope with it, while the traditional model-based fault-tolerant control method can only cope with the condition of 0 < α < 1, and has no way to deal with the condition of α > 1.

[0084] Figure 3 is the control effect of the present application when the system is in a normal condition. It can be found that both the fault-tolerant control method proposed in the present application and the traditional model-based fault-tolerant control method can achieve good control effect.

[0085] It can be found from Figure 4 that, in the face of the expected output curve, the system output curve of the traditional model-based fault-tolerant control method appears obvious jitter or even overshoot before the step, but it can also achieve tracking eventually. The output curve of the control method proposed in the present application is obviously smoother, which can track the expected output trajectory faster, and the tracking error image can also reflect the same conclusion. In addition, under the fault condition, it can be seen obviously that the input curve amplitude of the method proposed by us is large, which means that it can resist larger external disturbance.

[0086] Figure 5Therefore, it is shown that the method proposed in the application can track the expected output curve well, and the traditional fault-tolerant control method cannot do this. The tracking error curves of the two show that the method can still converge to 0 when α>1, and the tracking error of the traditional model-based fault-tolerant control method cannot do this, and the tracking error image of the latter oscillates more and more with the increase of time. The input curve comparison shows that the data-driven control method can quickly achieve control effect, while the input curve of the model-based fault-tolerant control method oscillates constantly. In addition, compared with the smooth external disturbance curve of the fault-tolerant control method proposed in the application, the external disturbance of the one-dimensional traditional fault-tolerant control increases with time, and the fluctuation amplitude also increases, which cannot play the effect of resisting external disturbance.

[0087] In summary, the present embodiment takes the injection molding process as an example to verify the effectiveness and feasibility of the control effect of the application. This fault-tolerant control method based on reinforcement learning can cope with the situation of actuator failure and external disturbance in the environment where the model parameters are unknown through data-driven method, and can achieve better control effect than the traditional model-based fault-tolerant control method. The experimental comparison results also show that the application has a wider range of application, and widens the range of actuator failure that can be coped with in the presence of external disturbance. Therefore, the application provides a new fault-tolerant control method for industrial processes in the environment with actuator failure and external disturbance, which can be more practical than the model-based fault-tolerant control method, has a wider range of application, ensures the safety and quality of industrial process manufacturing products, and can create greater value in actual industrial processes.

[0088] Finally, it should be noted that: the above only describes the preferred embodiments of the application and is not used to limit the application. For those skilled in the art, the technical solutions described in the above embodiments can be modified, or some technical features can be replaced with equivalent ones. Any modification, equivalent replacement, improvement, etc. within the spirit and principles of the application shall be included in the protection scope of the application.

Claims

1. A fault-tolerant control method for industrial process min-max optimization based on reinforcement learning, characterized in that: The method comprises the following steps: (1) establishing an augmented state space model containing tracking error and state increment on the basis of a state space model of an original system with actuator faults and external disturbances, and proposing a performance index function according to the augmented state space model; (2) proposing a value function and a Q function according to the performance index function, and constructing expressions of corresponding optimal control input, worst external disturbance, optimal control gain and worst external disturbance gain; (3) the initial control gain that stabilizes the system the external disturbance gain and collect data θ j (k) and ρ k j wherein, are the initial control gain and the external disturbance gain, respectively, θ j (k) and ρ k j is the data produced by the jth iteration that contains information about the system production; (4) updating the control gain K1 by reinforcement learning F , the external disturbance gain K2 F ; (5) if an iteration end condition is reached, the iteration ends, otherwise, returning to step (4) to continue iteration; the augmented state space model containing tracking error and state increment in step (1) is: wherein, x Δk+1 is the difference between the state of the original system at k+1 time and k time, y Δk+1 is the tracking error of the original system at k+1 time; x Δk is the difference between the state of the original system at k time and k-1 time, y Δk is the tracking error of the original system at k time; u Δk is the iterative update rate of the original system at k time; w Δk is the difference between the external disturbance of the original system at k time and the external disturbance at k-1 time; is the system matrix matched with the dimension of {Z k , u Δk , w Δk}, and the {A, B, C, D} constituting is the system matrix of the original system, I is the unit matrix; α is the fault coefficient; Z k is the state of the augmented state space model at k time, u Δk is the input of the augmented state space model at k time, w Δk is the external disturbance of the augmented state space model at k time; the performance index function proposed in step (1) on the basis of the augmented state space model is: wherein Z i is the state of the augmented state-space model at the i-th time, u Δi is the input of the augmented state-space model at the i-th time, i = k, k + 1,..., ∞, w Δi is the external disturbance of the augmented state-space model at the i-th time; Q1, R1 are positive definite matrices with dimensions matching the state Z i , the input u Δi ; γ ≥ 0, and γ represents the sustained disturbance attenuation level.

2. The fault-tolerant control method for reinforcement learning based industrial process minimax optimization according to claim 1, characterized in that: the value function in step (2) is: where Z k is the state of the augmented state space model at time k, is a symmetric positive definite matrix, and the value function satisfies the condition J* is a performance index J to be implemented.

3. The fault-tolerant control method for reinforcement learning based industrial process minimax optimization of claim 2, wherein: the Q function in step (2) is: Q*(Z k , u Δk , w Δk ) = Z k T Q2Z k + u Δk T R2u Δk - γ 2 w Δk T w Δk + V*(Z k+1 , u Δk+1 , w Δk+1 ) wherein Z k is the state of the augmented state-space model at time k, u Δk is the input of the augmented state-space model at time k, w Δk is the external disturbance of the augmented state-space model at time k, γ ≥ 0, γ represents the persistent disturbance attenuation level, Q2, R2 are positive definite matrices matching the dimensions of the state Z k , the input u Δk , respectively, V*(Z k+1 , u Δk+1 , w Δk+1 ) is the value function of the augmented state-space model at time k+1.

4. The fault-tolerant control method for reinforcement learning based industrial process minimax optimization of claim 3, wherein: the expressions of the optimal control input, worst external disturbance, optimal control gain and worst external disturbance gain in step (2) are: u Δk = (H uu -H uw (H ww ) -1 H wu ) -1 (H uw (H ww ) -1 H wZ -H uZ )Z k w Δk = (H ww -H wu (H uu ) -1 H uw ) -1 (H wu (H uu ) -1 H uZ -H wZ )Z k where the matrix H uu ,H uw ,H ww ,H wu ,H wZ ,H uZ is a component of the matrix H derived from the Q function.

5. The fault-tolerant control method for reinforcement learning based industrial process minimax optimization of claim 4, wherein: The step (3) is given an initial control gain (K1 F ) 0 and an external disturbance gain (K2 F ) 0 ; collect data θ j (k) and ρ k j , θ j (k) and ρ k j generated by the jth iteration, which contains the system production information.

6. The fault-tolerant control method for reinforcement learning based industrial process minimax optimization of claim 5, wherein: the control gain and external disturbance gain are updated in step (4) according to the following formula: θ j (k)L j+1 = p j (k) = θ j (k) x [(vec(L1 j+1 ) T (vec(L2 j+1 ) T (vec(L3 j+1 ) T (vec(L4 j+1 ) T (vec(L5 j+1 ) T (vec(L6 j +1 ) T ] T wherein L1 j+1 = P j+1 , L2 j+1 = H Zu j +1 , L3 j+1 = H Zw j+1 , L4 j+1 = H uu j+1 - R, L5 j+1 = (H uw j+1 ) T , L6 j+1 = H ww j+1 + γ 2 I, Z k is the state of the augmented state space model at time k, u Δk is the input of the augmented state space model at time k, w Δk is the external disturbance of the augmented state space model at time k, Q2, R3 are positive definite matrices with dimensions matching the state Z k , the input u Δk , respectively, P j+1 is the P obtained at the j+1th iteration, is a symmetric positive definite matrix, the matrix H uu j+1 , H uw j+1 , H ww j+1 , H Zw j+1 , H Zu j+1 is a component of the matrix H j+1 derived from the Q function, j+1 refers to the j+1th iteration, I is the identity matrix, (K1 F ) j is the control gain obtained at the jth iteration, (K2 F ) j is the external disturbance gain obtained at the jth iteration.

7. The fault-tolerant control method for reinforcement learning based industrial process minimax optimization of claim 6, wherein: the control gain and external disturbance gain in step (4) are updated according to the following formula: where (K1 F ) j+1 is the control gain obtained in the j+1th iteration, (K2 F ) j+1 is the external disturbance gain obtained in the j+1th iteration.

8. The fault-tolerant control method for reinforcement learning based industrial process minimax optimization of claim 7, wherein: the iteration end condition in step (5) is: l>0 where l is a very small positive integer, (K1 F ) j and (K1 F ) j+1 are the control gains resulting from the jth and (j+1)th iterations, respectively, (K2 F ) j and (K2 F ) j+1 are the external disturbance gains resulting from the jth and (j+1)th iterations, respectively.

Citation Information

Patent Citations

  • Batch process 2D constraint fault-tolerance control method through infinite time domain optimization

    CN107976942A

  • Industrial process fault-tolerant control method based on data-driven Q-learning

    CN114035523A