Q-learning optimal tracking control method for time-delay batch processes with actuator faults
By using a non-policy Q-learning algorithm based on reinforcement learning, an optimal controller for the injection molding process is established, solving the optimal tracking control problem under actuator failure and time delay, and achieving efficient control of the injection molding process.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-27
- Publication Date
- 2026-04-10
AI Technical Summary
In the injection molding process, existing technologies struggle to effectively utilize data to design controllers, especially in cases of actuator failure and time lag, making it difficult to achieve precise and optimal tracking control.
A non-policy Q-learning algorithm based on reinforcement learning is adopted to establish a new system model through state increment and output error, design a time delay performance index function, and iteratively learn the optimal control gain matrix to achieve the resistance to actuator failure and the acquisition of the optimal control law.
In the event of actuator failure and time delay, it can effectively track the set value, overcome the dependence of traditional methods on precise models, realize data-driven optimal control, and improve production efficiency and product quality.
Smart Images

Figure QLYQS_1 
Figure QLYQS_44 
Figure QLYQS_45
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of industrial process control, and particularly relates to a time-delay batch process Q-learning optimal tracking control method with actuator failure. BACKGROUND
[0002] In the globalized and highly competitive market environment, various industries need to continuously improve production efficiency and product quality, and reduce production costs to maintain competitiveness. Therefore, the industrial production mode using advanced technologies such as digitization, intelligentization, and automation has become an indispensable trend. This production mode not only meets the above production needs, but also realizes efficient use of resources and sustainable development of the environment.
[0003] The production mode of an industrial process can be divided into a continuous production process and a batch production process. From the input of raw materials to the output of products, it is continuous and uninterrupted, and such an industrial process is referred to as a continuous production process. However, with the continuous development of society, the products available for selection are becoming more and more extensive, and market demand is changing rapidly, and people are increasingly inclined to small-scale multi-process production mode. The batch production process meets these needs, and it is mainly a processing process that repeatedly operates to obtain products. Unlike the continuous production process, the batch process has the unique properties of repeatability, rapidity, and low cost, and is widely used in various fields, such as aerospace, transportation, manufacturing, and the like. The proportion of batch production processes in the chemical, food, beverage, and pharmaceutical fields is high. By effectively controlling the batch production process, various industries can improve production efficiency, reduce production costs, improve product quality, and enhance market competitiveness. Therefore, it is very meaningful to study the control method of the batch production process for various fields.
[0004] In view of the fact that the model is often difficult to obtain in the injection molding production process, a large amount of data is generated and stored, and these data contain data in the time direction and batch direction. Therefore, in the absence of an accurate process model, how to effectively utilize the data to directly design a controller for the injection molding production process is very important, and therefore, a method relying on data is more suitable for the time-delay batch problem. Reinforcement learning has the advantage of optimizing the control of a complex system only relying on data without prior information, and since its development, reinforcement learning has a certain research on the optimal tracking control problem of a state time-delay system. Therefore, for a batch process with state time delay, it is very important to study a control method based on reinforcement learning, which can make the system not rely on a model and only rely on data to continuously learn an optimal control law. SUMMARY
[0005] The application is a Q learning optimal tracking control method for time-delay batch process with actuator failure, belongs to the technical field of industrial process control, overcomes the time-varying dynamic parameter limitation of traditional control method, and the specific steps are as follows: step one: the state space expression form of time-delay batch process is described, and a new system model composed of state increment and output error is established on the basis; step two: a time-delay performance index function is introduced, and a control law capable of resisting partial actuator failure in the time-delay environment is designed; step three: a non-strategy Q learning algorithm with actuator failure is proposed, and the optimal control gain matrix is solved through continuous iterative learning.
[0006] The application is implemented through the following technical schemes:
[0007] The application is a Q learning optimal tracking control method for time-delay batch process with actuator failure, belongs to the technical field of industrial process control, overcomes the time-varying dynamic parameter limitation of traditional control method, and the specific steps are as follows: step one: the state space expression form of time-delay batch process is described, and a new system model composed of state increment and output error is established on the basis; step two: a time-delay performance index function is introduced, and a control law capable of resisting partial actuator failure in the time-delay environment is designed; step three: a non-strategy Q learning algorithm with actuator failure is proposed, and the optimal control gain matrix is solved through continuous iterative learning.
[0008] Step one: the state space expression form of time-delay batch process is described, and a new system model composed of state increment and output error is established on the basis;
[0009] Firstly, a kind of batch process with state time delay is considered:
[0010]
[0011] Wherein, t represents time, And u t =[u 1t u 2t ... u mt ] T ∈R mrespectively denote system states, system outputs, control inputs; A, A d , B, C denote system matrices of appropriate dimensions, R denotes a real matrix, n x , n y and m denote appropriate dimensions of real matrix R;
[0012] According to (1), the following iterative learning control law form is designed:
[0013] u t = u t-1 + u Δt (2)
[0014] where u Δt is the difference between the control input at time t and that at time t-1;
[0015] For the desired output trajectory y d , the tracking error variable at time t and the state error variable at time t-d, respectively, can be expressed as:
[0016] y Δt = y d - y t (3)
[0017] x Δt = x t - x t-1 (4)
[0018] x Δt-d = x t-d - x t-d-1 (5)
[0019] where x t , x t-1 , x t-d and x t-d-1 represent the state variables at time t, t-1, t-d and t-d-1, respectively;
[0020] According to (1) to (5), a new augmented model can be derived as follows:
[0021]
[0022] where, and are relevant matrices of appropriate dimensions, X t , X t-d and u Δtare the state variables of the new system model at time t, at time t-d and the control input at time t, respectively, R is a real matrix, and n+1 is the proper dimension of the real matrix R;
[0023] When the actuator fails, the control input u t is not always able to reach the desired value; for the case of actuator failure, it is mainly divided into three cases: partial failure, shutdown failure and stuck failure; this paper studies the phenomenon of partial actuator failure, defines the range of values of a to represent different types of failure, and uses the fault model as:
[0024]
[0025] wherein, α = diag[ α 1 α 2,…, α m ], It can be seen that, is the normal case of the actuator, a i = 0 is the complete failure case of the actuator, a i > 0 (a i ≠ 1) is the partial failure case of the actuator;
[0026] Then formula (6) can be rewritten as:
[0027]
[0028] Step 2: Introduce the time delay performance index function, and design a control law that can resist partial actuator failure in a time delay environment;
[0029] According to the above description of the time delay intermittent process with actuator failure, the following performance index can be designed:
[0030]
[0031] wherein, Q1 and Q2 are the weight matrices of states X i and X i-d , R is a positive definite matrix, and represents the weight of the control variable;
[0032] By finding the optimal control strategy, the system output y t can track the ideal reference trajectory y d ; therefore, the control strategy can be represented as:
[0033]
[0034] The Q function can be designed as follows:
[0035]
[0036] When the controller policy u Δt By comparing the value function and the Q function, it can be derived that the value function and the Q function are equal when the controller policy is optimal, as shown in equation (12):
[0037]
[0038] The quadratic function is defined as follows:
[0039] V(X t ,X t-d )=V(X t )+V(X t-d ) (13)
[0040] where V(X t )=X t T P1X t , δ(r,k)=X(t-d+i)-X(t-d+(i-1)),0≤i≤d;
[0041] When the controller policy is optimal, the quadratic form of the optimal time-delay value function can be described as:
[0042]
[0043] where, τ1=P1+d 2 P3,τ2=P2+(2dw-d)P3,τ n =P2+dP3,τ n+2 =-d 2 P3,τ 2n-1 =-dP3;
[0044] Similarly, the time-delay Q function can be expressed in the quadratic form as follows:
[0045]
[0046] Therefore, according to the relationship between equation (14) and equation (15), the relationship between the P matrix and the H matrix can be derived as follows:
[0047]
[0048] By analyzing the relationship between the P matrix and the H matrix, and substituting the expanded time-delay state-space equation (8) into equation (11) and equation (15), the H matrix can be expressed as:
[0049]
[0050] wherein, * represents the transpose value of the symmetric position, in order to simplify the expression, in the subscript of each component of the H matrix, x1 represents X t , x di represents X t-i ; wherein, (i=1, 2, …, d);
[0051] Step three: a non-strategy Q learning algorithm with actuator fault is proposed, and the optimal control gain matrix is solved by continuous iteration learning;
[0052] Using dynamic programming method, Bellman equation based on optimal Q function is obtained from formula (11), (12) and (15);
[0053]
[0054] In order to facilitate the expression, Bellman equation is further simplified as follows:
[0055]
[0056] In order to make full use of the data learned before, auxiliary variables are introduced in time delay system So the new time delay state space equation is:
[0057]
[0058] wherein, u Δt is the behavior strategy, which is used to generate system data is the target strategy, which is constantly optimized and updated by using the data generated by the behavior strategy, so that the target strategy converges to the optimal value;
[0059] Substitute formula (20) into formula (18) to get:
[0060]
[0061] wherein,
[0062] According to formula (20) and formula (21), the following can be further obtained:
[0063]
[0064] According to formula (22), the simplified optimal Bellman equation is as follows:
[0065]
[0066] According to the expression of Kronecker product, equation (23) can be rewritten as:
[0067]
[0068] where,
[0069]
[0070]
[0071]
[0072]
[0073]
[0074]
[0075]
[0076]
[0077]
[0078]
[0079]
[0080] Through the above calculation, the obtained controller gain matrix is as follows:
[0081]
[0082]
[0083] Data-based non-strategic Q-learning algorithm
[0084] Step 1: Implement the behavior policy on the system to generate the required data, and then store the data into and ;
[0085] Step 2: Select a suitable initial controller gain matrix, let and the initial value be zero;
[0086] Step 3: Use the data collected and stored in and , combined with equation (24) to solve L j+1 , update the controller gain matrix through equations (25) and (26) and
[0087] Step four: stop iteration when the condition and is satisfied, where ε (ε≠0) is a small positive number. If the condition is not satisfied, the algorithm jumps back to step two and continues.
[0088] Robust stability analysis
[0089] Theorem 1: For a time-delayed batch process with actuator faults described by equation (8), to design a robust optimal control input, the following conditions need to be established:
[0090]
[0091] where δ max (χ), λ min (χ) and λ max (χ) are the largest singular value, the smallest eigenvalue and the largest eigenvalue of χ, respectively, and P and W are two positive definite matrices and satisfy:
[0092]
[0093] where K F is obtained by combining and simplifying equations (25) and (26);
[0094] Proof: First, the actuator fault time-delayed batch process model described by equation (8) can also be equivalent to the following form:
[0095]
[0096] where is the delay operator;
[0097] The time-delayed batch process described by equation (8) is as follows:
[0098]
[0099] According to equation (29), the closed-loop time-delayed batch system is:
[0100]
[0101] The Lyapunov function is:
[0102]
[0103] Substituting equation (31) into equation (32), we can obtain:
[0104]
[0105] The above calculation process shows that:
[0106]
[0107] Therefore, combining the above inequalities can realize the inequality in (28). Wherein, Therefore, the conditions under which Theorem 1 holds were verified, and the robust stability analysis of the system was completed.
[0108] The advantages and effects of this invention are as follows:
[0109] This invention, primarily addressing time-delayed intermittent processes with actuator failures, proposes a novel non-strategy Q-learning algorithm. This algorithm learns the optimal controller gain matrix and thus the optimal control law solely from real-time measurement data generated by the system's dynamic process, without relying on system model parameters. The advantages of this algorithm are twofold: firstly, it overcomes the problem of traditional control methods relying on precise system model parameters; secondly, unlike other data-driven control methods, it finds the optimal controller gain matrix through trial and error even with unknown prior conditions. Furthermore, this method can still track the desired output trajectory and achieve good control performance even in situations with state delays and actuator failures. Attached Figure Description
[0110] Figure 1 This represents the convergence process of the H matrix when the fault factor α = 0.6;
[0111] Figure 2 This represents the convergence process of the H matrix when the fault factor α = 1;
[0112] Figure 3 This represents the convergence process of the H matrix when the fault factor α = 1.3;
[0113] Figure 4 This describes the convergence process of the controller gain K1 when the fault factor α = 0.6;
[0114] Figure 5 This describes the convergence process of the controller gain K2 when the fault factor α = 1;
[0115] Figure 6 This describes the convergence process of the controller gain K3 at the fault factor α = 1.3;
[0116] Figure 7 Tracking curve of control output y when fault factor α = 0.6;
[0117] Figure 8 Tracking curve of control output y when fault factor α = 1;
[0118] Figure 9 Tracking curve of the control output y at fault factor a = 1.3. DETAILED DESCRIPTION
[0119] In order to further illustrate the present application, it will be described in detail below in conjunction with the accompanying drawings and examples, but they should not be understood as limiting the scope of protection of the present application.
[0120] Example 1:
[0121] Injection molding is a typical intermittent production process, which plays an important role in the plastic product industry, which is mainly composed of three stages of filling, packaging and cooling, among which the packaging stage is an important stage to determine the quality of the product, and the injection speed is a key variable, which affects the flow behavior of the melt in the cavity. In order to ensure product quality, the injection speed should be controlled within a given profile. In this part, in order to verify that the non-strategic Q-learning algorithm proposed in this paper has good control effect in the time-delay intermittent process with actuator failure, in this part, taking injection molding as an example, the proposed control method is used to control the injection speed parameter in the injection molding production process.
[0122] According to a large number of experiments, the response of the injection speed to the proportional valve is identified as an autoregressive model, which is converted into a state space model as shown in equation (35):
[0123]
[0124] We will analyze the tracking control effect of the proposed algorithm on the time-delay intermittent process with actuator failure through three cases of a. Case 1: 0 < a < 1, i.e. the case where the actuator does not exist; Case 2: a = 1, i.e. the actual injection speed is faster than the optimal case; Case 3: a > 1, i.e. the injection speed is lower than the optimal case. Here, through a large number of experimental data, it can be determined that the parameters in the injection stage performance index are Q1 = Q2 = diag[2, 2, 2], R = 1, and the set value of the output is y d = 40. Given the initial H matrix, we find that under any condition, after multiple iterations of learning, the matrix H j , and gradually converges to the optimal matrix H, and The optimal matrix is represented as:
[0125] Case 1: a = 0.6
[0126]
[0127]
[0128]
[0129] Case 2: α = 1
[0130]
[0131]
[0132]
[0133] Case 3: α = 1.3
[0134]
[0135]
[0136]
[0137] exist Figure 1 , 2 In Figures 2 and 3, they represent the convergence results of the H matrix in the three cases of α = 0.6, α = 1, and α = 1.3, respectively. The results show that in these three cases, the H matrix gradually converges to the optimal H, almost unaffected by actuator failures. Figure 4 , 5 Figures 6 and 7 represent the controller gain for the cases where α = 0.6, α = 1, and α = 1.3, respectively. i=1,2 and the optimal controller gain The change in the absolute value of the difference between i=1 and i=2. The simulation results show that, regardless of the actuator failure condition... Both i=1 and i=2 show good convergence. Clearly, the controller gain matrix obtained in this paper is unaffected by actuator faults. The purpose of this paper is to verify that the proposed algorithm can track the ideal setpoint. Figure 7 , 8 Figures 9 and 1 represent the system output curves under three different actuator failure conditions. It can be seen that regardless of the failure condition, although the system output exhibits significant oscillations and overshoot in the initial stage, it eventually tracks the desired output trajectory as time increases. Simulation results demonstrate the feasibility and effectiveness of the proposed method.
[0138] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A time-delayed batch process Q-learning optimal tracking control method with actuator failure, the specific steps are as follows: Step one: the state space expression form of the time-delayed batch process is described, and a new system model composed of state increment and output error is established on this basis; First, consider a class of batch processes with state time delay: wherein Indicates time, , and These represent the system state, system output, and control input, respectively. Indicates the lag time. , , , Representing a system matrix of appropriate dimension, Represented as a real matrix, , and Represented as a real matrix The appropriate dimension; According to formula (1), the following iterative learning control law form is designed: wherein is the difference between the control input at the time instant For the desired output trajectory At time t, the tracking error variable and the state error variable, respectively, can be expressed as: At time t, the tracking error variable and the state error variable, respectively, can be expressed as: At time t, the tracking error variable and the state error variable, respectively, can be expressed as: wherein, , , and represent the state variables at the time instants , , and , respectively. According to formulas (1) to (5), a new augmented model can be derived as follows: wherein , , , , , , , and are appropriately dimensioned correlation matrices, , and are state variables of the new system model at time , at time and control input at time , is represented as a real matrix, is represented as a real matrix of appropriate dimension; When the actuator fails, the control input of the system The desired value is not always achieved; for the case of actuator failure, it is mainly divided into three cases: partial failure, shutdown failure and stuck failure; by studying the phenomenon of partial actuator failure, the value range of is defined to represent different types of failures, and the failure model is adopted: wherein, , , , , , represents the lower limit of actuator failure, represents the upper limit of actuator failure, it can be seen that, is the normal case of the actuator, is the case of complete failure of the actuator, and is the case of partial failure of the actuator; Then formula (6) can be rewritten as: Step two: a control law that can resist partial actuator failure in a time-delayed environment is designed by introducing a time-delay performance index function; According to the above description of the time-delayed batch process with actuator failure, the following performance index can be designed: wherein, and are weight matrices for states and respectively, is a positive definite matrix representing the control variable weight; By finding the optimal control strategy, to ensure that the system output is able to track the ideal reference trajectory ; thus, the control strategy can be expressed as: wherein represents a current time state-related control gain, represents a lag state-related control gain; By comparing the value function, the Q function can be designed as follows: When the controller policy At optimality, by comparing the value function with the Q-function, it can be derived that the value function and the Q-function are equal, as in equation (12): wherein, represents the optimal control strategy for the current time increment; Define the following quadratic function: wherein , , denotes a positive definite matrix related to the current time state, denotes a positive definite matrix related to the lag state, denotes a positive definite matrix related to the lagged incremental state, , ; When the controller strategy is optimal, the quadratic form of the optimal time delay value function can be described as: wherein , , , , , , ; Similar to the time delay value function, the time delay Q function can be expressed in the form of a quadratic form as follows: Thus, from the relationship between Equation (14) and Equation (15), it can be derived that The relationship between the matrices and The relationship between the matrices and By analyzing The matrix and The expanded time-delay state-space equation (8) is substituted into equation (11) and equation (15) by analyzing the relationship between the matrices, The matrix can be expressed as: in, , , , , , , The transpose value representing the symmetrical position is used for simplification. The subscripts of the components of the matrix are represented by... express ,use express ;in, ; Step three: a non-strategy Q-learning algorithm with actuator failure is proposed, and the optimal control gain matrix is solved by continuous iterative learning; Using the dynamic programming method, the Bellman equation based on the optimal Q function is obtained from formulas (11), (12) and (15); In order to facilitate the description, the Bellman equation is further simplified as follows: In order to make full use of the previously learned data, auxiliary variables are introduced in the time-delay system Thus, the new time-delay state space equation is obtained as wherein, , , , is a behavior strategy for generating system data, is a target strategy that is constantly optimized and updated by using the data generated by the behavior strategy, so that the target strategy converges to an optimal value; Substitute formula (20) into formula (18) to obtain: wherein ; According to formula (20) and formula (21), can be further obtained (22) According to formula (22), the simplified optimal Bellman equation is as follows: According to the expression of the Kronecker product, formula (23) can be rewritten as: wherein, Through the above calculation, the controller gain matrix obtained is as follows: This algorithm can obtain the optimal controller gain of the injection molding process through multiple learning of the data generated in the injection molding process. The optimal control law under the performance index can be obtained through the controller gain, and then it can be applied to the actuator control system to make the output of the system gradually track the set value.
Citation Information
Patent Citations
Multi-agent consistency reinforcement learning method and system based on improved Q function
CN114545777A
De-orbit strategy optimal tracking control method for two-dimensional state time-delay batch processing process
CN115327903A