Optimal tracking control of industrial processes based on multiscale singularly perturbed systems
By decomposing a multi-timescale singular perturbation system into fast and slow subsystems and utilizing ridge regression reinforcement learning, an optimal control strategy was designed. This solved the problems of unknown system model and unmeasurable slow state, achieving suboptimal tracking control performance for the system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-16
- Publication Date
- 2026-04-07
Smart Images

Figure CN116009403B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a tracking control industrial process method, and more particularly to an optimal tracking control industrial process method based on a multi-timescale singular perturbation system. Background Technology
[0002] Optimal tracking control has become a fundamental problem in control theory and engineering applications due to its practical significance, as it aims to enable the system to track an ideal trajectory by designing a controller with minimal cost and energy consumption. In industrial process applications of optimal tracking control, multi-unit devices and operating processes typically involve fast and slow modes, meaning that industrial processes exist across multiple time scales, such as mixing and separation thickening processes and catalytic cracking processes. To control singularly perturbed systems, singular perturbation theory and its variants have played a dominant role in designing controllers that completely separate the fast and slow subsystems.
[0003] Various methods for solving the optimal tracking control problem have indeed been reported, but most of them apply to normal linear systems with the same time scale. Designing optimal tracking control strategies for multi-timescale systems has attracted considerable attention from researchers and engineers. It is worth noting that most existing singular perturbation methods require prior knowledge of the system model. Even when the system model is known, the truly different time scales at which the system operates pose significant challenges to designing optimal tracking control strategies, and the challenge is even greater without information about the system model.
[0004] In recent years, some preliminary attempts have laid the foundation for designing optimal tracking controllers for fast and slow-mode systems using data-driven methods. In these studies, reinforcement learning techniques are combined with singular perturbation theory for setpoint tracking in industrial processes with multiple time scales. However, it is worth noting that the more complex inter-variable couplings and the unmeasurability of slow states have not been addressed. Summary of the Invention
[0005] The purpose of this invention is to provide an industrial process method for optimal tracking control of singularly perturbated systems across multiple time scales. This invention designs an optimal tracking control strategy for singularly perturbated systems, thereby enabling the system to achieve an ideal trajectory through a suboptimal method; thus solving the optimal tracking problem constrained by singularly perturbated systems. Furthermore, a novel data-driven ridge regression reinforcement control technique is proposed to reduce the influence of perturbation parameters on the system, and a control strategy for the decomposed slowly varying subsystem is designed using measurable data.
[0006] The objective of this invention is achieved through the following technical solution:
[0007] An industrial process method based on optimal tracking control of multi-timescale singular perturbation systems, the method comprising the following steps:
[0008] Step 1: Establish the dynamic model of the linear discrete-time multi-timescale high-order system and the linear dynamic model of the reference trajectory, and decompose the global system based on singular perturbation theory;
[0009] Step 2: For the ideal setpoint, a tracking control problem is proposed, and the problem is processed into an optimization control problem constrained by the subsystem. The optimal theoretical solution corresponding to the subsystem is designed.
[0010] Step 3: Considering the unknown slow state and corresponding dynamic information of the singular perturbation system, a novel data-based ridge regression non-policy reinforcement learning algorithm is designed to design the optimal control strategy by using the output and input data of multiple historical moments and introducing a non-policy iterative learning mode.
[0011] The aforementioned industrial process method for optimal tracking control based on multi-timescale singular perturbation systems, in step 1, considers linear discrete singular perturbation systems at two time scales:
[0012] (1)
[0013] in, and They represent The status and control input at any given time, For perturbation parameters, and Let these represent the system matrix and the control matrix, respectively, and the matrix... It is unknown; This indicates the fast sampling time.
[0014] The optimal tracking control industrial process method based on multi-timescale singular perturbation systems, wherein the dynamic model of the ideal reference trajectory in step 2 is as follows:
[0015] (9)
[0016] in, for Ideal reference trajectory at all times The coefficient matrix is the reference trajectory.
[0017] The aforementioned industrial process method for optimal tracking control of multi-timescale singular perturbation systems, in step 3 where the dynamics of the slow system are unknown, employs a data-driven online learning approach, introducing a behavioral strategy to generate data from the system, and a target strategy to be learned. The slow-variable augmented subsystem is rewritten as follows:
[0018] (twenty three)
[0019] in , It is a behavioral strategy. It is a target strategy.
[0020] The industrial process method for optimal tracking control based on multi-timescale singular perturbation systems, wherein step 3 proposes a non-policy ridge regression reinforcement learning algorithm, which includes the following steps:
[0021] Step 3.1: Collect overall system variable data and reference trajectory values at each time point by using behavior control strategies and adding detection noise;
[0022] Step 3.2: Determine the permissible initial control gain and set the iteration index. , ,from , start;
[0023] Step 3.3: Based on the system variable data and equations (28) and (30), evaluate the performance strategies of the corresponding subsystems respectively;
[0024] Step 3.4: Update the policy based on the parameter data learned from the policy evaluation and equations (29) and (31);
[0025] Step 3.5: Determine the learned control gain and Has the convergence returned to the previous time step? If so, proceed to the next step; otherwise... , Return to step 3.3.
[0026] The advantages and effects of this invention are:
[0027] This invention proposes a data-driven reinforcement learning technique for tracking systems with singular perturbations across multiple time scales, addressing the optimal tracking control problem. By employing singular perturbation theory and reinforcement learning, the system is decomposed into fast and slow subsystems. Optimal control strategies are designed for each subsystem, leading to a combined suboptimal control strategy for the singularly perturbated system. This solves the optimal tracking problem constrained by singular perturbations. Furthermore, a novel data-driven ridge regression reinforcement learning control technique is proposed, reducing the impact of perturbation parameters on the system. Specifically:
[0028] 1) A non-policy ridge reinforcement learning method combining singular perturbation theory is proposed, and an optimal tracking control strategy for the singular perturbation system is designed, so that the system can achieve the ideal trajectory through a suboptimal method;
[0029] 2) In terms of system model, the system is a discrete singular perturbation system, which means it is affected by some small parameters, and there is a mutual coupling relationship between the variables of the fast mode system and the slow mode system;
[0030] 3) Considering that in actual processes, the state variables in slow-mode systems are unmeasurable or have high measurement costs, some mathematical transformations and measurable output-input data of singular perturbation systems are used to replace the difficult-to-measure slow state data. The control strategy of the decomposed slow variable subsystem is designed using the measurable data. Attached Figure Description
[0031] Figure 1 This is a numerical example of the application of an embodiment of the present invention and a diagram illustrating the learning process of the controller gain of Algorithm 2.
[0032] Figure 2 This is a numerical example of the application of an embodiment of the present invention and a system output variable trajectory diagram of implementing Algorithm 2;
[0033] Figure 3 This is a numerical example of the application of the present invention and a diagram illustrating the learning process of the system control input for implementing Algorithm 2.
[0034] Figure 4 This is a practical example of the application of an embodiment of the present invention and a diagram illustrating the learning process of the controller gain in implementing Algorithm 2;
[0035] Figure 5 This is a practical example of the application of an embodiment of the present invention and a system output variable trajectory diagram for implementing Algorithm 2;
[0036] Figure 6 This is a practical example of the application of an embodiment of the present invention and a diagram illustrating the learning process of the system control input for implementing Algorithm 2. Detailed Implementation
[0037] An embodiment of the present invention will be further described below with reference to the accompanying drawings.
[0038] In this embodiment of the invention, a reinforcement learning technique for tracking multi-timescale singular perturbation systems based on data includes the following steps:
[0039] Step 1: Consider a linear discrete singular perturbation system with two time scales:
[0040] (1)
[0041] in, and They represent The status and control input at any given time, For perturbation parameters, and Let these represent the system matrix and the control matrix, respectively, and the matrix... It is unknown; This indicates the fast sampling time.
[0042] As can be seen, the system in equation (1) is a multi-rate singular perturbation system due to the presence of perturbation parameters. When solving the optimal tracking control problem, the presence of perturbation parameters can lead to ill-conditioned numerical problems, excessive dimensionality, and large computational load during the design process when designing the controller for the system. In order to reduce the impact of perturbation, we decompose it into a fast subsystem and a slow subsystem according to the singular perturbation theory, as follows:
[0043] By ignoring the rapid changes in the system, we have:
[0044] (2)
[0045] in, , Indicates in The system state of the slow subsystem at time t. This represents the control input of the slow subsystem.
[0046] To obtain the quasi-steady-state values of the fast state variables of the fast dynamic system in system equation (1), i.e. Requires the following assumptions: matrix It is a non-singular matrix, and has Based on the above conditions, we obtain the following formula:
[0047] (3)
[0048] Substituting equation (3) into equation (2), the slow subsystem is:
[0049] (4)
[0050] in , Since slow subsystems are difficult to handle if they are still operating on a fast timescale, they should therefore be operating on a slow timescale. Below, and have The slow subsystem equation (4) can first be rewritten as:
[0051] (5)
[0052] in , , By discretizing the system (5), the slow subsystem can be obtained:
[0053] (6)
[0054] in , .
[0055] Subtracting the slow subsystem dynamic equation (2) from (1), the fast subsystem dynamic equation is:
[0056] (7)
[0057] in This represents the control input of the rapidly changing subsystem.
[0058] After decomposition according to system formula (1), the variables of the slow subsystem, the variables of the fast subsystem, and the global system variables have the following relationships:
[0059] (8)
[0060] Step 2: For the ideal setpoint, propose a tracking control problem and verify the feasibility of solving the problem using the model-based singular perturbation decomposition method;
[0061] The dynamic model of the ideal reference trajectory is as follows: (9)
[0062] in, for Ideal reference trajectory at all times The coefficient matrix of the reference trajectory.
[0063] For the global system, the slowly changing subsystem, and the rapidly changing subsystem, their performance metrics are defined as follows:
[0064] (10)
[0065] (11)
[0066] (12)
[0067] in, , and These represent the performance metrics of the global system, the slowly changing subsystem, and the rapidly changing subsystem, respectively. , and They represent the slow-changing subsystem and the fast-changing subsystem respectively. Time control strategy , The discount factor is used to ensure the boundedness of performance. This represents the sum of all moments.
[0068] , and This indicates the global system, the slowly changing subsystem, and the quickly changing subsystem. The effect function at time t, Indicates that the slow subsystem is in Tracking error at any time, Indicates that the fast subsystem is in The system state at any given moment. , , , , This represents a symmetric matrix with appropriate dimensions.
[0069] And the following relationship exists: (3)
[0070] According to equation (6), the slow subsystem should be in a slow time scale. Therefore, the performance index corresponding to equation (6) for the slow subsystem is:
[0071] (14)
[0072] Augmented matrix processing is performed on the system variables and reference trajectory of the slow subsystem:
[0073] (15)
[0074] in .
[0075] Based on existing reinforcement learning results, the quadratic form of the value function of the slow subsystem is defined as follows: (16)
[0076] in, It is a positive definite matrix and satisfies the following algebraic Riccati equation:
[0077] (17)
[0078] And there are (18)
[0079] Based on dynamic programming theory, the Bellman equation corresponding to the slow subsystem is:
[0080] (19)
[0081] Define the Hamiltonian function of a slow subsystem based on the Bellman equation as follows:
[0082] (20)
[0083] by condition The optimal control strategy for a slow subsystem can be designed as follows:
[0084] (twenty one)
[0085] in Similarly, the optimal control strategy for the performance index of the fast subsystem is:
[0086] (twenty two)
[0087] in Positive definite matrix Satisfying the algebraic Riccati equations constrained by the tachyon system:
[0088] (twenty three)
[0089] Step 3: Use Algorithm 1 to learn the optimal control strategy when the system state and system dynamics model are all known. Considering that the slow state data of the system is not measurable, a non-policy-based ridge regression reinforcement learning algorithm 2 is adopted to learn the optimal control protocol through data generated by the global system. Based on the learned optimal control strategy, optimal tracking control is performed on the multi-timescale singular perturbation system, as follows:
[0090] Algorithm 1: Model-based policy iterative reinforcement learning
[0091] Step 3.1: Determine the iteration index , and stable initial allowable control gain , , from , start;
[0092] Step 3.2: Evaluate the strategy using the following Lyapunov equation; (twenty four)
[0093] (25)
[0094] Step 3.3: The policy update is iterated through the following equation;
[0095] (26)
[0096] (27)
[0097] in (28)
[0098] (29)
[0099] Step 3.4: When , ( , Stop when the number is a very small positive number; otherwise... , And return to step 3.2.
[0100] Based on the above algorithm, we can design a combined suboptimal controller. As can be seen from equation (27), the control strategy corresponding to the slow subsystem is still at a slow time scale, while the combined control strategy is at a fast time scale. Therefore, the control strategy of the slow system... Considered as being in the interval The innermost element is an invariant constant. Show no more than The largest integer.
[0101] We can see that Algorithm 1 requires the system dynamics to be known; however, in practice, due to the influence of environmental and other factors, the system dynamics... , , Since the system model is unknown or difficult to establish, we can use a data-driven online learning approach. This involves introducing behavioral policies to generate data, and the target policy to be learned. The slow-varying augmented subsystem can then be transformed into:
[0102] (30)
[0103] in, , It is a behavioral strategy. It is a target strategy.
[0104] Based on the quadratic form of the value function, Lyapunov's equation, and dynamic programming theory, we have:
[0105] (31)
[0106] It is worth noting the system status It cannot be measured, even if it is used through equation (8). Its approximate replacement, at the same time, the singular perturbation system (1) is also a multi-rate system, which poses a challenge to solving the unmeasurability of state variables.
[0107] To handle system state To address the problem of unmeasurability, we replace the unusable slow subsystem state with input-output data from multiple past historical moments using Equation (8) and the dynamic equation (15) of the slow-varying augmented subsystem. The state of the slowly varying augmented subsystem is:
[0108] (32)
[0109] in , , , Based on this, the augmented state in the Bellman equation, after being replaced, becomes:
[0110] (33)
[0111] The least squares method is commonly used to solve for the optimal control policy in equation (33). However, the least squares method has a drawback: overfitting. Furthermore, the particular solution obtained by the least squares method cannot be uniquely determined because the system cannot be easily and continuously excited using probe noise. Therefore, we add a disturbance to equation (33) and learn the control policy based on the newly defined Bellman equation using the least squares method. Specifically:
[0112] The rewritten expression (33) is: (34)
[0113] in , , , , , , , Add a perturbation In equation (34), It can be solved as:
[0114] (35)
[0115] Meanwhile, the controller can be designed as follows:
[0116] (36)
[0117] in .
[0118] Algorithm 2: Ridge Reinforcement Learning Based on Non-Policy
[0119] Step 3.1: Collect overall system variable data and reference trajectory values at each time point by using behavior control strategies and adding detection noise;
[0120] Step 3.2: Determine the permissible initial control gain and set the iteration index. , ,from ,
[0121] start;
[0122] Step 3.3: Based on the collected global system variable data and equation (35), evaluate the performance strategy corresponding to the slow subsystem, and evaluate the performance strategy corresponding to the fast subsystem according to equation (24);
[0123] Step 3.4: Update the strategy according to equations (36) and (28);
[0124] Step 3.5: Determine the learned control gain and Has the convergence returned to the previous time step? If so, proceed to the next step; otherwise... , Return to step 3.3;
[0125] In embodiments of the present invention, such as Figures 2-6 As shown, in order to more intuitively demonstrate the effectiveness of the data-based reinforcement learning method for solving the tracking problem of multi-timescale singular perturbation systems proposed in this invention, MATLAB software is used and two experimental examples are used to verify the effectiveness of the method proposed in this invention.
[0126] Numerical example: The discrete linear singular perturbation system is as follows: (37)
[0127] in , , , , , , , .
[0128] Practical example: Mixing, separation, and thickening process: Its continuous system is as follows:
[0129] (38)
[0130] In response to the discrete system corresponding to this invention, the continuous system can be discretized as follows:
[0131] (39)
[0132] In this embodiment of the invention, Figure 1 This is the learning process of the controller gain by applying numerical examples and implementing Algorithm 2. Figure 2 This is a system output variable trajectory graph that applies a numerical example and implements Algorithm 2. Figure 3 It is the learning process of the system control input applied to numerical examples and the implementation of Algorithm 2, from Figure 3It can be seen that the multi-timescale singular perturbation system proposed in this invention can keep up well with the ideal reference trajectory. Figure 4 This involves applying a practical example and implementing the learning process for the controller gain of Algorithm 2. Figure 5 This is a system output variable trajectory graph that applies a practical example and implements Algorithm 2. Figure 6 It is the learning process of system control input that applies practical examples and implements Algorithm 2. Similarly, from Figure 6 It can also be seen that Algorithm 2 proposed in this invention enables multi-timescale singular perturbation systems to keep up well with the ideal reference trajectory. This further enables slow-state, non-measurable multi-timescale singular perturbation systems to achieve tracking control in an approximately optimal manner.
Claims
1. An industrial process method based on optimal tracking control of multi-timescale singular perturbation systems, characterized in that, The method includes the following steps: Step 1: Establish the dynamic model of the linear discrete-time multi-timescale high-order system and the linear dynamic model of the reference trajectory, and decompose the global system based on singular perturbation theory; Step 2: For the ideal setpoint, a tracking control problem is proposed, and the problem is processed into an optimization control problem constrained by the subsystem. The optimal theoretical solution corresponding to the subsystem is designed. Step 3: Considering the unknown slow state and corresponding dynamic information of the singular perturbation system, a novel data-based ridge regression non-policy reinforcement learning algorithm is designed to design the optimal control strategy by using the output and input data of multiple historical moments and introducing a non-policy iterative learning mode. The dynamics of the slow system in step 3 are unknown. A data-driven online learning approach is used, introducing behavioral policies to generate data from the system. The target policy is then learned. The slow augmented subsystem is rewritten as follows: ; in , It is a behavioral strategy. It is a target strategy; The non-policy-based ridge regression reinforcement learning algorithm proposed in step 3 includes the following steps: Step 3.1: Collect overall system variable data and reference trajectory values at each time point by using behavior control strategies and adding detection noise; Step 3.2: Determine the permissible initial control gain and set the iteration index. , ,from , start; Step 3.3: Evaluate the performance strategies of the corresponding subsystems based on the system variable data, control gain, and Lyapunov equation; Step 3.4: Update the policy based on the parameter data learned from the policy evaluation, the slow-varying augmented subsystem, and the state of the slow-varying augmented subsystem; Step 3.5: Determine the learned control gain and Has the convergence returned to the previous time step? If so, proceed to the next step; otherwise... , Return to step 3.
3.
2. The industrial process method for optimal tracking control based on a multi-timescale singular perturbation system according to claim 1, characterized in that, Step 1 considers a linear discrete singular perturbation system at two time scales: (1) in, , and They represent The status and control input at any given time, For perturbation parameters, , and , Let these represent the system matrix and the control matrix, respectively, and the matrix... It is unknown; , This indicates the fast sampling time.
3. The industrial process method for optimal tracking control based on a multi-timescale singular perturbation system according to claim 1, characterized in that, The dynamic model of the ideal reference trajectory in step 2 is as follows: (9) in, for Ideal reference trajectory at all times The coefficient matrix is the reference trajectory.
Citation Information
Patent Citations
Uncertain singular perturbation system online multi-time-scale rapid adaptive control method
CN112506057A
Multi-time scale system optimal tracking control method based on reinforcement learning
CN115453884A