Batch process two-dimensional model-free minimum-maximum control method with high-order nonlinearity and non-repetitive interference

Through the modelless minimum maximum control method of two-dimensional off-orbit strategy, reinforcement learning and signal compensation technology are used to solve the problems of high-order nonlinearity and non-repetitive interference during the intermittent process, fast and accurate control strategy design is achieved, and the system's control effect and production efficiency are improved.

CN120406130APending Publication Date: 2025-08-01LIAONING UNIVERSITY OF PETROLEUM AND CHEMICAL TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510525047.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

Due to the deviation of control effects caused by higher-order nonlinear and non-repetitive interference during intermittent processes, it is difficult for traditional methods to establish a process model that combines accuracy and universality. Frequent changes in market demand lead to frequent adjustment of production parameters, and existing control strategies are difficult to adapt.

Method used

The modelless minimum control method of two-dimensional off-track strategy is adopted. Through reinforcement learning, unmodeled dynamics are solved and signal compensation is performed, the dependence of the system model is reduced, and the optimal control strategy is designed to improve system control and tracking performance.

Benefits of technology

Effectively reduce the impact of unmodeled dynamics on control strategies, improve the control and tracking performance of the system, accelerate convergence speed, reduce system model dependence, and improve production efficiency and cost optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005375028060000021
    Figure BDA0005375028060000021
  • Figure BDA0005375028060000022
    Figure BDA0005375028060000022
  • Figure BDA0005375028060000033
    Figure BDA0005375028060000033
Patent Text Reader

Abstract

A batch process two-dimensional model-free minimum-maximum control method with high-order nonlinearity and non-repetitive interference belongs to the technical field of industrial process control, and comprises the following specific steps: step 1, establishing a state-space equation of a batch process, and expanding the state-space equation into a linear model and an unmodeled dynamic state; 2, defining a Lyapunov function, namely a two-dimensional value function; and 3, expanding the optimal Q function into a quadratic form, and deriving the optimal Q function to obtain an optimal value. And 4, solving the optimal control strategy through iteration to enable the optimal control strategy to be finally converged to an ideal set value. And 5, designing a zero-sum game control algorithm with unmodeled dynamic compensation, and solving an optimal control law. According to the method, the problem that a traditional method for processing complex characteristics of a two-dimensional system is too tedious is solved, the unmodeled dynamic state is solved and compensated through the method based on reinforcement learning and independent of a system model, the influence of the unmodeled dynamic state on the control performance is reduced, and the control effect is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of industrial process control, and particularly relates to a two-dimensional model-free min-max control method for batch processes with high-order nonlinearity and non-repetitive disturbances, which reduces the influence of unmodeled dynamics on the system control effect through signal compensation. Background Art

[0002] An intermittent process is a production method that processes materials in discrete batches. Its core feature is the periodic operation of "start - run - stop", which plays a crucial role in the fields of fine chemicals, biopharmaceuticals, and food and beverage manufacturing. This production method involves processing a certain amount of substances in a set processing sequence through multiple stages of a machine to produce a fixed quantity of products. In the chemical industry, intermittent processes account for approximately 40 - 50% of the total, while in the pharmaceutical industry, this proportion is even higher than 80%, and in the fine chemical field, it is almost 100%. Intermittent processes are known for their rapid response, low-cost input, repeatability, diversity, and multi-stage operation. Compared with continuous processes where control methods are mature and diverse, the development of control methods for intermittent processes is still in an ascending stage.

[0003] However, in practical applications, the batch processing process is not as simple as we imagine. Intermittent process manufacturing processes generally have significant non-linear characteristics. When using linearized modeling methods, if the model accuracy is insufficient, significant deviations will occur, thereby weakening the operating efficiency of the control system. In an intermittent production system, process parameters often exhibit time-varying characteristics. Although the entire production system can maintain a dynamic stable state, it always lacks a fixed steady-state operating point, which brings significant difficulties to the formulation of control strategies. Such production processes essentially belong to a multi-disciplinary intersection field, and it is necessary to integrate professional knowledge such as machinery and chemical engineering and coordinate multiple internal and external variables during the modeling process. It should be noted that with the dynamic changes in market demand, product specifications and production capacity requirements often need to be adjusted synchronously. In actual operation, not only do we need to frequently replace production materials and process equipment, but we also need to dynamically adjust operating parameters to meet new production requirements. This makes it extremely challenging to construct a process model with both accuracy and universality. Therefore, in view of the above situation, a model-free min-max control method for batch processes with high-order nonlinearity and non-repetitive disturbances is designed to reduce the influence of unmodeled dynamics on the control effect through signal compensation without relying on process models and system initial parameters. Summary of the Invention

[0004] The present invention proposes a two-dimensional off-track strategy model-free minimax control method for batch processes with high-order nonlinearity and non-repetitive disturbances. This method can effectively solve the influence of high-order nonlinearity and non-repetitive disturbances existing in the system on the control effect, reduce the model dependence of the system, and continuously learn relying on the data in the time direction and batch direction. By combining reinforcement learning and signal compensation, the unmodeled dynamics are solved, a more accurate control strategy is established, the optimal control strategy is obtained, the control and tracking performance of the system are improved, and the convergence speed is accelerated.

[0005] The present invention is implemented through the following technical solutions:

[0006] The present invention proposes a two-dimensional off-track strategy model-free minimax control method for batch processes with high-order nonlinearity and non-repetitive disturbances. The unmodeled dynamics are solved by reinforcement learning and compensated to reduce their influence on the control strategy. First, a nonlinear state-space equation of the batch process with strong nonlinearity and non-repetitive disturbances is established. Second, the tracking error is extended as a state variable into the performance index to construct the performance index of the two-dimensional nonlinear system. Based on the designed performance index, the two-dimensional zero-sum game value function is derived. Subsequently, according to the relationship between the value function and the Q function, the Bellman equation is constructed. By solving the Bellman equation containing the Kronecker product, the optimal controller is obtained. Through adaptive compensation, the system effectively reduces the influence of the unmodeled dynamics. Further, a convergence analysis is carried out to prove that the proposed two-dimensional model-free method converges to the optimal solution obtained by solving the two-dimensional game algebraic Riccati equation. The present invention can effectively solve the problem that two-dimensional reinforcement learning is difficult to apply in nonlinear systems, greatly reduce the excessive dependence on the system model, solve the unmodeled dynamics through the signal compensation method, quickly and accurately obtain the optimal control strategy, improve the optimal performance of the system, and accelerate the convergence speed.

[0007] Step 1: Establish the state-space equation of the batch process, which is expanded into a linear model and unmodeled dynamics;

[0008] The batch process with unknown system dynamics is represented by a nonlinear state-space equation, and its expression form is as follows:

[0009]

[0010] where κ represents the time direction, τ represents the batch direction, x(κ, τ) ∈ R n represents the system state, u(κ, τ) ∈ R m represents the system control input, f(x(κ, τ), u(κ, τ), d(κ, τ)) ∈ R nThe function f denoted as x(κ,τ), C represents a system matrix with appropriate dimensions, R represents a real matrix, and n and m represent the appropriate dimensions of the real matrix R;

[0011] The batch model with high-order nonlinearity and non-repetitive disturbances is expanded into a combination of a low-order linear model and unmodeled dynamics through Taylor expansion. The nonlinear terms and non-repetitive disturbance terms are expanded into low-order linear terms and unmodeled dynamics. The tracking error is extended into the state space to construct a performance index for a two-dimensional nonlinear system with unmodeled dynamics:

[0012]

[0013] where, represents the state of the τ-th batch at the κ-th moment, Δ b x(κ,τ) = x(κ,τ) - x(κ,τ - 1) represents the increment of the state in the batch direction, V(κ - 1,τ) represents the value of the τ-th batch at the (κ - 1)-th moment, e(κ,τ) = y r (κ,τ) - y(κ,τ) represents the error between the set value and the current value of the τ-th batch at the κ-th moment, represents the state of the (τ - 1)-th batch at the (κ + 1)-th moment, V(κ,τ - 1) represents the value of the (τ - 1)-th batch at the κ-th moment, e(κ + 1,τ - 1) represents the error between the set value and the current value of the (τ - 1)-th batch at the (κ + 1)-th moment, ω1(κ,τ) = u(κ,τ) - u(κ,τ - 1) represents the input difference between the τ-th batch and the (τ - 1)-th batch at the κ-th moment, ω2(κ,τ) = V(κ,τ) - V(κ,τ - 1) represents the unmodeled dynamics difference between the τ-th batch and the (τ - 1)-th batch at the κ-th moment, I represents a system parameter matrix with appropriate dimensions;

[0014] When the output of the system can track the expected value, find the optimal control strategy K 11 , K 12 , K 21 , K 22 :

[0015]

[0016] Step 2: Define the Lyapunov function, that is, the two-dimensional value function, and assume it is equal to the two-dimensional optimal value function accumulated in the time and batch directions. Under the framework of the two-dimensional system, decompose the value function into horizontal and vertical components, and use the principle of asymptotic stability of the Lyapunov function to obtain the two-dimensional difference expression of the time-batch direction of the two-dimensional value function:

[0017]

[0018] The value function is split into an incremental form and accumulated to obtain the following form;

[0019]

[0020]

[0021] where, and represent the value functions in the time direction and batch direction respectively, and are the lengths of time and batch respectively,

[0022] Step 3: By expanding the optimal Q - function into a quadratic form and taking its derivative to obtain the optimal value; In a dynamic system with practical constraints, the optimal control problem based on the zero - sum game theory is formulated as a minimax optimization problem. Considering the case where the control input ω1 is in the optimal strategy (i.e., trying to minimize the cumulative cost) and the unmodeled dynamics ω2 present the worst - case disturbance (i.e., trying to maximize the cumulative cost), the optimal value function of the system can be described by the Hamilton - Bellman - Jacobi (HBJ) equation. Specifically, in the two - dimensional state space the explicit expression of the value function can be obtained by solving this partial differential equation, which usually has a quadratic form;

[0023]

[0024] where Q1, Q2 and R are and the weight matrices of ω1(κ,τ) respectively, and γ is the attenuation factor;

[0025] To find the optimal control law, by taking the derivative of the Q - function with respect to the control input strategy ω1 and the unmodeled dynamic input strategy ω2 and setting its derivative equal to zero at the extreme point, the optimal control input strategy and the unmodeled dynamic input strategy obtain the following equations:

[0026]

[0027]

[0028] where are all non - negative matrices;

[0029] The optimal control gain can be derived as:

[0030]

[0031]

[0032] Step 4: Iteratively solve the optimal control strategy to make it finally converge to the ideal set value; To solve the optimal control strategy and make it finally converge to the ideal set value, the method of policy iteration is adopted, where is the target policy, and numerical optimization is carried out in combination with the Bellman equation:

[0033]

[0034]

[0035] Step 5: Design a zero-sum game control algorithm with unmodeled dynamics compensation to solve the optimal control law; Solve it by iterating using production data. In a two-dimensional state space, if the system dynamics satisfy the Kronecker product structure, the matrix decomposition characteristics can be used to efficiently solve the Bellman equation, and finally the optimal control strategy is obtained;

[0036]

[0037]

[0038] Unbiased analysis of the proposed algorithm:

[0039] In the behavior strategy of the two-dimensional off-policy zero-sum game self-learning algorithm, exploration noises σ1(κ,τ) and σ2(κ,τ) are respectively added. Whether the added exploration noises are equal to zero or not, the value of the H matrix will not be affected. After adding the exploration noises, the state space expression becomes the following form:

[0040]

[0041]

[0042] By shifting the terms on both sides of the equation and substituting equation (16) into equation (15), the following can be obtained:

[0043]

[0044] By comparing equation (15) with equation (17), it can be found that the two are equal, and adding the exploration noises will not affect the final value, which proves the unbiasedness of the algorithm.

[0045] The advantages and effects of the present invention are:

[0046] This study focuses on the problem of model-free control for batch processes with unmodeled dynamics and innovatively proposes an unmodeled dynamics compensation method based on two-dimensional model-free min-max control. Different from traditional precise modeling methods and one-dimensional control modes, this method utilizes information from two dimensions, time and batch, in the production process. By reinforcement learning, the unmodeled dynamics are solved and compensated, reducing their impact on the control strategy. Based on the designed performance index, a two-dimensional zero-sum game value function is derived. Subsequently, according to the relationship between the value function and the Q function, the Bellman equation is constructed. By solving the Bellman equation containing the Kronecker product, the optimal controller is obtained. Through adaptive compensation, the system effectively reduces the impact of unmodeled dynamics. In particular, the designed zero-sum game self-learning algorithm is free from the constraints of initial system parameters and can directly extract the optimal control strategy from process data, improving industrial production efficiency and cost optimization while ensuring tracking performance. Description of the Drawings

[0047] Figure 1 Control gains for different batches under the two-dimensional model-free min-max control method Convergence process;

[0048] Figure 2 Control gains for different batches under the two-dimensional model-free min-max control method Convergence process;

[0049] Figure 3 Control gains for different batches under the two-dimensional model-free min-max control method Convergence process;

[0050] Figure 4 Control gains for different batches under the two-dimensional model-free min-max control method Convergence process;

[0051] Figure 5 Control output tracking curves for different batches under the two-dimensional model-free min-max control method;

[0052] Figure 6 Control input curves for different batches under the two-dimensional model-free min-max control method

[0053] Figure 7 Unmodeled dynamic input curves for different batches under the two-dimensional model-free min-max control method. Detailed Implementation Modes

[0054] To further illustrate the present invention, the present invention will be described in detail below in conjunction with the drawings and examples, but they should not be construed as limiting the protection scope of the present invention.

[0055] Example 1:

[0056] Injection molding technology is the core pillar technology of global plastic manufacturing. Its ability to produce quickly, efficiently, and flexibly update products makes it an indispensable process in the field of mass processing. This process mainly includes four stages: melting, injection, holding pressure, and cooling. Among them, the injection stage directly determines the quality and appearance of the final product. During the injection process, the precise control of injection speed and pressure is particularly crucial. These two mutually coupled parameters jointly dominate the filling behavior of the melt in the mold. Specifically, the injection pressure ensures complete filling of the mold cavity by affecting the density distribution of the plastic melt. To prevent defects such as short shots or flash, precise control of the nozzle pressure is required, which requires precise adjustment of the valve opening. This study proposes a two-dimensional offline self-learning algorithm incorporating unmodeled dynamic compensation, with the precise speed tracking in the injection stage as the control objective. This technology is expected to optimize production efficiency while improving product quality.

[0057] Based on a large amount of experimental data, the correlation equation between the nozzle pressure (NP) and the valve opening (VO) in the holding pressure stage of two-dimensional injection molding is established as follows:

[0058]

[0059] where NP(κ,τ) is the nozzle pressure at time κ and batch τ; VO(κ,τ) is the valve opening at time κ and batch τ; NP(κ - 1,τ) is the nozzle pressure at time κ - 1 and batch τ; VO(κ - 1,τ) is the valve opening at time κ - 1 and batch τ; d(κ,τ) is the unmodeled dynamic value at time κ and batch τ.

[0060] Definition: ω1(κ,τ) = VO(κ,τ) is the control input, and Y(κ,τ) = NP(κ,τ) is the system output.

[0061] Then, according to formula (18), the state-space model of the holding pressure stage is established:

[0062]

[0063] where After repeated experiments, the parameters of the holding pressure stage are determined as Q1 = Q2 = diag[2, 0, 0, 0, 1], R = 0.01, and γ = 1.5. However, in the actual batch production process, it is also very difficult to obtain the initial parameters of the system. Therefore, in the following research, we propose a zero-sum game self-learning algorithm based on a two-dimensional off-track strategy to design the optimal controller for the batch process.

[0064] From Figure 1-4It can be seen that in the third batch, the convergence trend of the control strategy is not ideal; by the sixth batch, there has been a significant improvement; after thirteen batches, it has completely converged and approached the optimal value. From the simulation results of the two-dimensional off-orbit strategy zero-sum game self-learning algorithm, it can be seen that this algorithm exhibits excellent tracking performance, and the tracking error shows a gradually decreasing trend as the number of batches increases.

[0065] Figure 5 The system output curve under the action of the proposed algorithm is shown. It can be seen from the figure that the output set value of the two-dimensional batch system is 20 bar. The current output has not yet reached the preset target value, and there is an error between the current output and the set value, but this error decreases as the number of batches increases. At the same time, the system shows enhanced tracking performance over time, and by the thirteenth batch, the system output completely matches the set value. Figure 6 The input curve of the two-dimensional system controller signal is shown, Figure 7 The results of the unmodeled dynamic input signal are presented. It can be observed from both figures that the input trajectory of the control variable shows a stable state. As the number of iterations accumulates, the proposed method gradually optimizes the control gain and finally approaches the optimal value.

[0066] In summary, taking the batch process control design with high-order nonlinearity and non-repetitive interference as an example, the present invention verifies the effectiveness and feasibility of the proposed control method. Considering that traditional one-dimensional reinforcement learning does not take into account batch information and cannot improve control performance as the number of batches increases, and at the same time, most of the existing two-dimensional reinforcement learning methods are for linear systems and are difficult to handle the strong nonlinearity and non-repetitive interference commonly existing in actual batch processes, a method for solving strong nonlinearity and non-repetitive interference under the two-dimensional reinforcement learning framework is proposed. The batch process with strong nonlinearity and interference is transformed into a composition of a low-order linear model and unmodeled dynamics, and a reinforcement learning method is used to solve it, and a composite controller with signal compensation is designed. This method first expands the state space equation with two-dimensional characteristics, then sets the two-dimensional value function through the performance index, deduces the correlation between the two-dimensional value functions, thereby revealing the corresponding two-dimensional Bellman equation, and solves the Bellman equation to obtain the optimal control gain. When solving the controller gain, the mutual connection between the time direction and the batch direction is considered, and while obtaining a more accurate control strategy, the convergence speed is accelerated and the optimal performance is enhanced.

[0067] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. Two-dimensional model-free min-max control method for batch processes with high-order nonlinearity and non-repetitive disturbances, the specific steps are as follows: Step 1: Establish the state-space equation of the batch process, and expand it into a linear model and unmodeled dynamics; The batch process with unknown system dynamics is represented by a nonlinear state-space equation, and its expression form is as follows: Among them, $\kappa$ represents the time direction, $\tau$ represents the batch direction, and $x(\kappa,\tau)\in\mathbb{R}$ n represents the system state, and $u(\kappa,\tau)\in\mathbb{R}$ m represents the system control input, and $f(x(\kappa,\tau),u(\kappa,\tau),d(\kappa,\tau))\in\mathbb{R}$ n represents the $f$ function of $x(\kappa,\tau)$, $C$ represents the system matrix with appropriate dimensions, $R$ represents the real matrix, and $n$ and $m$ represent the appropriate dimensions of the real matrix $R$; The batch model with high-order nonlinearity and non-repetitive disturbances is expanded by Taylor series into a combination of a low-order linear model and unmodeled dynamics: Among them, state space represents the state of batch τ at time κ, Δ b x(κ,τ) = x(κ,τ) - x(κ,τ - 1) represents the increment of the state in the batch direction, V(κ - 1,τ) represents the value of batch τ at time κ - 1, e(κ,τ) = y r (κ,τ) - y(κ,τ) represents the error between the set value and the current value of batch τ at time κ, represents the state of batch τ - 1 at time κ + 1, V(κ,τ - 1) represents the value of batch τ - 1 at time κ, e(κ + 1,τ - 1) represents the error between the set value and the current value of batch τ - 1 at time κ + 1, ω1(κ,τ) = u(κ,τ) - u(κ,τ - 1) represents the input difference between batch τ and batch τ - 1 at time κ, ω2(κ,τ) = V(κ,τ) - V(κ,τ - 1) represents the unmodeled dynamic difference between batch τ and batch τ - 1 at time κ; I represents a system parameter matrix with a moderate dimension; Find the optimal control strategy K when the output of the system can track the desired value 11 , K 12 , K 21 , K 22 : Step 2: Define the Lyapunov function, that is, the two-dimensional value function; Under the framework of the two-dimensional system, the value function is decomposed into horizontal and vertical components, and using the principle of asymptotic stability of the Lyapunov function, the difference expressions of the two-dimensional value function in the time-batch direction are obtained; The value function is split into an incremental form and accumulated to obtain the following form: Among them, and represent the time direction and batch direction value functions respectively, and are the lengths of time and batch respectively. In a dynamic system with practical constraints, the optimal control problem based on the zero-sum game theory is formulated as a minimax optimization problem. In a two-dimensional state space the expression of the value function can be obtained by solving the partial differential equation, which usually has a quadratic form: where Q1, Q2, and R are respectively and the weight matrices of ω1(κ, τ), and γ is the decay factor; Step 3: Expand the optimal Q function into a quadratic form and take its derivative to obtain the optimal value; To find the optimal control law, by taking the derivatives of the Q function with respect to ω1 and ω2 and setting their derivatives equal to zero at the extreme points, the optimal control input strategy and the unmodeled dynamic input strategy yield the following equations: wherein are all non-negative matrices; Optimal control gain can be derived as follows: Step 4: Iteratively solve the optimal control strategy to make it finally converge to the ideal set value; Combined with the Bellman equation for numerical optimization, in the two-dimensional state space, if the system dynamics satisfy the Kronecker product structure, the matrix decomposition characteristics can be used to efficiently solve the Bellman equation, and finally the optimal control strategy is obtained; Among them is the target strategy; Step 5: Design a zero-sum game control algorithm with unmodeled dynamics compensation to solve the optimal control law; By using production data, the equation with the Kronecker product is iterated to solve the optimal control law.