Batch process two-dimensional model-free signal compensation control method with network packet loss and unmodeled dynamics
By introducing a two-dimensional model-free signal compensation method based on the Smith Predictor and reinforcement learning algorithm in the network control system, the problems of network packet loss and unmodeled dynamics are solved, the control performance and stability of the batch process are improved, and the convergence time is shortened.
Patent Information
- Application Number
- CN202510970308.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-09-12
AI Technical Summary
In networked control systems, the unreliability and unmodeled dynamics of network transmission lead to packet loss, which affects the control performance and stability of industrial batch processes. Existing technologies are difficult to effectively address this problem.
A reinforcement learning algorithm integrated with the Smith Predictor is used to establish a two-dimensional model-free signal compensation control method to compensate for network packet loss and unmodeled dynamics, reducing the impact on the control strategy. The reinforcement learning algorithm is used to solve the unmodeled dynamics and compensate for them, reducing dependence on the system model.
The control and tracking performance of the system is improved, the convergence speed is shortened, the effective compensation for unmodeled dynamics is achieved, and the optimal performance of the system is improved.
Smart Images

Figure BDA0005499671860000021 
Figure BDA0005499671860000034 
Figure BDA0005499671860000046
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of industrial process control technology, and specifically relates to a two-dimensional model-free signal compensation control method for batch processes with network packet loss and unmodeled dynamics. By establishing a Smith predictor to compensate for the packet loss problem existing in the network transmission process, a reinforcement learning algorithm is used to solve the unmodeled dynamics and compensate for them, thereby reducing the impact on the control strategy. Background Art
[0002] With the increasing complexity of control tasks and structures, the surge in demand for information sharing and exchange between system components, and the rapid development and widespread application of computer, communication, sensor, and network technologies, a new type of intelligent, distributed, and networked control system has emerged: the Networked Control System (NCS). These systems construct closed-loop control loops over dedicated or general-purpose communication networks, using the network as a feedback channel for information transmission and exchange. However, in real-world industrial applications, network transmission unreliability, bandwidth constraints, and congestion often lead to unavoidable packet loss. Under certain conditions, these phenomena can severely degrade system performance and even cause system instability. It is generally believed that developing advanced control strategies to improve system control performance relies on more accurate mathematical models. However, given the complexity of industrial products, which are rife with high-order, strongly nonlinear, tightly coupled, and uncertain factors that cannot be accurately modeled (collectively referred to as unmodeled dynamics), developing accurate models that meet the desired controller requirements is often challenging.
[0003] Intermittent production utilizes a batch processing approach, with the same equipment completing multiple production steps sequentially over a period of time. This process requires rigorous preparation before starting the task, and a single batch of materials undergoes multiple steps to complete the processing cycle. Its core advantage lies in its adaptability to the production schedule of multiple product categories and batches. Compared to continuous processes, intermittent processes have the following key advantages:
[0004] (1) Periodic batch production. The process of intermittent production is different from the continuous production process. Intermittent production is a periodic production process, and it requires production to be organized according to the production sequence, time period and operating parameters specified in the recipe.
[0005] (2) Material states and operating parameters are dynamic. Dynamic characteristics are the essence of intermittent processes. In intermittent production processes, material composition, temperature, pressure, required heat and cooling, and other operating parameters will change over time.
[0006] (3) Strong flexible production capacity. In the intermittent production process, different process operations can be completed on a fixed device according to different formulas, using different raw materials and operating parameters, which is conducive to the production of small batches of multiple varieties of products. This operating characteristic brings many random and unstable factors to safe production.
[0007] In light of the above, this paper designs a two-dimensional model-free signal compensation control method for batch processes that is suitable for network packet loss and unmodeled dynamics. This method uses a reinforcement learning algorithm integrated with a Smith Predictor to compensate for packet loss and the actual unmodeled dynamics during data transmission. This method does not rely on process models or initial system parameters, and the compensation mechanism simultaneously suppresses the impact of packet loss and unmodeled dynamics on control performance. Summary of the Invention
[0008] This paper invents a two-dimensional model-free signal compensation control method for batch processes with network packet loss and unmodeled dynamics. This method uses a reinforcement learning algorithm integrated with a Smith Predictor to compensate for packet loss and unmodeled dynamics during data transmission. The algorithm continuously learns from data in both time and batch directions. By combining reinforcement learning with signal compensation, the unmodeled dynamics are resolved, enabling a more precise control strategy to be established, resulting in an optimal control strategy. This improves the system's control and tracking performance and accelerates convergence.
[0009] The present invention is achieved through the following technical solutions:
[0010] This study proposes a two-dimensional model-free signal compensation control method for batch processes with network packet loss and unmodeled dynamics. A Smith Predictor is developed to compensate for packet loss in network transmission. Subsequently, a reinforcement learning algorithm is used to solve and compensate for the unmodeled dynamics, minimizing their impact on the control strategy. First, a nonlinear state-space equation for the batch process with network packet loss and unmodeled dynamics is established. Second, the tracking error is extended as a state variable to the performance metric. A packet loss model is constructed in the network environment, and a two-dimensional Smith Predictor with packet loss compensation is introduced to compensate for packet loss in network transmission. Subsequently, the Bellman equation is constructed based on the relationship between the value function and the Q function. An optimal controller is obtained by solving the Bellman equation involving the Kronecker product. Through adaptive compensation, the system effectively mitigates the impact of the unmodeled dynamics. This study can effectively solve the problem of the difficulty in applying two-dimensional reinforcement learning in nonlinear systems, greatly reducing the excessive dependence on the system model. At the same time, the two-dimensional Smith predictor is introduced to compensate for the packet loss of production data during network transmission. By solving the unmodeled dynamics through the signal compensation method, the optimal control strategy can be obtained quickly and accurately, thereby improving the optimal performance of the system and accelerating the convergence speed.
[0011] Step 1: Describe the state space equations of the batch process under packet loss, and expand them into a linear model and unmodeled dynamics;
[0012] The batch process with unknown system dynamics is represented by nonlinear state-space equations, which are expressed as follows:
[0013]
[0014] Among them, κ represents the time direction, τ represents the batch direction, Indicates the system status, represents the system control input, System output, ξ(x κ,τ ,u κ,τ ,d κ,τ ) is represented by x κ,τ ξ function, C represents the system matrix with appropriate dimensions, R represents the real matrix, N u , N y and N V is represented as a real matrix R of appropriate dimensions;
[0015] The batch model with network packet loss is Taylor expanded into an incremental combination of a low-order linear model and unmodeled dynamics, resulting in the expanded state space equation:
[0016]
[0017] in, State Space represents the state of batch τ at time κ, V(κ-1,τ) represents the value of batch τ at time κ-1, represents the state of the batch τ-1 at time κ+1, V(κ,τ-1) represents the value of the batch τ-1 at time κ, and y r (κ+1,τ) represents the output value set by the system, and y(κ,τ-1) represents the actual output value of the system. a system matrix representing the appropriate dimensions;
[0018] Step 2: Construct a packet loss model in a network environment and introduce a two-dimensional Smith Predictor with packet loss compensation.
[0019] In a wireless communication environment, considering the impact of data packet loss on control signal transmission during data transmission, the dynamic characteristics of the system including data loss after wireless link transmission can be expressed as:
[0020]
[0021] in, Indicates the process control status obtained after transmission via wireless network. Indicates whether the transmission is successful, and the value can be 0 or 1. , it indicates that data packets are lost during transmission. When , it means the transmission is successful and no data packet loss occurs during the transmission process;
[0022] The system status expression of the receiver containing network loss after transmission through the wireless network is as follows:
[0023]
[0024] in, is the number of consecutive packet losses during data transmission, and satisfies The value range of is the maximum number of consecutive packet losses, then the following formula can be obtained to predict the state quantity at the current time;
[0025]
[0026] Among them, A and B are system matrices with appropriate dimensions under network packet loss;
[0027] In the case of TCP or UDP protocols, the number of packet losses can be considered to be known, and the incremental state space μ containing packet losses is κ,τ The Smith Predictor can be constructed as follows:
[0028]
[0029] in,
[0030] When the system output can track the expected value, find the optimal control strategy:
[0031]
[0032] Among them, K 11 , K 12 , K 21 , K 22 is the control strategy gain after integrating the Smith predictor;
[0033] Step 3: Expand the optimal Q function into a quadratic form and derive the optimal value;
[0034] In the framework of a two-dimensional system, the value function is decomposed into horizontal and vertical components. Using the asymptotic stability principle of the Lyapunov function, the differential accumulation expression of the two-dimensional value function in the time batch direction is obtained.
[0035]
[0036] in, and Represent the time direction and batch direction value functions respectively, and time and batch length respectively,
[0037] In a dynamic system with practical constraints, the optimal control problem based on zero-sum game theory is expressed as a minimax optimization problem in a two-dimensional state space. By solving the partial differential equation, we can get the expression of the value function, which usually has a quadratic form:
[0038]
[0039] Among them, Q1, Q2 and R are and The weight matrix, γ is the attenuation factor, is the unmodeled dynamic input increment;
[0040] Step 4: Design a control algorithm with signal compensation;
[0041] Combined with the Bellman equation for numerical optimization, in the two-dimensional state space, by making the Q function u κ,τ -u κ,τ-1 With V κ,τ -V κ,τ-1 Derivative, and set its derivative equal to zero at the extreme point. If the system dynamics satisfies the Kronecker product structure, its matrix decomposition characteristics can be used to efficiently solve the Bellman equation, and the optimal control gain can be derived;
[0042]
[0043]
[0044] in, represents the target policy for controlling the input, represents the target policy without modeling dynamic input, G represents the state matrix of the system;
[0045] Step 5: Integrate packet loss compensation into the control algorithm to solve the optimal control law;
[0046] According to the introduced Smith predictor (6), a control strategy based on Smith compensation can be constructed:
[0047]
[0048] in
[0049] Similarly, after the Smith predictor is introduced, equation (13) is substituted into equation (12) to obtain the updated two-dimensional Bellman equation. First, the behavioral strategy is used to act on the system to generate two-dimensional data in the time direction and batch direction, and the data is stored in the system. Then, the initial controller gain that can stabilize the system is given. By iteratively learning using production data, the equation with the Kronecker product is iterated to solve the optimal control law.
[0050]
[0051]
[0052] The advantages and effects of the present invention are:
[0053] This study addresses the model-free control problem of batch processes in the presence of network packet loss and unmodeled dynamics. We innovatively propose a two-dimensional model-free signal compensation control method incorporating a Smith Predictor. Unlike traditional precise modeling methods and one-dimensional control schemes, this method compensates for packet loss in network transmission by establishing a Smith Predictor. Subsequently, a reinforcement learning algorithm is used to solve and compensate for the unmodeled dynamics, minimizing their impact on the control strategy. First, a nonlinear state-space equation for the batch process with network packet loss and unmodeled dynamics is established. Second, the tracking error is extended as a state variable to the performance metric. A packet loss model is then constructed in the network environment. A two-dimensional Smith Predictor with packet loss compensation is introduced to compensate for packet loss in network transmission. Finally, the Bellman equation is constructed based on the relationship between the value function and the Q function. The optimal controller is obtained by solving the Bellman equation involving the Kronecker product. Through adaptive compensation, the system effectively mitigates the impact of the unmodeled dynamics. The proposed method can effectively solve the problem of the difficulty of applying two-dimensional reinforcement learning in nonlinear systems, greatly reducing the excessive dependence on the system model. At the same time, a two-dimensional Smith predictor is introduced to compensate for the packet loss of production data during network transmission. By solving the unmodeled dynamics through the signal compensation method, the optimal control strategy can be obtained quickly and accurately, thereby improving the optimal performance of the system and accelerating the convergence speed. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 The control gains of different batches under the two-dimensional model-free signal compensation method fused with Smith Predictor Convergence process;
[0055] Figure 2 The control gains of different batches under the two-dimensional model-free signal compensation method fused with Smith Predictor Convergence process;
[0056] Figure 3 The control gains of different batches under the two-dimensional model-free signal compensation method fused with Smith Predictor Convergence process;
[0057] Figure 4 The control gains of different batches under the two-dimensional model-free signal compensation method fused with Smith Predictor Convergence process;
[0058] Figure 5 The control output tracking curves of different batches under the two-dimensional model-free signal compensation method integrated with the Smith Predictor;
[0059] Figure 6 These are the control input curves for different batches using the two-dimensional model-free signal compensation method fused with the Smith Predictor.
[0060] Figure 7 These are the unmodeled dynamic input curves of different batches under the two-dimensional model-free signal compensation method integrated with the Smith Predictor. DETAILED DESCRIPTION
[0061] In order to further illustrate the present invention, the present invention is described in detail below with reference to the accompanying drawings and examples, but they should not be construed as limiting the scope of protection of the present invention.
[0062] Example 1:
[0063] Injection molding technology is the core pillar technology of global plastic manufacturing. Its rapid production, high efficiency and flexible product update capabilities make it an indispensable process in the field of batch processing. The process mainly includes four stages: melting, injection, holding pressure and cooling. The injection stage directly determines the quality and appearance of the final product. During the injection process, precise control of injection speed and pressure is particularly critical. These two mutually coupled parameters jointly dominate the filling behavior of the melt in the mold. Specifically, the injection pressure ensures that the mold cavity is completely filled by affecting the density distribution of the plastic melt. In order to prevent defects such as short shots or flash, precise control of the nozzle pressure is required, which requires precise adjustment of the valve opening. This study proposes a two-dimensional model-free signal compensation algorithm that integrates the Smith Predictor, with precise speed tracking in the injection stage as the control target. This technology is expected to improve product quality while optimizing production efficiency.
[0064] Based on a large amount of experimental data, the correlation equation between nozzle pressure (NP) and valve opening (VO) in the holding pressure stage of two-dimensional injection molding is established as follows:
[0065]
[0066] Among them, NP κ,τis the nozzle pressure of the batch at time κ; VO κ,τ is the valve opening of batch τ at time κ; NP κ-1,τ is the nozzle pressure of batch τ at time κ-1; VO κ-1,τ is the valve opening of batch τ at time κ-1; d κ,τ is the unmodeled dynamic value of the batch τ at time κ.
[0067] definition: u κ,τ -u κ,τ-1 =VO(κ,τ) is the control input, Y κ,τ =NP κ,τ Output of the system.
[0068] According to formula (16), the state space model of the pressure holding stage is established:
[0069]
[0070] in After repeated experiments, we determined the parameters for the pressure-holding phase to be Q1 = Q2 = diag[2, 0, 0, 0, 1], R = 0.01, and γ = 1.5. However, in actual batch production, obtaining the initial system parameters is difficult. Therefore, in subsequent research, we proposed a two-dimensional model-free signal compensation algorithm that incorporates a Smith Predictor to design an optimal controller for batch processes.
[0071] from Figure 1-4 As can be seen, the control strategy's convergence trend was suboptimal in the fourth and fifth batches, but it showed significant improvement by the twelfth batch, and after twenty batches, it had fully converged and approached the optimal value. Simulation results of the two-dimensional model-free signal compensation algorithm incorporating the Smith Predictor show that the algorithm exhibits excellent tracking performance, with the tracking error gradually decreasing with each batch.
[0072] Figure 5 The system output curves for the proposed algorithm are shown. As can be seen, the output of the two-dimensional batch system is set at 20 bar. The current output has not yet reached the target value, and there is an error between the current output and the set value. However, this error decreases with the number of batches. Furthermore, the system's tracking performance improves over time, and by the thirteenth batch, the system output fully matches the set value. Figure 6 shows the input curve of the controller signal of the two-dimensional system, Figure 7 Results for an unmodeled dynamic input signal are presented. Both figures show that the input trajectory of the control variable exhibits a stable state. As the number of iterations accumulates, the proposed method gradually optimizes the control gain and eventually approaches the optimal value.
[0073] In summary, this study used a batch process control design with network packet loss and unmodeled dynamics as an example to demonstrate the effectiveness and feasibility of the proposed control method. As an industrial production model, batch processes have attracted considerable attention for their stable and efficient operation. Given the complex dynamics of batch processes and the difficulty in modeling the controlled process, model-based control methods face practical challenges. This study proposed a model-free algorithm that uses reinforcement learning to solve the two-dimensional optimal tracking control problem for a linear batch process with packet loss and disturbances. The algorithm can be used to determine the optimal gains of the designed controller. Compared to model-based methods, this algorithm does not require prior knowledge of the actual process, relying solely on historical data in the time and batch dimensions. This method compensates for packet loss during network transmission by establishing a Smith predictor. Subsequently, a reinforcement learning algorithm is used to solve and compensate for the unmodeled dynamics, minimizing their impact on the control strategy. The interrelationship between the time and batch dimensions is taken into account when solving for the controller gains, resulting in a more accurate control strategy while accelerating convergence and enhancing optimal performance.
[0074] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.
Claims
1. A two-dimensional model-free signal compensation control method for batch processes with network packet loss and unmodeled dynamics. The specific steps are as follows: Step 1: Describe the state space equations of the batch process under packet loss, and expand them into a linear model and unmodeled dynamics; The batch process with unknown system dynamics is represented by nonlinear state-space equations, which are expressed as follows: in, κ represents the time direction, τ represents the batch direction, Indicates the system status, represents the system control input, System output, ξ(x κ,τ ,u κ,τ ,d κ,τ ) is represented by x κ,τ ξ function, C represents the system matrix with appropriate dimensions, R represents the real matrix, N u , N y and N V is represented as a real matrix R of appropriate dimensions; The batch model with network packet loss is Taylor expanded into an incremental combination of a low-order linear model and unmodeled dynamics, resulting in the expanded state space equation: in, State Space represents the state of batch τ at time κ, V(κ-1,τ) represents the value of batch τ at time κ-1, represents the state of the batch τ-1 at time κ+1, V(κ,τ-1) represents the value of the batch τ-1 at time κ, and y r (κ+1,τ) represents the output value set by the system, and y(κ,τ-1) represents the actual output value of the system. a system matrix representing the appropriate dimensions; Step 2: Construct a packet loss model in a network environment and introduce a two-dimensional Smith Predictor with packet loss compensation. In a wireless communication environment, considering the impact of data packet loss on control signal transmission during data transmission, the dynamic characteristics of the system including data loss after wireless link transmission can be expressed as: in, Indicates the process control status obtained after transmission via wireless network. Indicates whether the transmission is successful, and the value can be 0 or 1. , it indicates that data packets are lost during transmission. When , it means the transmission is successful and no data packet loss occurs during the transmission process; The system status expression of the receiver containing network loss after transmission through the wireless network is as follows: in, is the number of consecutive packet losses during data transmission, and satisfies The value range of is the maximum number of consecutive packet losses, then the following formula can be obtained to predict the state quantity at the current time; Among them, A and B are system matrices with appropriate dimensions under network packet loss; In the case of TCP or UDP protocols, the number of packet losses can be considered to be known, and the incremental state space μ containing packet losses is κ,τ The Smith Predictor can be constructed as follows: in, When the system output can track the expected value, find the optimal control strategy: Among them, K 11 , K 12 , K 21 , K 22 is the control strategy gain after integrating the Smith predictor; Step 3: Expand the optimal Q function into a quadratic form and derive the optimal value; In the framework of a two-dimensional system, the value function is decomposed into horizontal and vertical components. Using the asymptotic stability principle of the Lyapunov function, the differential accumulation expression of the two-dimensional value function in the time batch direction is obtained. in, and Represent the time direction and batch direction value functions respectively, and time and batch length respectively, In a dynamic system with practical constraints, the optimal control problem based on zero-sum game theory is expressed as a minimax optimization problem in a two-dimensional state space. By solving the partial differential equation, we can get the expression of the value function, which usually has a quadratic form: Among them, Q1, Q2 and R are and The weight matrix, γ is the attenuation factor, is the unmodeled dynamic input increment; Step 4: Design a control algorithm with signal compensation; Combined with the Bellman equation for numerical optimization, in the two-dimensional state space, by making the Q function u κ,τ -u κ,τ-1 With V κ,τ -V κ,τ-1 Derivative, and set its derivative equal to zero at the extreme point. If the system dynamics satisfies the Kronecker product structure, its matrix decomposition characteristics can be used to efficiently solve the Bellman equation, and the optimal control gain can be derived; in, represents the target policy for controlling the input, represents the target policy without modeling dynamic input, G represents the state matrix of the system; Step 5: Integrate packet loss compensation into the control algorithm to solve the optimal control law; According to the introduced Smith predictor (6), a control strategy based on Smith compensation can be constructed: in Similarly, after the Smith predictor is introduced, equation (13) is substituted into equation (12) to obtain the updated two-dimensional Bellman equation. First, the behavioral strategy is used to act on the system to generate two-dimensional data in the time direction and batch direction, and the data is stored in the system. Then, the initial controller gain that can stabilize the system is given. By iteratively learning using production data, the equation with the Kronecker product is iterated to solve the optimal control law.
Citation Information
Cited By
Multi-rate hierarchical learning control method for dense medium coal preparation process under unreliable communication
CN121704193A