A Predictive Control Method for Intelligent Models in the Reinforcement Learning-Based Mixing Process

By employing a reinforcement learning-based intelligent model predictive control method, the problem of adjusting process parameters during rubber mixing using traditional control methods was solved, thereby improving the quality and service life of rubber products. Through state-space representation and reinforcement learning compensation control, the adjustment of process parameters was optimized, improving the stability and control accuracy of the system.

CN120406110BActive Publication Date: 2026-05-26NANJING TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NANJING TECH UNIV
Filing Date
2025-03-12
Publication Date
2026-05-26

Smart Images

  • Figure CN120406110B_ABST
    Figure CN120406110B_ABST
Patent Text Reader

Abstract

This invention discloses an intelligent model predictive control method for intensive mixing processes based on reinforcement learning. Addressing the degradation of control performance caused by model mismatch due to modeling errors, abrupt state changes, and component degradation in traditional model predictive control, this invention introduces reinforcement learning into the model predictive control framework. The method designs an intelligent model predictive control approach based on reinforcement learning to solve a standard quadratic programming problem with respect to optimization variables. Reinforcement learning selects compensation terms to compensate for model deviations caused by model mismatch, combining this with the controller output and applying it to the system. Finally, the reward is calculated and updated to optimize the control compensation selection for future time steps. As the control cycle progresses, new information is continuously integrated, model predictions are updated, and control inputs are optimized to adapt to potential changes in system behavior or external disturbances. This improves system stability and control accuracy, ensuring that the intensive mixing process can track a preset trajectory.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the mixing process in rubber tire manufacturing, and more specifically, but not limited to, a method for predictive control of the mixing process using an intelligent model based on reinforcement learning. Background Technology

[0002] Internal mixing is a necessary step in rubber processing, and its quality directly determines subsequent processing and the physical properties of the finished product. Poor mixing not only affects the normal progress of processes such as vulcanization, but also reduces the quality, performance, and service life of rubber products. Rubber internal mixing is a typical rapid, intermittent production process, and traditional control methods are insufficient to achieve the desired results under different material and process conditions. Due to the diverse parameters affecting the mixing process, process control and adjustment of these parameters to optimize the mixing effect has become a key direction for the intelligent control of internal mixers.

[0003] In actual production, the main process parameters affecting rubber compounding quality include main engine power, top bolt pressure, and rotor speed. While PLC equipment configures all parameters for the same batch by selecting the rubber compound formula number, different formulas and different batches require timely adjustments to process parameters. Controlling the discharge time solely by observing the discharge temperature still relies on manual adjustments based on operator experience, rendering traditional control methods ineffective. The internal mixing process is highly complex and nonlinear; changes in the viscoelastic properties of the rubber compound primarily depend on the intense mechanical friction within the mixing chamber, and the plasticity of the rubber compound increases during heating. Therefore, temperature and pressure are key factors affecting discharge quality indicators. Precise control scheme design enables smoother operation and more accurate control of key process parameters, ensuring they remain within reasonable ranges. Combining intelligent control algorithms with the actual internal mixing process mechanism and production requirements becomes an effective technical means for lean product quality control. Summary of the Invention

[0004] To address issues such as modeling errors, sudden state changes, component degradation, and external disturbances in the prediction model of the mixing process, which prevent the actual values ​​of process parameters from accurately tracking the given target, an intelligent model predictive control method based on reinforcement learning is proposed. This method introduces reinforcement learning on the basis of model predictive control to correct model deviations caused by model mismatch and reconstructs the model prediction algorithm and function, thereby improving the stability and control accuracy of the system and ensuring that the mixing process can track the preset trajectory.

[0005] To achieve the technical objective of this invention, the technical solution adopted is as follows:

[0006] A predictive control method for an intelligent model of a refining process based on reinforcement learning, the method comprising the following steps:

[0007] S1. Establish a state-space representation of the refining process control model;

[0008] S2. Design a model predictive control method based on reinforcement learning and determine the objective function;

[0009] S3. Solve the standard quadratic programming problem with respect to the optimization variables;

[0010] S4. Reinforcement learning selects compensation terms, which are combined with the controller output and applied to the rubber mixing system;

[0011] S5. Calculate and update the reward value, and optimize the control compensation selection for future time steps.

[0012] To optimize the above technical solution, the specific measures also include:

[0013] Based on the linearization and discretization of differential equations, S1 yields the following state-space model, where the rubber compound temperature, the pressure of the top plug, and the cooling water temperature are selected as state variables x = [T1 P T2]. T Rotor speed, electro-hydraulic proportional directional valve opening, and cooling water flow rate are selected as control variables u = [nx] u Q c ] T The expression for obtaining the state is:

[0014]

[0015] Where k is the number of training iterations; These represent the system's state variables, input variables, and output variables, respectively; A, B, and C are matrices of corresponding dimensions, and their specific expressions are as follows:

[0016]

[0017] Where T s Sampling time; K1 is the equivalent heat transfer coefficient; K2 is the thermal resistance coefficient of the flow rate; C3 is the equivalent heat capacity of the rubber compound in the mixing chamber; C4 is the equivalent heat capacity of the mixing chamber wall; C5 is the equivalent heat capacity of the cooling water in contact with the mixing chamber wall; R6 represents the degree of heat transfer obstruction between the mixing medium and the mixing chamber wall; R7 represents the degree of heat transfer obstruction between the mixing chamber wall and the contacting cooling water; R8 represents the degree of heat transfer obstruction between the mixing chamber wall and the air; V1 is the initial pressure of the top plug; V2 is the inlet temperature of the cooling water; x d u d These are the reaction equilibrium points, obtained from the equipment operation data collected during the internal mixing production process.

[0018] The above S2, for the process control model of the internal mixing process, considers the reinforcement learning-based compensation control δ(k), thus improving the prediction model as follows:

[0019]

[0020] Where: k = 1, 2, 3... represents the number of training iterations, x(k) represents the state vector at time k, u(k) represents the control vector at time k, y(k) represents the output vector at time k, A, B, and C are coefficient matrices of the corresponding dimensions, and δ(k) represents the control compensation term at the current time.

[0021] First, define the prediction output across N prediction intervals. The training sequence of the predictive controller input u(k) is as follows:

[0022]

[0023] By combining compensation control and system process control models, the prediction model is derived:

[0024]

[0025] Where M x C u Let G be the initial state coefficient matrix and the input coefficient matrix of the prediction model, respectively, and let G be the coefficient matrix of the compensation term. The expression is:

[0026]

[0027] A quadratic performance index is designed based on the prediction model. This index consists of the error between the output reference trajectory and the predicted output, and the control input value at the time step, which are used to measure the impact of the output error and the amplitude of the control input, respectively. A weight matrix is ​​added to each term, and the objective function most suitable for the intensive mixing process is determined by adjusting the weight coefficients.

[0028] J = e k (k+1|k) T Qe k (k+1|k)+u k (k+1|k) T Ru k (k+1|k)

[0029] The final expression is:

[0030]

[0031] in, Indicates system state x k Regarding the change in control input u k The objective function is given by k, where k represents the current time. H is the weighting matrix for controlling the input, representing the penalty applied to the input; U T Ex k U represents the effect of state error on the cost function, where E is the interaction term between the state and control inputs;T U extra It is a term related to the difference between the output reference trajectory and the derivation that the system output moves closer to the target output, U. extra It is an additional linear term that includes the output reference value; n x It is the maximum time-domain step size for predicting the state, n u It is the maximum time-domain step size of the predictive control input.

[0032] Using optimality conditions Solving for the variable input control:

[0033]

[0034] Among them, Q bar R bar These are the expanded state weight matrix and the control input weight matrix, Y. ref It is a reference state sequence.

[0035] That is, the control input is:

[0036] U = -H -1 (Ex k +U extra )

[0037] Within each time step, only the first control action from the calculated optimal sequence is used. Therefore, the control input signal at time k is:

[0038]

[0039] Further design the reinforcement learning parameters in S4, using the process control model in S1 as the virtual interactive environment.

[0040] The core of reinforcement learning in problem-solving lies in finding the optimal strategy through the interaction between reinforcement learning and the environment. In the complex system of the mixing process, reinforcement learning, as a key technical means, is crucial for solving system process control problems. The key lies in determining the optimal control strategy through the close interaction between reinforcement learning and the mixing environment. Given that the mixing process system control model has been constructed as a discrete state-space form, the environmental state s can be composed of factors such as the discharge temperature, the actual output value of the top plug pressure, the setpoint, and the control signal input. t These environmental states, closely related to the refining process, are input into reinforcement learning. Based on pre-designed reinforcement learning parameters such as learning rate, exploration rate, and discount factor, a corresponding policy π(a) is generated. t |s t To determine the optimal process control strategy, the objective function is defined as follows:

[0041]

[0042] Where π→a represents selecting action a according to strategy π, Q(s) t ,a t ) is in state s t Take action a t The Q value represents the expected future reward that can be obtained. This indicates selecting the action with the maximum Q value in the current state. By determining the optimal control strategy, the current state is transitioned to the next new state s. t+1 The Q-value function Q(s) t ,a t This can be expressed using the Bellman equation:

[0043]

[0044] in, This represents the expected value of all future rewards after starting action a from state s. During this process, a reward value r(t) closely related to the current working environment is also obtained. Based on this, the network parameters are continuously learned and updated to optimize the strategy; this is used to evaluate the current environment s. t For action a t Given the reward value r(t), the state-action value function is updated as follows:

[0045] Q(s t ,a t )←Q(s t ,a t )+α(r(t)+γ·maxQ(s t+1 ,a t+1 )-Q(s t ,a t ))

[0046] Where γ represents the discount factor, α is the learning rate, and maxQ(s) t+1 ,a t+1 ) indicates that in strategy π(a t+1 |s t+1 The maximum expected value of all possible actions under the given conditions. The Q-value is updated based on the currently selected control action, the current reward, and the maximum Q-value for the next time step. The compensation control term can then be obtained through updating the Q-table.

[0047] δ(k)=Q(s t ,a t )

[0048] Based on the environmental state space information in reinforcement learning algorithms Using strategy selection as an action in reinforcement learning, with a k (t)=π(s k(t) indicates that, subsequently, a k (t) serves as the compensation control δ(k) and the highly optimized model predictive controller u after multiple training iterations. k (t) is combined to obtain a new feedback state and an immediate reward.

[0049] In predictive control, the reward function aims to minimize tracking error, thereby ensuring that the system output closely follows the desired reference trajectory. The reward at each time step can be defined as...

[0050] r(t) = -(yy ref ) T ·diag(w)·(yy ref )

[0051] Where y represents the predicted output, y ref This represents the target reference trajectory. Updated reward values ​​continuously influence the Q-value update to optimize the control strategy, reduce errors from the target state, and thus achieve better control.

[0052] The main advantage of this invention lies in its ability to perform reasonable model predictive control of the mixing process. By adjusting the weighting coefficients, the control performance of the mixing process can be improved, achieving optimal control input under a given state and quickly tracking the control target.

[0053] The main advantages of this invention are that by combining model predictive control and reinforcement learning methods, the system gradually learns control strategies that can optimize system behavior, improves the stability of the mixing process control, significantly enhances the system's adaptability to model mismatch and unknown disturbances, and effectively improves the tracking control performance of the mixing process. Attached Figure Description

[0054] The accompanying drawings are provided to further illustrate the invention and, together with the description, serve to explain embodiments of the invention, but do not constitute a limitation thereof. In the drawings:

[0055] Figure 1 The basic framework for controlling the mixing process of the present invention is shown.

[0056] Figure 2 This illustrates the principle of combining model predictive control with reinforcement learning in this invention.

[0057] Figure 3 A flowchart of the intelligent model predictive control method based on reinforcement learning of the present invention is shown.

[0058] Figure 4 The effect of predictive control based on reinforcement learning intelligent models is demonstrated.

[0059] Figure 5The reward value update for the reinforcement learning-optimized control policy is shown.

[0060] Figure 6 The control inputs for temperature control during the internal mixing process are shown. Detailed Implementation

[0061] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are based on the technical solution of the present invention and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.

[0062] Combination Figure 3 This invention presents an intelligent model predictive control method based on reinforcement learning, the specific process of which is as follows:

[0063] S1, such as Figure 1 As shown, the main component of the upper top bolt, the pressure block 1, is located above the feeding hopper 2. The cooling water pipe 3 is located inside the wall of the internal mixer chamber. A pair of rotors 4 with a speed ratio, driven by a motor 5, complete the main mixing process. After completion, the rubber compound is flipped out of the internal mixer chamber through the lower top bolt 6 and discharged through the discharge door 7. For subsequent training and implementation of the predictive control algorithm, the differential equation needs to be transformed into a state-space expression, taking the state variable x = [T1PT2]. T Control variable u = [nx u Q c ] T At sampling time T s , (x d ,u d The linearized and discretized state-space representation of a point is as follows:

[0064]

[0065] Where k is the number of training iterations; These represent the system's state variables, input variables, and output variables, respectively; A, B, and C are matrices of the corresponding dimensions. The coefficient matrix is ​​expressed as:

[0066]

[0067] Based on the actual reaction process principle and the properties of the rubber compound, Table 1 gives the parameter values ​​in the above coefficient matrix:

[0068] Table 1. Process Mechanism Model Parameters

[0069]

[0070] S2. Using the above dynamic model and its parameter values, construct a prediction model and solve the objective function:

[0071] J = e k(k+1|k) T Qe k (k+1|k)+u k (k+1|k) T Ru k (k+1|k)

[0072]

[0073] Prediction time domain n x =10, control time domain n u =10, weight coefficients Q=I, R=0.1I. The control sequence is calculated using the optimization principle:

[0074]

[0075] S3, such as Figure 2 As shown in Table 2, the output of the model predictor controller is combined with the actions of reinforcement learning, and the training hyperparameters of reinforcement learning are set as follows:

[0076] Table 2 Parameters of Reinforcement Learning Algorithm

[0077]

[0078] To address issues such as modeling errors, abrupt state changes, component degradation, and external disturbances in the predictive model for the internal mixing process, which prevent actual process parameters from accurately tracking the given target, a reinforcement learning-based compensating control is added to the system based on model predictive control design. This corrects model biases and reconstructs the model prediction algorithm and cost function to improve system stability and control accuracy, ensuring the internal mixing process can track the preset trajectory. Considering the actual temperature control scenario in internal mixing, the reasonable rotational speed range is 0-150 r / min, the control input constraint is set to [0,150], and the target temperature is set to 160 degrees Celsius. Figure 4 The results show that the reinforcement learning-based intelligent model predictive control method proposed in this paper has fast and stable tracking performance. Figure 5 As shown, the strategy proposed in this paper continuously explores the virtual environment, optimizes control performance, and achieves the maximum cumulative reward within about 30 training rounds, thereby rapidly approaching zero reward value with high speed and accuracy. Figure 6 This demonstrates that during the internal mixing temperature control process, the control input speed is always kept within a reasonable range. After an initial period of upper limit input, the speed continues to stably track the reference trajectory, ensuring the stable operation of the control system.

Claims

1. A predictive control method for an intelligent model of a refining process based on reinforcement learning, characterized in that, The method is applied to a rubber mixing system and specifically includes the following steps: S1. Establish a state-space representation of the mixing process control model; S1, based on the linearization and discretization of differential equations, yields the following state-space model, with the rubber compound temperature selected accordingly. Top bolt pressure Cooling water temperature State variables Select rotor speed Electro-hydraulic proportional directional valve opening and cooling water flow rate For control variables The expression for obtaining the state is: in It refers to the number of training sessions; ℝ , ℝ , ℝ These are the system's state variables, input variables, and output variables, respectively; ℝ ℝ represents a real vector with the same dimension as the state variable. Represents a real vector with the same dimension as the input, ℝ Represents a real number vector with the same dimension as the output; , , It is a matrix of the corresponding dimension, and the specific expression is as follows: in Sampling time, The equivalent heat transfer coefficient, The thermal resistance coefficient is the flow velocity. The equivalent heat capacity of the rubber compound in the internal mixing chamber. The equivalent heat capacity of the wall of the mixing chamber, The equivalent heat capacity of the cooling water in contact with the wall of the mixing chamber; This represents the degree to which heat transfer between the mixing medium and the walls of the mixing chamber is hindered. This represents the degree of obstruction in heat transfer between the mixing chamber wall and the cooling water in the contact area. This represents the degree to which heat transfer between the walls and air of the mixing chamber is obstructed. The initial pressure of the top bolt. This refers to the inlet temperature of the cooling water. , These are the reaction equilibrium points, obtained from equipment operation data collected during the internal mixing production process. S2. Design a model predictive control method based on reinforcement learning, and determine the objective function; the objective function is: in, For the error sequence at prediction time k+1, Indicates system state Regarding the change in control input The objective function, Indicates the current moment. and It is a weight matrix of appropriate dimensions, used to quantify the impact of state bias and control input on the cost function; n x It is the maximum time-domain step size for predicting the state, n u It is the maximum time-domain step size of the predictive control input; S3. Solve the standard quadratic programming problem with respect to the optimization variables; S3 utilizes the optimality condition. Solve for the control input: in, , These are the expanded state weight matrix and the control input weight matrix, respectively. It is a reference state sequence; within each time step, only the first control action in the calculated optimal sequence is used, then the... The control input signal for the time is: It is an identity matrix of appropriate dimensions, and the subsequent zero matrix is ​​used to ignore subsequent control inputs in the control sequence and only keep the first one; S4. Reinforcement learning selects compensation terms, combines them with the controller output, and applies them to the rubber mixing system; S4 interacts with the system process control model established in S1 by designing parameters in the reinforcement learning algorithm, and performs compensation control on the control system. S5. Calculate and update the reward value to optimize the control compensation selection for future time steps; wherein, the negative norm of the error between the output reference trajectory and the predicted output is used as the reward value.

2. The method according to claim 1, characterized in that, The S2 process control model for the internal mixing process incorporates reinforcement learning-based compensatory control. The improved prediction model is represented as follows: in: Represents the number of training sessions. This represents the state vector at time k. This represents the control vector at time k. This represents the output vector at time k. , , It is the coefficient matrix of the corresponding dimension. This represents the control compensation term at the current moment; First, define the prediction output across N prediction intervals. Training Predictive Controller Input The sequence is as follows: By combining compensation control and system process control models, the prediction model is derived: in , These are the initial state coefficient matrix and the input coefficient matrix of the prediction model, respectively. The coefficient matrix of the compensation term is expressed as follows: Based on the prediction model, a quadratic performance index is designed. This index consists of the error between the output reference trajectory and the predicted output, and the control input value at the time step, which are used to measure the impact of the output error and the amplitude of the control input, respectively. A weight matrix is ​​added to each item, and the objective function most suitable for the intensive mixing process is determined by adjusting the weight coefficients.

3. The method according to claim 1, characterized in that, Given that the process control model of the internal mixing system has been constructed as a discrete state-space form, the actual output value, set value, and control signal input of the discharge temperature and the top plug pressure are selected to constitute the environmental state. These environmental conditions, closely related to the intensive refining process, are input into reinforcement learning. The reinforcement learning then generates corresponding strategies based on pre-designed learning parameters, including the learning rate, exploration rate, and discount factor. To determine the optimal process control strategy, the objective function is defined as follows: in, Indicates according to strategy Select Action , In the state Take action below The Q value represents the expected future reward that can be obtained; This indicates selecting the action with the maximum Q value in the current state; by determining the optimal control strategy, the current state is transitioned to the next new state. ; where the Q-value function Expressed using the Bellman equation: in, Indicates from state Start executing the action Then, the expected value of all future rewards; in the process, a reward value closely related to the current working conditions will also be obtained. Based on this, we continuously learn and update network parameters, and continuously optimize strategies; to assess the current environment Action Reward value The design state-action value function is updated as follows: in, Indicates the discount factor. It's the learning rate. Indicating in strategy The maximum expected value of all possible actions; the Q value is updated based on the currently selected control action, the current reward, and the maximum Q value for the next time step; thus, the compensation control term can be obtained through updating the Q table: Based on the environmental state space information in reinforcement learning algorithms Using strategy selection as an action in reinforcement learning, It is stated that, subsequently, As a compensatory control With a highly optimized model predictive controller after multiple training iterations The system combines these elements and applies them precisely to the controlled object during the refining process. At the same time, the system will quickly provide feedback on new status information and provide immediate rewards, thus forming a closed-loop, continuously optimized process control loop. Within each control cycle, reinforcement learning can continuously adjust its action selection and strategy optimization based on the latest feedback information to adapt to various problems that may arise in the refining process control.

4. The method according to claim 3, characterized in that, The reward function is calculated using the following formula: in, Indicates the predicted output. This represents the target reference trajectory; the updated reward value continuously affects the Q-value update to optimize the control strategy, reduce the error with the target state, and thus achieve better control. The above steps are continuously performed to form a closed-loop control strategy. As the control cycle progresses, new information is continuously integrated, model predictions are updated, and control inputs are optimized to adapt to possible changes in system behavior or external disturbances.