A reinforcement learning-based iterative learning predictive control method for mixing temperature

By using reinforcement learning and iterative learning predictive control methods, the mixing temperature control was optimized, solving the problem of traditional control methods relying on experience, and improving the quality and service life of rubber products.

CN119758731BActive Publication Date: 2025-11-28NANJING TECH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411930998.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-11-28
Estimated Expiration
2044-12-26

AI Technical Summary

Technical Problem

Traditional internal mixing process control methods are difficult to achieve the desired results under different material and process conditions. The adjustment of process parameters depends on the operator's experience, which leads to unstable mixing quality and affects the quality and service life of rubber products.

Method used

An iterative learning predictive control method based on reinforcement learning is adopted, which combines a two-dimensional iterative learning model and state-space representation. The control strategy is optimized by the Soft Actor-Critic algorithm, and the error is compensated by data-driven capability, thereby improving the system stability and robustness.

Benefits of technology

It achieves precise control of mixing temperature, improves the system's tracking performance under nonlinear and disturbed conditions, and ensures the quality stability and service life of rubber products.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119758731B_ABST
    Figure CN119758731B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on reinforcement learning's mixing temperature iterative learning prediction control method.It contains: S1, based on reaction energy conservation principle establishes mixing process system temperature control model;S2, establish the system temperature control model of state space expression form;S3, the batch characteristic of process is designed iterative learning control law, establishes iterative axis error model;S4, design two-dimensional system iterative learning model prediction control method, determine cost performance index function;S5, solve the standard quadratic programming problem about optimization variable;S6, set the hyperparameter of Soft Actor-Critic algorithm, the policy network of training completion is combined with ILMPC controller output, and act on system;S7, according to the sample data of buffer area updates network parameter.The application optimizes the control robustness of control system under non-repetitive disturbance and unknown dynamic change, effectively improves the stability of mixing process glue discharge temperature.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the rubber tire processing mixing process, and in particular but not limited to a mixing temperature iterative learning predictive control method based on reinforcement learning. BACKGROUND

[0002] Mixing is a necessary link in rubber processing, and its quality directly determines the physical properties of subsequent processing and products. Poor mixing not only affects the normal progress of processes such as vulcanization, but also reduces the quality, performance and service life of rubber products. The rubber mixing process is a typical rapid intermittent production process, and the traditional control means is difficult to achieve the expected effect under different materials and process conditions. New. Due to the diversification of parameters affecting the mixing process, process control and adjustment of process parameters to optimize the mixing effect become the key direction of intelligent control of the mixing machine.

[0003] In actual production, the process parameters affecting the quality of rubber mixing mainly include main machine power, top bolt pressure, rotor speed, etc. By selecting the formula number of the rubber compound, all parameters of the same batch are configured by the PLC device. Different formulations and different vehicle times need to be adjusted in a timely manner. The discharge time is controlled by observing the high and low temperature of the discharge, and the process control still depends on the experience of the operator for manual adjustment, which makes the traditional control means unable to achieve the expected control effect. The mixing process has high complexity and nonlinearity, and the change of the viscoelastic properties of the rubber compound mainly relies on the severe mechanical friction in the mixing chamber, and the plasticity of the rubber compound increases during the heating process. Therefore, temperature is the most critical factor affecting the quality indicators of the discharge, and an accurate control scheme design can make the operation more stable and accurately control the key process parameters to ensure that they remain within a reasonable range. Combining intelligent control algorithms with actual mixing process mechanisms and production needs becomes an effective technical means for product quality control and precision. SUMMARY

[0004] In view of the problems of rapid, lagging, batch difference and the like in the mixing process, the actual value of the process parameter cannot accurately track the given target, based on the two-dimensional iterative learning model predictive control and reinforcement learning, a dynamic performance index containing two-dimensional error information is introduced, the open-loop drawbacks of iterative learning in the time domain are solved, the parameter mismatch of the model and the non-repeating external disturbance are effectively suppressed by the reinforcement learning algorithm, the adverse effects caused by the error are compensated by the data driving capability, and the stability and robustness of the system are improved. The discharge temperature of each rubber mixing batch can track the preset trajectory.

[0005] To achieve the technical purpose of the present application, the technical scheme adopted is:

[0006] A mixing temperature iterative learning predictive control method based on reinforcement learning, the method comprising the following steps:

[0007] S1, a temperature control model of the mixing process system is established based on the reaction energy conservation principle;

[0008] S2, a system temperature control model in the form of state space is established;

[0009] S3, an iterative learning control law is designed for the batch characteristics of the process, and an iterative axis error model is established;

[0010] S4, a two-dimensional system iterative learning model predictive control method is designed, and a cost performance index function is determined;

[0011] S5, a standard quadratic programming problem about optimization variables is solved;

[0012] S6, the hyperparameters of the Soft Actor-Critic algorithm are set, the trained policy network is combined with the ILMPC controller output, and the combination is applied to the system

[0013] S7, the network parameters are updated according to the sample data in the cache area.

[0014] To optimize the above technical solutions, the following specific measures are taken:

[0015] S1 establishes a temperature control model with discharge temperature as the controlled object according to the reaction thermodynamics in the mixing chamber:

[0016]

[0017] Where n is the rotor speed; K1 is the equivalent heat transfer coefficient; Q c is the cooling water flow rate; K2 is the flow rate heat resistance coefficient; T1 is the temperature of the rubber in the mixing chamber, the equivalent heat capacity C3, T2 represents the temperature of the mixing chamber wall, the equivalent heat capacity C4; T3 represents the temperature of the cooling water in contact with the mixing chamber wall, the equivalent heat capacity C5; R6 represents the degree of heat transfer resistance between the medium in the mixing chamber and the mixing chamber wall, R7 represents the degree of heat transfer resistance between the mixing chamber wall and the cooling water in contact, R8 represents the degree of heat transfer resistance between the mixing chamber wall and the air, and R9 represents the degree of heat transfer resistance between the cooling water in contact with the furnace wall and the cooling water not in contact. The values of the above parameters are determined according to the material characteristics and equipment data in the actual production process.

[0018] The state space model in S2 is derived, taking state variables x = [T1 T2 T3] T , control variables u = [n Q c ] T , and linearizing and discretizing the state space at sampling time T s , (x d , u d ) point:

[0019]

[0020] where k is the iteration number; are the system state, input and output, respectively; f(·) is the uncertainty non-repetitive feature with batch characteristics v k (t). The coefficient matrix is expressed as:

[0021]

[0022] Further derivation obtains a two-dimensional incremental generalized error model containing the iteration domain and time domain dynamic characteristics:

[0023]

[0024] where

[0025]

[0026] δu k (t) = u k (t) - u k (t - 1), w k (t) = f(x k (t), u k (t), v k (t)) - f(x k-1 (t - 1), u k-1 (t - 1), v k-1 ).

[0027] The disturbance w k (t) is simulated as the difference between two consecutive evaluations of the function f(·), representing the nonlinear system dynamics affected by the state x k , the control input u k and the external parameters v k . This representation captures the inherent non-repetitive, time-varying characteristics of batch processes. The error model on the batch axis established by S3 above is used to update the iterative learning control law:

[0028]

[0029] where:

[0030] Combining the error information on adjacent batch axes and model predictive control, the output prediction is obtained:

[0031]

[0032] where:

[0033]

[0034] δu k (t)=u k (t)-u k (t-1),w k (t)=f(x k (t),u k (t),ν k (t))-f(x k-1 (t),u k-1 (t),ν k-1 ),

[0035]

[0036] The error model of the predicted output and the real output of the system is:

[0037] For repetitive disturbance:

[0038] The predicted value of the error sequence of the kth iteration at time t is:

[0039]

[0040] Determine the S4 cost performance index function:

[0041]

[0042] Where t represents the current time, R x and R u are weight matrices of appropriate dimensions, which are used to explain the influence of state and input action on the cost function, and the optimal condition Solve the control increment:

[0043]

[0044] In each time step, only the first control action in the calculated optimal sequence is used, so the control input signal at time t of the kth batch is:

[0045]

[0046] Further design the parameters of the value network and the policy network in the reinforcement learning Soft Actor-Critic (SAC) algorithm in S6, taking the temperature control model in S2 as a virtual interactive environment.

[0047] The core of solving problems by reinforcement learning is to find the optimal policy through the interaction between the agent and the environment. Assuming that in a discrete time process, the state s t is input to the agent, and the resulting policy π(a t |st ) the action a is selected again to the environment state t , the current state is transferred to the next new state s t+1 At the same time, the agent will obtain a reward value r(t) to learn a new strategy.

[0048] The state-action value function is used to measure the expected cumulative reward from state s t Start and take action a t , its expression is:

[0049]

[0050] Where γ represents the discount factor, α is the entropy factor, The expected value of state-action under the policy π(a t |s t ) is denoted by ρ π Represents the probability of the trajectory appearing state s t And action a t Under the current policy.

[0051] Further define the objective function about the optimal policy:

[0052]

[0053] Where, Is the information entropy function of the policy π(a t |s t ) under the current state. By adjusting the entropy parameter, the randomness of the policy can be changed, so as to change the size of the policy exploration ability.

[0054] Introduce the parameterized policy network, and use the least square error minimization method to train the target network:

[0055]

[0056] Where, Is the distribution of previously sampled states and actions, or the replay buffer; θ' represents the target network parameters, Indicates the state value function:

[0057]

[0058] Update the Q function by minimizing the Bellman residual, and add an entropy term to ensure exploration:

[0059]

[0060] Where, Q soft (s t ,a t() indicates the Q value to be updated. s t+1 and a t+1 The Q-value is calculated using the target network θ during the training process. i 'and the original network θ i By minimizing the policy network J π (φ) Training the policy network:

[0061]

[0062] The objective function J π (φ) Estimated gradient with respect to parameter φ. Based on the environmental state-space information in the SAC algorithm. The output of the policy network serves as the agent's action, represented by a. k (t)=π(s k (t)|φ) represents the expression, where a k (t) and the predictive controller Δδu of the iterative learning model k (t) is combined and applied to the controlled process, while obtaining a new feedback state of the process and an immediate reward:

[0063] In iterative predictive control, the reward function aims to minimize the tracking error, thereby ensuring that the system output closely follows the desired reference trajectory. The reward at each time step can be defined as...

[0064] r(s k (t),a k (t))=-l*‖e k (t)‖ 2

[0065] Among them, e k =y k (t)-y k-1 (t) represents the tracking error of the k-th batch at time t, indicating the difference between the system output and the reference output, where l is a positive constant. This reward function penalizes deviations from the reference trajectory, encouraging the strategy to minimize tracking error. The parameters of the target network described in S5 are updated and optimized using the exponential moving average method.

[0066]

[0067] Where τ is the soft update rate, used to stabilize the training process of the target network.

[0068] The effect of the application mainly reflects that the two-dimensional structure considering time and batch is combined with the iterative learning control method and the generalized predictive control method, the rapidity and robustness of the reaction process of the batch bioreactor are comprehensively optimized, the control performance of the batch bioreactor can be improved by adjusting the weight coefficient, and the multi-batch rapid tracking control target under multiple disturbances can be realized.

[0069] The effect of the application mainly reflects that the iterative learning predictive control and the reinforcement learning method are combined, the stability and robustness of the temperature control of the internal mixer are optimized, the adaptability of the system under non-repetitive disturbance and unknown dynamic change is significantly improved, and the tracking control performance of the glue discharging temperature of the internal mixing process is effectively improved. BRIEF DESCRIPTION OF DRAWINGS

[0070] The accompanying drawings are used to provide a further understanding of the application, together with the description, to explain the embodiments of the application, and do not constitute a limitation on the application. In the drawings:

[0071] Figure 1 It is the basic framework diagram of the temperature control of the internal mixing process of the application.

[0072] Figure 2 It is the principle diagram of the combination of the iterative learning model predictive control and the reinforcement learning of the application.

[0073] Figure 3 It is a comparison diagram of the control effect of the application and the traditional iterative learning predictive control under the action of repetitive disturbance.

[0074] Figure 4 It is a comparison diagram of the control effect of the application and the traditional iterative learning predictive control under the action of non-repetitive disturbance.

[0075] Figure 5 It is a learning curve diagram of the intelligent agent under the action of non-repetitive disturbance.

[0076] Figure 6 It is a comparison diagram of the tracking performance of the application and the traditional iterative learning under the action of non-repetitive disturbance.

[0077] Figure 7 It is a whole process schematic diagram of the application. DETAILED DESCRIPTION

[0078] The application will be described in detail below in combination with the drawings and specific embodiments. The embodiments are implemented on the premise of the technical scheme of the application, and detailed implementation modes and specific operation processes are given, but the protection scope of the application is not limited to the following embodiments.

[0079] In combination Figure 7 A kind of internal mixing temperature iterative learning predictive control method based on reinforcement learning, specifically includes:

[0080] S1, such as Figure 1 The present invention illustrates the basic framework for temperature control in the internal mixing process. The main component of the upper top plug, the pressure block 1, is positioned above the feeding hopper 2. Cooling water pipes 3 are located inside the internal mixer chamber wall. A motor 5 drives a pair of rotors 4 with a specific speed ratio to complete the main mixing process. After completion, the rubber compound is flipped out of the internal mixer chamber via the lower top plug 6 and discharged through the discharge door 7. The rotor speed is selected as the manipulated variable, and the temperature of the rubber compound at the discharge port is the controlled variable. For the subsequent implementation of the iterative learning predictive control algorithm, the differential equation needs to be transformed into a state-space representation, with the state variable x = [T1 T2 T3]. T Control variable u = [n Q c ] T At sampling time T s , (x d ,u d The linearized and discretized state-space representation of a point is as follows:

[0081]

[0082] Where k is the number of iterations; These are the system's state variables, input variables, and output variables, respectively; f(·) is a batch-characteristic v k The uncertainty and non-repetition characteristics of (t). The coefficient matrix is ​​expressed as:

[0083]

[0084] Based on the actual reaction process principle and the properties of the rubber compound, Table 1 gives the parameter values ​​in the above coefficient matrix:

[0085] Table 1. Process Mechanism Model Parameters

[0086]

[0087] S2. Using the above dynamic model and its parameter values, construct a prediction model and solve for the two-dimensional performance index function:

[0088]

[0089] Prediction time domain n x =10, control time domain n u =10, weighting coefficient R x =I,R u =0.1I. The control sequence is calculated using the principle of optimization:

[0090]

[0091] S3, such as Figure 2The output of the iterative learning model predictive controller is combined with the action of the reinforcement learning agent, and the training hyperparameters of the agent are set as shown in Table 2:

[0092] Table 2 Hyperparameters of Soft Actor-critic algorithm

[0093]

[0094] Based on the design of the two-dimensional iterative learning model predictive control, two different types of disturbances are added to the system: 1) From the fourth batch, the cooling water temperature is increased by 1℃; 2) From the fourth batch, a random load noise with amplitude in [-1, 1] is added, and the dynamic response performance index of the introduced learning system is analyzed to evaluate the effectiveness of the design scheme. Figure 3 and Figure 4 The results show that the iterative learning model predictive control method based on reinforcement learning introduced in this paper has faster and more stable tracking performance compared to traditional iterative learning model predictive control as the batch increases, and has superior robustness under non-repetitive and unknown noise disturbance. As shown in Figure 5 The tracking error decreases with the increase of batch, and the root mean square error and absolute error can converge to 0 at 20 batches, highlighting the superior tracking performance. Figure 6 The learning curve of this method under non-repetitive disturbance is shown, and the strategy proposed in this paper continuously explores the virtual environment, optimizes the control performance, and realizes the maximum cumulative reward within about 20 training rounds, so that the reward value quickly approaches zero, and the speed and accuracy are high.

Claims

1. A reinforcement learning based iterative learning predictive control method for mixing temperature, characterized in that, The method acts on a rubber mixing system and specifically comprises the following steps: S1, establishing a mixing process system temperature control model based on the reaction energy conservation principle; S2, establishing a system temperature control model in the form of a state space expression; S3, designing an iterative learning control law for the batch characteristics of the process, and establishing an iterative axis error model; S4, designing a two-dimensional system iterative learning model predictive control method, and determining a cost performance index function; The S4 is based on two-dimensional system information in the time domain and the iterative domain, realizes the complementarity of the iterative domain performance and the time domain performance, and guarantees the two-dimensional stability of the batch process control; First define the temperature prediction output on N samples The iterative prediction controller input Δδu and external disturbance w k The sequence is as follows: w k (t) = f(x k (t),u k (t),v k (t))-f(x k-1 (t-1),u k-1 (t-1),v k-1 ), Combined with the iterative axis error model and the system temperature control model, a prediction model is derived: wherein denotes the incremental form of each variable, is the initial state change along the batch axis, G, F are the input coefficient matrix and initial state coefficient matrix of the prediction model, respectively, and the expression is: Based on the prediction model, a quadratic cost performance index is designed, which is composed of the error between the output reference trajectory and the predicted output, the batch, and the time axis control input value, and is used to measure the amplitude of the control input, the convergence robustness of the batch domain, and the dynamic stability of the time domain, respectively; In each item, a weight matrix is added, and by adjusting the size of the weight coefficient, the control performance function most suitable for the mixing process is determined: wherein represents the system state x k (t) about the control input variation Δδu k (t) the objective function, t denotes the current time, is the error sequence for the prediction time t+1, R x and R u are weight matrices of appropriate dimension, to quantify the relative impact of state bias and control increment on the cost function; n x is the maximum time-domain step size for the predicted state, n u is the maximum time domain step size of the prediction control input; S5, solving a standard quadratic programming problem about optimization variables; The S5 utilizes optimality conditions Solving control increments: In each time step, only the first control action in the calculated optimal sequence is used, and the control input signal of the kth batch at time t is: S6, setting the hyperparameters of the Soft Actor-Critic algorithm, combining the trained policy network with the ILMPC controller output, and acting on the rubber mixing system; S7, updating the network parameters according to the sample data in the cache area.

2. The method of claim 1, wherein, The S1 establishes a system temperature control model with the discharge temperature as the controlled object based on the reaction energy conservation principle occurring in the mixing chamber: where n is the rotor speed; K1 is the equivalent heat transfer coefficient; Q c is the cooling water flow rate; K2 is the flow rate thermal resistance coefficient; T1 is the temperature of the compound in the mixing chamber, equivalent heat capacity C3, T2 represents the temperature of the mixing chamber wall, equivalent heat capacity C4; T3 represents the temperature of the cooling water in contact with the mixing chamber wall, equivalent heat capacity C5; R6 represents the degree of obstruction of heat transfer between the medium in the mixing chamber and the mixing chamber wall, R7 represents the degree of obstruction of heat transfer between the mixing chamber wall and the cooling water in contact, R8 represents the degree of obstruction of heat transfer between the mixing chamber wall and the air, and R9 represents the degree of obstruction of heat transfer between the cooling water in contact with the furnace wall and the cooling water not in contact, and the values of the above parameters are determined according to the material characteristics and equipment data in the actual production process.

3. The method of claim 2, wherein, The S2 is based on the linearization and the discretization processing of differential equation, obtains following state space model, selects the rubber temperature, the room wall temperature, the cooling water temperature as state variable x = [T1T2 T3] T , selects the rotor rotating speed and the cooling water flow as control variable u = [n Q c ] T , obtains state expression as: where k is the iteration number; are the state, input and output quantities of the system, respectively; f(·) is the non-repetitive feature with batch characteristics v k (t) of the uncertainty; T s is the sampling time; A, B, C are matrices of respective dimensions, expressed as follows: where T s is the sampling time, V2 is the inlet temperature of the cooling water, x d , u d are the reaction equilibrium points, obtained from the data acquisition of the internal mixer production process equipment.

4. The method of claim 3, wherein, The S3 considers the non-repetitive characteristics of the batch process and establishes an iterative axis error model based on the difference between adjacent batches for different formula specifications and different mixing stages of the mixing process: The S6 interacts with the system temperature control model established in S1 by designing the parameters of the value network and the policy network in the Soft Actor-Critic algorithm, and iteratively updates the network parameters; k = 1,2,3... represents the number of iteration batches, t = 0,1,2,3... represents the sampling time of each batch; x k (t) represents the n x dimensional state vector of the kth batch at t time, u k (t) represents the n u dimensional control vector of the kth batch at t time, y k (t) represents the n y dimensional output vector of the kth batch at t time, A, B, C are coefficient matrices of corresponding dimensions; Δ represents the adjacent iteration axis error operator, and δ represents the adjacent time axis error operator.

5. The method of claim 1, wherein, In order to maintain the stability of the network during training, the S7 adopts an exponential moving average EMA update strategy of target network parameters for optimization; In this method, the parameters of the target network are not directly replaced by the parameters of the current network, but gradually approach the parameters of the current network to achieve more smooth and stable updates; This optimization process controls the update amplitude by introducing a decay factor τ, thereby reducing the volatility in the training process: In the complex system of the mixing process, the key to solving the core problem of system temperature control by reinforcement learning as a key technical means lies in determining the optimal temperature control strategy through the close interaction between the agent and the mixing environment; Given that the temperature control model of the mixing process system has been constructed as a discrete state space form, the actual output value, set value, control signal input, etc. of the discharge temperature can be selected to form the environment state s t ; These environment states closely related to the mixing conditions are input to the agent, and the agent generates the corresponding strategy π(a t |s t ) according to the pre-designed Soft Actor-Critic network parameters, in order to define the objective function: where H (π (·|s t )) is the information entropy function of the current state policy π (a t |s t ), r (s t ,a t ) is the reward value of the current action of the environment, represents the expected value of the state-action under the policy π (a t |s t ), and α is the entropy factor; by adjusting the entropy parameter, the size of the policy exploration ability can be changed, so as to transfer the current glue-out temperature state to the next new state s t+1 ; the training process minimizes the policy network J π (φ) to train the policy network: π φ (a t ∣s t ) represents the policy when the network parameter is φ, represent the state value when the network parameter is θ i , i = 1, 2; in this process, the agent also obtains a reward value r(t) closely related to the current working environment, and continuously learns and updates the network parameters based on this to continuously optimize the policy; in order to evaluate the current environment s t The reward value r(t) of the action a t is designed as a state-action value function: where γ denotes the discount factor, and a is the entropy factor, represents the expected value of state-action under policy t |s t ) under policy π represents the probability of the trajectory to reach state s t and action a t under the current policy; the Q-function is updated by minimizing the Bellman residual while adding an entropy term to ensure exploration: Among them, Q soft (s t ,a t ) represents the Q value to be updated, E a+1~π This indicates that in strategy π(a) t |s t Next action a t+1 Expected value s t+1 and a t+1 The Q-value is set below; the target Q-network is trained using the mean square error minimization method. where D is a distribution of previously sampled states and actions, or a replay buffer; θ' represents the target network parameters, V represents the state value function; According to the environmental state space information in the SAC algorithm The output of the policy network is used as the action of the agent, and the action is denoted as a k (t) = π(s k (t) | φ) is represented, and then a k (t) is organically combined with a model predictive controller Δδu k (t) that is highly optimized after multiple iterations of learning, and precisely acts on the controlled object of the mixing process; at the same time, the system rapidly feeds back new state information and immediate rewards, thereby forming a closed-loop, continuously optimized temperature control loop; in each control period, the agent can continuously adjust the parameters of the policy network and the value network according to the latest feedback information, so as to adapt to various working condition changes that may occur in the mixing process, such as material batch differences, equipment performance fluctuations, and environmental temperature changes, and always ensure that the temperature of the mixing process system is within the ideal control range, thereby providing solid and reliable technical support for the high-quality production of the mixing product.

6. The method of claim 1, wherein, ​ where θ Q represents the parameters of the Q network, θ′ Q are the parameters of the target Q network; φ π represents the parameters of the policy network, φ π ' are the parameters of the target policy network; the update coefficient has a value ranging from 0 to 1 to ensure that the target network parameters gradually approach the current network parameters without causing a mutation.

Citation Information

Patent Citations

  • Chemical batch process fuzzy iterative learning control method

    CN108829058A

  • Robust iterative learning model predictive control method applied to intermittent stirred tank reactor

    CN110045611A