Banbury mixing process intelligent model prediction control method based on reinforcement learning

Through the intelligent model prediction and control method based on reinforcement learning, the problem of inaccurate process parameter adjustment during traditional refining is solved, and the stability and service life of rubber products are improved.

CN120406110AActive Publication Date: 2025-08-01NANJING TECH UNIV

Patent Information

Application Number
CN202510291870.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-08-01
Estimated Expiration
2045-03-12

AI Technical Summary

Technical Problem

Traditional intensive process control methods rely on manual experience, making it difficult to achieve accurate process parameter adjustment under different materials and process conditions, resulting in poor mixing effect and affecting the quality and service life of rubber products.

Method used

The intelligent model prediction and control method based on reinforcement learning is adopted, combined with model prediction and reinforcement learning, and through state space expression and reinforcement learning compensation control, process parameter adjustment is optimized to achieve accurate tracking and stable control of the intensive process.

Benefits of technology

It improves the control accuracy and stability of the intensive refining process, can quickly track control targets, and improves the quality and service life of rubber products.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120406110A_ABST
    Figure CN120406110A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent model predictive control method for an internal mixing process based on reinforcement learning, and aims at solving the problem that the control effect becomes poor due to model mismatch caused by modeling errors, sudden state change, part degradation and the like in traditional model predictive control, and reinforcement learning is introduced on the basis of model predictive control. And designing an intelligent model prediction control method based on reinforcement learning, and solving a standard quadratic programming problem about optimization variables. And performing reinforcement learning to select a compensation item, performing compensation control on model deviation caused by model mismatch, combining with the output of the controller, and acting on the system. And finally, calculating a reward and updating a reward value, and optimizing control compensation selection of a future time step. Along with the proceeding of a control cycle, new information is continuously integrated, model prediction is updated, and control input is optimized so as to adapt to possible system behavior change or external interference. The stability and the control precision of the system are improved, and a preset track can be tracked in the internal mixing process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the kneading process in rubber tire processing, and particularly but not limited to an intelligent model predictive control method for the kneading process based on reinforcement learning. Background Art

[0002] Kneading is an essential process in rubber processing, and its quality directly determines the physical properties of subsequent processing and products. Poor mixing not only affects the normal progress of processes such as vulcanization but also reduces the quality, performance, and service life of rubber products. The rubber kneading process belongs to a typical fast intermittent production process, and traditional control means are difficult to achieve the expected effect under different materials and process conditions. Due to the diversification of parameters affecting the mixing process, the process control and adjustment of process parameters to optimize the mixing effect have become the key direction for the intelligent control of kneading machines.

[0003] In actual production, the process parameters affecting the quality of rubber mixing mainly include the main engine power, upper plug pressure, rotor speed, etc. By selecting the rubber compound formula number, the PLC device configures all parameters of the same batch. For different formulas and different batches, the process parameters need to be adjusted in a timely manner. Only by observing the discharge temperature and controlling the discharge time, the process control still relies on the experience of operators for manual adjustment, which makes the traditional control means unable to achieve the expected control effect. The kneading process has high complexity and nonlinearity. The change in the viscoelastic properties of the rubber compound mainly relies on the intense mechanical friction in the kneading chamber, and the plasticity of the rubber compound increases during the heating process. Therefore, temperature and pressure are the key factors affecting the discharge quality index. The design of an accurate control scheme can make the operation more stable, more precisely control the key process parameters, ensure that they are kept within a reasonable range, and combining intelligent control algorithms with the actual kneading process mechanism and production requirements has become an effective technical means for lean product quality control. Summary of the Invention

[0004] Aiming at problems such as modeling errors, state mutations, component degradation, and external disturbances in the prediction model of the kneading process, where the actual values of process parameters cannot accurately track the given target, an intelligent model predictive control method based on reinforcement learning, on the basis of model predictive control, introduces reinforcement learning to correct the model deviation caused by model mismatch, and reconstructs the model prediction algorithm and function to improve the stability and control accuracy of the system, ensuring that the kneading process can track the preset trajectory.

[0005] To achieve the technical objectives of the present invention, the technical solutions adopted are as follows:

[0006] An intelligent model predictive control method for the kneading process based on reinforcement learning, the method comprising the following steps:

[0007] S1. Establish a kneading process control model in the form of state space expression;

[0008] S2. Design a model predictive control method based on reinforcement learning to determine the objective function;

[0009] S4. Solve the standard quadratic programming problem for the optimization variables;

[0010] S7. The reinforcement learning selects the compensation term, combines it with the controller output, and acts on the rubber internal mixer system;

[0011] S5. Calculate the reward and update the reward value to optimize the control compensation selection for future time steps.

[0012] To optimize the above technical solutions, the specific measures taken also include:

[0013] The above S1 is based on the linearization and discretization of the differential equation to obtain the following state space model. The rubber temperature, upper plug pressure, and cooling water temperature are respectively selected as the state variables x = [T1 P T2] T , and the rotor speed, opening of the electro-hydraulic proportional reversing valve, and cooling water flow rate are selected as the control variables u = [n x u Q c T , and the state expression is obtained as:

[0014]

[0015] where k is the number of training times; are respectively the state quantity, input quantity, and output quantity of the system; A, B, and C are matrices of corresponding dimensions, and the expressions are specifically as follows:

[0016]

[0017] where T s is the sampling time, K1 is the equivalent heat transfer coefficient, K2 is the thermal resistance coefficient of the flow rate; C3 is the equivalent heat capacity of the rubber in the internal mixer, C4 is the equivalent heat capacity of the internal mixer wall, C5 is the equivalent heat capacity of the cooling water in contact with the internal mixer wall; R6 represents the degree of heat transfer obstruction between the medium in the internal mixer and the internal mixer wall, R7 represents the degree of heat transfer obstruction between the internal mixer wall and the contact part of the cooling water, R8 represents the degree of heat transfer obstruction between the internal mixer wall and the air; V1 is the initial pressure of the upper plug, V2 is the inlet temperature of the cooling water, x d , u d are respectively the reaction equilibrium points, which are obtained according to the equipment operation data collection in the internal mixing production process.

[0018] The above S2 considers the compensation control δ(k) of reinforcement learning for the process control model of the internal mixing process, and the improved prediction model is expressed as:

[0019] ​

[0020] Where: k = 1, 2, 3... represents the number of training times, x(k) represents the state vector at the k-th moment, u(k) represents the control vector at the k-th moment, y(k) represents the output vector at the k-th moment, A, B, and C are coefficient matrices of corresponding dimensions, and δ(k) represents the control compensation term at the current moment.

[0021] First, define the predicted output on N prediction intervals. The input sequence u(k) of the training prediction controller is as follows:

[0022]

[0023] Combining the compensation control and the system process control model, the prediction model is derived as:

[0024]

[0025] Where M x , C u are respectively the initial state coefficient matrix and the input coefficient matrix of the prediction model, G is the coefficient matrix of the compensation term, and the expression is:

[0026]

[0027] Based on the prediction model, a quadratic performance index is designed. This index consists of the error between the output reference trajectory and the predicted output, and the control input value at the time step, which are respectively used to measure the influence of the output error and the amplitude of the control input; weight matrices are added to each item, and by adjusting the size of the weight coefficients, the objective function most suitable for the internal mixer process is determined:

[0028] J = e k (k + 1|k) T Qe k (k + 1|k) + u k (k + 1|k) T Ru k (k + 1|k)

[0029] The final expression is:

[0030]

[0031] Among them, represents the objective function of the system state x k with respect to the change in the control input u k , k represents the current moment, is the penalty for the control input, H is the weighted matrix of the control input; U T Ex k represents the influence of the state error on the cost function, and E is the interaction term between the state and the control input; UT U extra It is the difference related to the output reference trajectory, which is used to deduce that the system output is close to the target output. extra is an additional linear term containing the output reference value; n x is the maximum time domain step size of the predicted state, n u is the maximum time domain step size of the predictive control input.

[0032] Using optimality conditions Solve for the control variable input:

[0033]

[0034] Among them, Q bar 、R bar They are the expanded state weight matrix and control input weight matrix, Y ref is the reference state sequence.

[0035] That is, the control input is:

[0036] U=-H -1 (Ex k +U extra )

[0037] In each time step, only the first control action in the calculated optimal sequence is used, and the control input signal at the kth moment is:

[0038]

[0039] The reinforcement learning parameters in S4 are further designed, and the process control model in S1 is used as the virtual interactive environment.

[0040] The core of reinforcement learning to solve problems lies in finding the optimal strategy through the interaction between reinforcement learning and the environment. In the complex system of the mixing process, reinforcement learning is a key technical means. The core of solving the system process control problem lies in determining the optimal control strategy through the close interaction between reinforcement learning and the mixing environment. Given that the control model of the mixing process system has been constructed in the form of a discrete state space, the actual output value of the discharge temperature, the upper bolt pressure, the set value, the control signal input, etc. can be selected to form the environmental state s t After these environmental states closely related to the mixing conditions are input into the reinforcement learning, the corresponding strategy π(a t |s t ), in order to clarify the optimal process control strategy, the objective function is defined as:

[0041]

[0042] where, π→a denotes selecting action a according to policy π, and Q(s t ,a t ) is the Q-value of taking action a t in state s t , representing the expectation of future rewards expected to be obtained. denotes selecting the action with the maximum Q-value in the current state. By determining the optimal control policy, the current state is transferred to the next new state s t+1 . Among them, the Q-value function Q(s t ,a t ) can be represented by the Bellman equation:

[0043]

[0044] where, represents the expectation of all future rewards after executing action a starting from state s. During this process, a reward value r(t) closely related to the current working condition environment is also obtained, and based on this, the network parameters are continuously learned and updated, and the policy is continuously optimized; to evaluate the reward value r(t) of the current environment s t for action a t , the state-action value function is designed to be updated as follows: [[ID=3*]]

[0045] Q(s t ,a t )←Q(s t ,a t )+α(r(t)+γ·maxQ(s t+1 ,a t+1 )-Q(s t ,a t ))

[0046] where, γ represents the discount factor, α is the learning rate, and maxQ(s t+1 ,a t+1 ) represents the maximum expectation of all possible actions under policy π(a t+1 |s t+1 ). The Q-value is updated according to the currently selected control action, as well as the current reward and the maximum Q-value at the next time step. Thus, the compensation control term can be obtained through the update of the Q-table:

[0047] δ(k)=Q(s t ,a t )

[0048] According to the environmental state space information in the reinforcement learning algorithm taking the selection of the policy as the action of reinforcement learning, using a k (t)=π(s k(t)) indicates that subsequently, a k (t) is used as the compensation control δ(k) and combined with the highly optimized model predictive controller u k (t) to obtain new feedback states and immediate rewards simultaneously.

[0049] In training predictive control, the purpose of the reward function is to minimize the tracking error, thereby ensuring that the system output closely follows the desired reference trajectory. The reward at each time step can be defined as

[0050] r(t) = -(y - y ref ) T ·diag(w)·(y - y ref )

[0051] where y represents the predicted output and y ref represents the target reference trajectory. The updated reward value continuously acts on the update of the Q value to optimize the control strategy, reduce the error from the target state, and thus achieve better control.

[0052] The effects of the present invention are mainly manifested in performing reasonable model predictive control on the internal mixer process. By adjusting the weight coefficients, the control performance of the internal mixer process can be improved, the optimal control input under a given state can be achieved, and the control target can be quickly tracked.

[0053] The effects of the present invention are mainly reflected in combining the model predictive control and the reinforcement learning method, gradually learning the control strategy that can optimize the system behavior, optimizing the stability of the internal mixer process control, significantly enhancing the adaptability of the system under model mismatch and unknown disturbances, and effectively improving the tracking control performance of the internal mixer process. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] The drawings are used to provide a further understanding of the present invention and, together with the description, are used to explain the embodiments of the present invention and do not constitute a limitation to the present invention. In the drawings:

[0055] Figure 1 shows the basic framework of the internal mixer process control of the present invention.

[0056] Figure 2 shows the principle of the combination of model predictive control and reinforcement learning of the present invention.

[0057] Figure 3 shows the flowchart of the intelligent model predictive control method based on reinforcement learning of the present invention.

[0058] Figure 4 shows the effect of the intelligent model predictive control based on reinforcement learning.

[0059] Figure 5Shows the update of the reward value of the reinforcement learning optimization control strategy.

[0060] Figure 6 Shows the control input of the temperature control in the internal mixer process. Detailed implementation manners

[0061] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and the detailed implementation manners and specific operation processes are given, but the protection scope of the present invention is not limited to the following embodiments.

[0062] Combined with Figure 3 , the present invention provides an intelligent model predictive control method based on reinforcement learning, and the specific process is as follows:

[0063] S1. As Figure 1 shown, the main component pressing weight 1 of the upper plug is arranged above the feeding hopper 2, the cooling water pipe 3 is arranged inside the wall of the internal mixer, and a pair of rotors 4 with a speed ratio are driven by a motor 5 to complete the main mixing process. After the process is completed, the rubber compound is turned out of the internal mixer through the lower plug 6 and discharged by the discharge door 7. For the implementation of the subsequent training and learning of the predictive control algorithm, the differential equation needs to be transformed into the state space expression form. Take the state variable x = [T1 PT2] T , the control variable u = [nx u Q c T , at the sampling time T s , the state space expression linearized and discretized at the point (x d , u d ) is:

[0064]

[0065] where k is the number of training times; are respectively the state quantity, input quantity and output quantity of the system; A, B, and C are matrices of corresponding dimensions. The coefficient matrix is expressed as:

[0066]

[0067] According to the actual reaction process principle and the characteristics of the rubber compound, Table 1 gives the parameter values in the above coefficient matrix:

[0068] Table 1 Process mechanism model parameters

[0069]

[0070] S2. Using the above kinetic model and its parameter values, construct a prediction model and solve the objective function:

[0071] J = e k(k+1|k) T Qe k (k+1|k)+u k (k+1|k) T Ru k (k+1|k)

[0072]

[0073] Prediction horizon n x = 10, control horizon n u = 10, weight coefficient Q = I, R = 0.1I. Calculate the control sequence through the optimization principle:

[0074]

[0075] S3. As Figure 2 shown, combine the output of the model predictive controller with the actions of reinforcement learning, and set the training hyperparameters of reinforcement learning as shown in Table 2:

[0076] Table 2 Reinforcement learning algorithm parameters

[0077]

[0078] For problems such as the modeling error, state mutation, component degradation, and external interference of the prediction model in the internal mixer process, the actual values of the process parameters cannot accurately track the given target. Based on the model predictive control design, add a compensation control given by reinforcement learning to the system to correct the model deviation, and reconstruct the model prediction algorithm and cost function to improve the stability and control accuracy of the system, ensuring that the internal mixer process can track the preset trajectory. Considering the actual control scenario of the internal mixer temperature, the reasonable value range of the rotational speed is 0 - 150 r / min, set the control input constraint as [0, 150], and set the target temperature as 160 degrees Celsius. Figure 4 The results show that the intelligent model predictive control method based on reinforcement learning introduced in this paper has fast and stable tracking performance. As Figure 5 shown, the strategy proposed in this paper continuously explores the virtual environment, optimizes the control performance, and achieves the maximum cumulative reward within about 30 training episodes, so that the reward value quickly approaches zero, and the speed is fast and the accuracy is high. Figure 6 It shows that during the internal mixer temperature control process, the control input rotational speed always remains within a reasonable range. After the upper limit input for an initial period of time, the rotational speed continuously and stably changes with the tracking of the reference trajectory, ensuring the stable operation of the control system.

Claims

1. An intelligent model predictive control method for the kneading process based on reinforcement learning, characterized in that, The method acts on a rubber internal mixer system and specifically includes the following steps: S1. Establish a control model for the internal mixing process in the form of a state space representation; S2. Design a model predictive control method based on reinforcement learning and determine the objective function; S3. Solve the standard quadratic programming problem for the optimization variables; S4. The reinforcement learning selects a compensation term, combines it with the controller output, and acts on the rubber internal mixer system; S5. Calculate the reward and update the reward value to optimize the control compensation selection for future time steps.

2. The method according to claim 1, characterized in that, Based on the linearization and discretization of the differential equation, S1 obtains the following state-space model. The rubber temperature T1, the upper plug pressure P, and the cooling water temperature T2 are selected as the state variables x = [T1 P T2] T , the rotor speed n, the opening degree x of the electro-hydraulic proportional directional valve u and the cooling water flow rate Q c are selected as the control variables u = [n x u Q c T , and the state expression is obtained as follows:​ where k is the number of training times; are the state quantity, input quantity, and output quantity of the system, respectively; represents a real vector with the same dimension as the state quantity, represents a real vector with the same dimension as the input quantity, represents a real vector with the same dimension as the output quantity; A, B, and C are matrices of corresponding dimensions, and the specific expressions are as follows: where T s is the sampling time, K1 is the equivalent heat transfer coefficient, and K2 is the thermal resistance coefficient of the flow rate; C3 is the equivalent heat capacity of the rubber compound in the internal mixer, C4 is the equivalent heat capacity of the internal mixer wall, and C5 is the equivalent heat capacity of the cooling water in contact with the internal mixer wall; R6 represents the degree of heat transfer obstruction between the medium in the internal mixer and the internal mixer wall, R7 represents the degree of heat transfer obstruction between the internal mixer wall and the cooling water in the contact part, and R8 represents the degree of heat transfer obstruction between the internal mixer wall and the air; V1 is the initial pressure of the upper ram, V2 is the inlet temperature of the cooling water, x d , u d are the reaction equilibrium points respectively obtained from the collection of equipment operation data during the internal mixer production process.

3. The method according to claim 2, characterized in that, In S2, for the process control model of the internal mixing process, the compensation control δ(k) of reinforcement learning is considered, and thus the improved prediction model is expressed as: Where: k = 1, 2, 3... represents the number of training times, x(k) represents the state vector at the k-th moment, u(k) represents the control vector at the k-th moment, y(k) represents the output vector at the k-th moment, A, B, and C are coefficient matrices of corresponding dimensions, and δ(k) represents the control compensation term at the current moment; First, define the predicted output over N prediction intervals The input sequence u(k) of the training predictive controller is as follows: Combining the compensation control and the system process control model, the prediction model is derived: Among which M x and C u are respectively the initial state coefficient matrix and the input coefficient matrix of the prediction model, and G is the coefficient matrix of the compensation term, and the expression is: Based on the prediction model, a quadratic performance index is designed. This index consists of the error between the output reference trajectory and the predicted output and the control input value at the time step, and is used to measure the influence of the output error and the amplitude of the control input respectively; Weight matrices are added to each item, and by adjusting the magnitude of the weight coefficients, the objective function most suitable for the internal mixing process is determined: J = e k (k + 1|k) T Qe k (k + 1|k) + u k (k + 1|k) T Ru k (k + 1|k) Among them, e k (k + 1|k) is the error sequence at the prediction time k + 1, and J represents the system state x k with respect to the change in the control input u k of the objective function, k represents the current time, Q and R are weight matrices of appropriate dimensions, used to quantify the influence of state deviation and control input on the cost function; n x is the maximum time domain step of the predicted state, and n u is the maximum time domain step of the predicted control input.

4. The method according to claim 3, wherein The S3 uses the optimality condition to solve for the control input: Among them, Q bar , R bar are respectively the expanded state weight matrix and the control input weight matrix, and Y ref is the reference state sequence; within each time step, only the first control action in the calculated optimal sequence is adopted, then the control input signal at the k-th moment is: I is an identity matrix of a suitable dimension, and the subsequent zero matrix is used to ignore the subsequent control inputs in the control sequence and only retain the first one.

5. The method according to claim 4, wherein In S4, by designing the parameters in the reinforcement learning algorithm, data interaction is carried out with the system process control model established in S1, and compensation control is performed on the control system; Given that the process control model of the internal mixer system has been constructed in a discrete state space form, the actual output value, set value, and control signal input of the discharge temperature and upper plug pressure are selected to jointly constitute the environmental state s t ; After these environmental states closely related to the internal mixer working conditions are input into the reinforcement learning, the reinforcement learning generates the corresponding policy π(a t |s t ) according to the pre-designed learning parameters including the learning rate, exploration rate, and discount factor. In order to clarify the optimal process control strategy, the objective function is defined as follows: Among them, π→a means selecting action a according to policy π, and Q(s t ,a t ) is the Q-value of taking action a t in state s t , representing the expectation of future rewards that are expected to be obtained; means selecting the action with the maximum Q-value in the current state; by determining the optimal control policy, the current state is transferred to the next new state s t+1 ; where the Q-value function Q(s t ,a t ) is represented by the Bellman equation: Among them, represents the expected value of all future rewards after executing action a starting from state s; during this process, a reward value r(t) closely related to the current working condition environment is also obtained, and based on this, the network parameters are continuously learned and updated to continuously optimize the strategy; to evaluate the current environment s t for action a t the reward value r(t) is designed to update the state-action value function as follows: Q(s t ,a t ) ← Q(s t ,a t ) + α(r(t) + γ·maxQ(s t+1 ,a t+1 ) - Q(s t ,a t )) Among them, γ represents the discount factor, α is the learning rate, and maxQ(s t+1 ,a t+1 ) represents the maximum expected value of all possible actions under the policy π(a t+1 |s t+1 ); update the Q value according to the currently selected control action, as well as the current reward and the maximum Q value at the next time step; thus, the compensation control term can be obtained through the update of the Q table: δ(k) = Q(s t , a t ) According to the environmental state space information in the reinforcement learning algorithm Taking the selection of the policy as the action of reinforcement learning, denoted by a k (t) = π(s k (t)), subsequently, a k (t) is combined with the compensation control δ(k) and the highly optimized model predictive controller u k (t) after multiple training and learning, and accurately acts on the controlled object in the internal mixer process; at the same time, the system will quickly feedback new state information and immediate rewards, thus forming a closed-loop and continuously optimized process control loop; within each control cycle, reinforcement learning can continuously adjust its action selection and policy optimization according to the latest feedback information to adapt to various problems that may occur in the internal mixer process control.

6. The method according to claim 5, wherein In S5, in order to maintain the stability of the control system during the training process and perform effective model deviation correction, the negative norm of the error between the output reference trajectory and the predicted output is used as the reward value; The reward function can be obtained by calculating the following formula: r(t) = -(y - y ref ) T ·diag(w)·(y - y ref ) Among them, y represents the predicted output, and y ref represents the target reference trajectory; the updated reward value continuously acts on the update of the Q value to optimize the control strategy and reduce the error from the target state, thereby achieving better control: Q(s t , a t ) ← Q(s t , a t ) + α(r(t) + γ · maxQ(s t+1 , a t+1 ) - Q(s t , a t )) Continuously perform the above steps to form a closed-loop control strategy; As the control cycle progresses, continuously integrate new information, update the model prediction, and optimize the control input to adapt to possible changes in system behavior or external disturbances.

Citation Information

Patent Citations

  • Reinforced learning intermittent process control method based on improved AC algorithm

    CN116520703A

  • Banburying process model prediction control method based on bonding graph and neural network

    CN117055510A

  • Full-automatic online PAT intermittent crystallization control method, medium and system

    CN119015736A

  • Hybrid modelling and control of industrial batch distillation processes

    WO2024126785A1

Cited By

  • Adaptive optimization control method and system for energy consumption of internal mixer and storage medium

    CN121447781A