Industrial process control method based on reinforcement learning of second-order value gradient model

The second-order value gradient model enhances reinforcement learning for industrial processes by improving data utilization and control precision, addressing inefficiencies in existing methods and enabling effective continuous action control.

CN120315380APending Publication Date: 2025-07-15SHENYANG INST OF AUTOMATION - CHINESE ACAD OF SCI
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202410046760.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-12
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing reinforcement learning industrial process control methods have problems such as inefficient data utilization, inability to perform continuous action control, and limited application scenarios, especially in complex industrial processes, which are difficult to achieve efficient automatic control.

Method used

The reinforcement learning method based on the second-order value gradient model is adopted. By adding second-order gradient information of the state value function in the model training process, combining deep neural networks and high-state value sampling strategies, the learning iteration efficiency and the accuracy of the model are improved, and virtual control data is used to train the agent for automatic control.

Benefits of technology

It realizes efficient continuous action control in complex industrial processes, improves model prediction accuracy and robustness, reduces the number of environmental interactions, and improves data utilization and control effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120315380A_ABST
    Figure CN120315380A_ABST
Patent Text Reader

Abstract

The invention relates to an industrial process control method based on reinforcement learning of a second-order value gradient model. Aiming at the characteristics of nonlinearity, strong coupling, hysteresis and the like of industrial process control, the invention provides a control method based on reinforcement learning of a second-order value gradient model on the basis of a value perception model. Firstly, in the model training process of the method, second-order gradient information of a state value function is added, so that the method has more accurate function approximation capability and higher robustness, and the learning iteration efficiency is higher; and secondly, by adopting a new state sampling strategy, the model can be more efficiently utilized to carry out strategy learning. Finally, an experiment in a penicillin production fermentation industrial process scene shows that compared with a traditional maximum likelihood estimation model, an environment model prediction error based on a second-order value gradient model is remarkably reduced; the learning efficiency of the reinforcement learning method based on the second-order value gradient model is superior to that of an existing strategy optimization method based on the model, better control performance is achieved, and the oscillation phenomenon in the control process is reduced. The method has theoretical and practical significance for automatic control of the industrial process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of industrial process automation control, and specifically relates to a control method based on model reinforcement learning. Background Art

[0002] With the continuous development of modern industry, process control technology has gradually become an indispensable part of industrial processes. Control methods that are accurate, stable and reliable, and can quickly respond to and adapt to complex changes can improve product quality, production efficiency, safety, and create better economic benefits. In most industrial processes, such as the penicillin fermentation production process, the food production and processing process, etc., control methods face many challenges: the complex dynamic characteristics, hysteresis, and strong non-linear coupling between variables of industrial process systems. Therefore, in the face of these challenges, reasonable control methods need to be studied.

[0003] In order to improve the control performance of complex industrial process systems, researchers have proposed using reinforcement learning methods to control industrial processes. For non-linear strongly coupled systems, Luo Ao et al. applied the execution-evaluation structure in reinforcement learning to the control strategy and achieved good control effects. However, in the face of multi-input constraints, there is a problem of decreased control performance; for complex lagging industrial processes, Zhang Danyang improved the Deep Deterministic Policy Gradient (DDPG) method using an intrinsic curiosity reward generation method and applied it to the control of the beer fermentation process to obtain an optimal control method. However, the simulation experimental environment used in this method is relatively ideal and needs to be verified in more industrial process scenarios; for the combined optimization control of the load distribution of chillers, the frequency of cooling tower fans, and the frequency of cooling water pumps, Ma Shuai et al. proposed an improved double-pool DQN algorithm. This method can effectively reduce energy consumption, but this method can only perform discrete control and cannot perform continuous action control; for the scheduling problem of textile machine manufacturing workshops with sequence-related setup times in complex process environments, Ji Zhiyong et al. proposed a reinforcement learning training algorithm with multiple action spaces, but ignored the influence of multiple factors on the scheduling goal; Ren Anni et al. optimized traffic signal control using a reinforcement learning method based on the attention mechanism, but this method relies on a large amount of training data.

[0004] The above methods are all model-free reinforcement learning control methods proposed for complex industrial process systems. These methods require a large amount of data generated by the interaction between the agent and the environment to train the policy network, which limits their applicability in complex real-world scenarios. The model-based reinforcement learning control method can solve the problem of inefficient data utilization in the model-free reinforcement learning control method. The model-based reinforcement learning control method constructs an explicit model of the environment. By using the learned environment model, even in a high-dimensional state space, the agent can interact with the environment model and optimize its policy, thereby reducing the number of real environment interactions required and improving data utilization. Summary of the Invention

[0005] Aiming at the problems of inefficient data utilization, inability to perform continuous action control, and limited application scenarios in the existing reinforcement learning industrial process control methods, based on the value gradient model, the present invention provides an industrial process control method based on second-order value gradient model reinforcement learning. First, during the model training process, the method adds the second-order gradient information of the state value function, has more accurate function approximation ability and higher robustness, and higher learning iteration efficiency; second, by adopting a new state sampling strategy, the model can be used more efficiently for policy learning.

[0006] The technical solution adopted by the present invention to achieve the above object is: an industrial process control method based on second-order value gradient model reinforcement learning, which iteratively obtains an agent with an optimal control strategy through the following steps. The agent is used to control the industrial process to make the control effect optimal, and includes the following steps:

[0007] In the penicillin fermentation production industrial process, use the agent strategy for control to obtain actual control data;

[0008] Use the actual control data and, according to the second-order value gradient loss function formula, iteratively train the deep neural network of the environment model;

[0009] Use the state sampling strategy to sample an initial state from the actual control data. Starting from this state, use the agent to interact with the environment model, that is, use the agent for control, so that the environment model predicts and outputs virtual control data;

[0010] Use the virtual control data to iteratively update the parameters of the agent, including the parameters of the double value network module, the double target network module, and the policy network module, to obtain an updated agent;

[0011] Use the updated agent to return to step 1 and continue the iterative loop until the control effect of the agent reaches the optimal.

[0012] In the industrial process of penicillin fermentation production, the cold water value is used as the controlled variable a in the actual control process t , and the oxygen concentration, cell concentration, penicillin concentration, culture medium volume, carbon dioxide concentration, fermenter reaction temperature, and temperature difference at the current moment are used as state variables s t , and the control objective is to keep the temperature in the fermenter at 297.5K;

[0013] Each time the cold water value is controlled, a piece of control data is recorded, where the control data includes: the current state quantity s t , the control quantity a t , the benefit r t , and the next state quantity s t+1 , that is, (s t ,a t ,r t ,s t+1 ); among them, the benefit r t is

[0014]

[0015] Among them, err is the difference between the current temperature in the fermenter and 297.5K, and σ1, σ2, σ3 are the thresholds of the temperature difference.

[0016] The deep neural network of the described environmental model includes:

[0017] Model input and output dimensions: batchsize×n×m features ;

[0018] Among them, batchsize is the batch size set during the batch training of the model, n is the number of integrated models, and m features is the number of dimensions of the observed state plus the number of dimensions of the control state, and is a positive integer;

[0019] Based on the current state-action pair (s t ,a t ), determine the distribution of the next state

[0020] Among them, N represents the Gaussian distribution, μ θ (.) represents the mean matrix, ∑ θ (.) represents the variance matrix, and μ θ (.) and ∑ θ (.) are constructed using a neural network.

[0021] The formula of the second-order value gradient loss function is as follows:

[0022]

[0023] Among them, H(x) is the Hessian matrix of the value function, S is the state vector, and V(s) represents the value function. Represents the first-order gradient of the value function. The point s′ is a reference point near the state s, and f θ (s i , a i ) represents the environment model, and θ is the environment model parameter.

[0024] The training termination condition is that the root mean square error between the numerical value of the virtual control data output by the model and the actual numerical value is within the threshold range.

[0025] For the state sampling strategy, Boltzmann probability distribution sampling is used, and the formula is as follows:

[0026] p(s) ∝ e βV(s)

[0027] Among them, β is a hyperparameter that controls the proportion of high-value states, and V(s) is the estimated value of the state value through the value network of the agent. Substitute the current state quantity s in all actual control data t into the above formula, calculate their probabilities, and then sample according to the probability calculation results to obtain an initial state s t .

[0028] Updating the parameters of the agent includes the parameters of the dual value network module, the dual target network module, and the policy network module, and includes the following steps:

[0029] The parameters of the dual value network module are ω1 and ω2 respectively, and they respectively correspond to the dual target network module. Among them, for the two value networks, a network with a smaller value is selected each time; the update method of the two target networks is

[0030] Taking the virtual control data as the input, training and updating the parameters of the policy π network module according to the Soft Actor-Critic method to make the loss function converge, and obtaining the updated agent.

[0031] The present invention has the following beneficial effects and advantages:

[0032] The present invention relates to an industrial process control method based on reinforcement learning of a second-order value gradient model, which is a method based on a state perception model. By using the idea of a value perception loss function and high-state value sampling, the environment model can learn more accurate function approximation capabilities and higher robustness, with higher learning iteration efficiency. It clearly solves the mismatch between the learning objectives of the model and the policy, effectively improves the accuracy of the prediction model, thereby obtaining a large amount of virtual control data to train the agent, and using the agent to automatically control the industrial process. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 It is a schematic diagram of an industrial example of the method of the present invention.

[0034] Figure 2 It is a flowchart of the method of the present invention.

[0035] Figure 3 It is a comparison chart of the prediction errors of different models on the test set.

[0036] Figure 4 It is a comparison chart of the control effects of different methods in the industrial process. DETAILED DESCRIPTION OF THE INVENTION

[0037] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe in detail the specific implementation method of the present invention with reference to the accompanying drawings. Many specific details are set forth in the following description in order to fully understand the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the invention. Therefore, the present invention is not limited by the specific implementations disclosed below.

[0038] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention.

[0039] An industrial process control method based on reinforcement learning of a second-order value gradient model preprocesses and standardizes industrial process data to form three-dimensional model input data. By combining a deep neural network environment model with a second-order value perception loss function and high-state value sampling, the environment model can learn more accurate function approximation capabilities and higher robustness, with higher learning iteration efficiency, clearly solving the mismatch in learning objectives between the model and the strategy, effectively improving the accuracy of the prediction model, and thus obtaining a large amount of virtual control data using the environment model. Finally, a large amount of virtual control data is used to train the control strategy of the intelligent agent, so as to automatically control the industrial process using the intelligent agent. The programming languages used in the program execution steps of the present invention are not limited to MATLAB, Python, Java, etc.

[0040] As Figure 1 and Figure 2 shown, it is a schematic diagram of an industrial example of the present invention. In the industrial process of penicillin fermentation production, it is necessary to collect acid flow, alkali flow, hot water, and cold water so that a reaction is carried out in the fermenter to prepare penicillin. In the present invention, it is assumed that the flow conditions of acid flow, alkali flow, hot water, etc. remain unchanged, and the cold water value is used as the controlled variable a in the actual control process t , for preparation. During the preparation process, 7 variables such as the oxygen concentration, cell concentration, penicillin concentration (in grams per liter), culture medium volume (in liters), carbon dioxide concentration, fermentation tank reaction temperature, and temperature difference at the current moment are collected by chemical methods or sensors as the state variable s t , and the control objective is to keep the temperature at 297.5K.

[0041] The specific steps of the present invention are as follows:

[0042] Step 1: Use the intelligent agent strategy to control in the industrial process of penicillin fermentation production to obtain actual control data;

[0043] Step 1-1: Use the intelligent agent to control in the industrial process;

[0044] Step 1-2: Record each piece of control data, and each piece of control data includes: the current state quantity, the control quantity, the reward, and the next state quantity.

[0045] In the industrial process of penicillin fermentation production, the cold water value is used as the controlled variable a in the actual control process t , and the 7 variables such as the oxygen concentration, cell concentration, penicillin concentration (in grams per liter), culture medium volume (in liters), carbon dioxide concentration, fermentation tank reaction temperature, and temperature difference at the current moment are used as the state variable s t , and the control objective is to keep the temperature in the fermenter at 297.5K;

[0046] Each time the cold water value is controlled, a control data record is made. Among them, the control data includes: the current state quantity s t , the control quantity a t , the revenue r t , the next state quantity s t+1 , that is, (s t , a t , r t , s t+1 ); among them, the revenue r t is

[0047]

[0048] Among them, err is the difference between the current temperature and 297.5K, and σ1, σ2, σ3 are the thresholds of the temperature difference.

[0049] Step 2: Collect all the actual control data;

[0050] Step 3: Use the actual control data and, according to the second-order value gradient loss function formula, iteratively train the deep neural network of the environment model;

[0051] Step 3-1: Determine the model input dimension: batchsize×n×m features , where batchsize is the batch size set during the batch training of the model, n is the number of integrated models, here n is generally 7, and m features is the number of dimensions of the observation state plus the number of dimensions of the control state, generally a positive integer; let the environment model be With its parameter being θ, then based on the current state-action pair (s t , a t ), the distribution of the next state can be written as Among them, N represents the Gaussian distribution, μ θ (.) represents the mean matrix, and ∑ θ (.) represents the variance matrix. Here we can use a neural network to construct μ θ (.) and ∑ θ (.).

[0052] Step 3-2: According to the model input dimension in Step 3-1, determine the structural parameters of the environment model;

[0053] Step 3-3: Use the actual control data and train the model by the gradient descent method. The loss function during training is the second-order value gradient loss function, as follows:

[0054]

[0055] Among them, H(x) is the Hessian matrix of the value function, s is the state vector, V(s) represents the value function, represents the first-order gradient of the value function, the reference point of point s′ near state s, f θ (s i ,a i ) represents the environmental model, and θ is the parameter of each probability model;

[0056] When the loss function converges, the training stops. At this time, the root mean square error between the value of the virtual control data output by the model and the actual value is within the threshold range.

[0057] Step 4: Use the state sampling strategy to sample an initial state from the actual control data. Starting from this state, use the agent to interact with the environmental model, that is, use the agent to perform control in the environmental model and output virtual control data.

[0058] The state sampling strategy uses the Boltzmann probability distribution for sampling, and the formula is as follows:

[0059] p(S)∝e βV(s)

[0060] Among them, β is the hyperparameter that controls the proportion of high-value states, V(s) is the estimated value of the state value through the value network of the agent. Substitute the current state quantity s in all actual control data t into the above formula, calculate their probabilities, and then sample according to the probability calculation values to obtain an initial state s t .

[0061] Step 5: Collect all the virtual control data under the initial state s t ;

[0062] Step 6: Use the virtual control data in the database to update the parameters of the agent, including the parameters of the dual value network module, the dual target network module, the policy network module, etc., to obtain the updated agent; among them, the dual value network module Q ω (s t ,a t ) takes the state s and the control action a as inputs, and then outputs a scalar, representing the value obtained by taking the action a in the state s. The input of the dual target network module is the state s and the control action a, and then it outputs a scalar, which is also the value obtained by taking the action a in the state s. The structure of the target network is the same as that of the Q network, but its parameters are updated with a delay. The input of the policy π network module π θ is the state s, and the output is the probability distribution of the control action in this state, and then the action is sampled through the action probability distribution.

[0063] Step 6-1: Update the parameters of the dual-value network module and the dual-target network module;

[0064] The above two modules model two action-value functions, namely the Q-value function (parameters are ω1 and ω2) and a policy function π (parameter is θ). Based on the idea of Double DQN, two Q-value networks are used, but each time a Q-value network is used, a network with a smaller Q-value is selected, thus alleviating the problem of overestimation of values. The loss function of any Q-value function (Q-value network) is: where R is the data collected by the policy in the past, that is, all the virtual control data collected in Step 6, (s t ,a t ,r t ,s t+1 ) is one of the data, γ is the discount factor, the capital letter E represents taking the mathematical expectation, Q ω (s t ,a t ) represents the Q-value network with parameter ω, represents the target network, corresponding one-to-one with the Q-value network, the ~ line represents random sampling according to the probability distribution, and π θ represents the policy network with parameter θ. α represents the entropy regularization term coefficient. In a state where the optimal action is uncertain, the value of entropy should be a bit larger; while in a state where a certain optimal action is relatively certain, the value of entropy can be smaller. The larger α is, the stronger the exploration ability, which helps to accelerate subsequent policy learning and reduce the possibility of the policy falling into a poor local optimum.

[0065] To make the training more stable, the target Q-value network is used here There are also two target Q networks, corresponding one-to-one with the two Q-value networks. The update method of the two target Q networks is soft update where τ is the soft update parameter.

[0066] Step 6-2: Update the parameters of the policy network module.

[0067] The loss function of the policy π is obtained from the KL divergence, and the calculation formula is: Q ω (s t ,a t ) represents the Q-value network with parameter ω, π θ represents the policy network with parameter θ, α represents the entropy regularization term coefficient, and R is the data collected by the policy in the past.

[0068] Step 7: At this time, all the parameters of the policy network module in the agent have been updated. Use the updated agent to enter Step 1, that is, use the agent's policy for control, and the obtained control data is the control data using the new policy. The whole process loops continuously until the agent meets the expected control effect of the industrial process, that is, the profit within a certain period meets the industrial expected requirements.

[0069] The results of the above method are as Figure 3 , Figure 4 shown. Figure 3 It is the comparison result of model errors in different model-based control methods during the actual penicillin fermentation production process. To further verify the industrial process control effect of the method of the present invention, we adopt the same initial state, set the initial fermentation temperature to 298.35K, and use the policy networks obtained by PETS, SAC, MBPO, VaGraM, and the method based on the second-order value gradient model to control the cold water flow value respectively. The goal is to make the penicillin fermentation environment more suitable and stable, and the temperature adjustment goal is 297.5K. Figure 4 It is a comparison chart of the actual control effects of each method. As can be seen from Figure 4 , the VaGraM method and the method based on the second-order value gradient model reach near 297.5K first and can maintain stability. Among them, after approaching 297.5K, the method based on the second-order value gradient model is more stable than VaGraM and does not show fluctuations. The adjustment granularity of the method based on the second-order value gradient model is finer. Compared with classical reinforcement learning methods such as PETS, SAC, and MBPO, both the deviation amount and the overall smoothness have been significantly improved.

[0070] In summary, an industrial process control method based on second-order value gradient model reinforcement learning proposed by the present invention uses the ideas of value-aware loss function and high-state value sampling, enabling the environment model to learn more accurate function approximation ability and higher robustness, with higher learning iteration efficiency. It clearly solves the learning objective mismatch between the model and the policy, effectively improves the accuracy of the prediction model, thereby obtaining a large amount of virtual control data to train the agent, and using the agent to automatically control the industrial process.

[0071] For the present invention, aiming at the defect of low sampling efficiency of model-free deep reinforcement learning and the problem of inaccurate model prediction in model-based reinforcement learning, a training method based on the second-order value gradient model is proposed. Aiming at the problems of overshoot and oscillation that usually exist in traditional control in multi-input single-output industrial processes, an industrial process control method based on second-order value gradient model reinforcement learning is proposed. It has theoretical and practical significance for the automatic control of industrial processes.

[0072] The embodiments described in the above description will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several changes and improvements can be made without departing from the concept of the present invention. These all fall within the protection scope of the present invention.

Claims

1. An industrial process control method based on reinforcement learning of a second-order value gradient model, characterized in that An agent that obtains an optimal control strategy through repeated iterations of the following steps, where the agent is used to control an industrial process to achieve optimal control effects, including the following steps: In the industrial process of penicillin fermentation production, use the agent strategy for control to obtain actual control data; Use the actual control data and, according to the second-order value gradient loss function formula, iteratively train the deep neural network of the environment model; Use the state sampling strategy to sample an initial state from the actual control data. Starting from this state, use the agent to interact with the environment model, that is, use the agent for control, so that the environment model predicts and outputs virtual control data; Use the virtual control data to iteratively update the parameters of the agent, including the parameters of the dual value network module, the dual target network module, and the policy network module, to obtain an updated agent; Use the updated agent to return to step 1 and continue the iterative loop until the control effect of the agent reaches the optimal; 2. The industrial process control method based on second-order value gradient model reinforcement learning according to claim 1, characterized in that, In the industrial process of penicillin fermentation production, the cold water value is used as the controlled variable a in the actual control process t , and the oxygen concentration, cell concentration, penicillin concentration, medium volume, carbon dioxide concentration, fermenter reaction temperature, and temperature difference at the current moment are used as the state variables s t , and the control objective is to keep the temperature in the fermenter at 297.5K; Each time cold water value control is performed, a control data record is made. Among them, the control data includes: the current state quantity s t , the control quantity a t , the revenue r t , the next state quantity s t+1 , that is, (s t , a t , r t , s t+1 ); among them, the revenue r t is Where err is the difference between the current temperature in the fermenter and 297.5K, and σ1, σ2, σ3 are the thresholds of the temperature difference.

3. The industrial process control method based on second-order value gradient model reinforcement learning according to claim 1, characterized in that The deep neural network of the environment model includes: Model input and output dimensions: batchsize × n × m features ; Among them, batchsize is the batch size set during the batch training of the model, n is the number of integrated models, and m features is the sum of the dimension of the observation state and the dimension of the control state, and is a positive integer; Based on the current state-action pair (s t , a t ), determine the distribution of the next state Among them, N represents the Gaussian distribution, and μ θ (.) represents the mean matrix, and ∑ θ (.) represents the variance matrix, and a neural network is used to construct μ θ (.) and ∑ θ (.).

4. The industrial process control method based on second-order value gradient model reinforcement learning according to claim 1, wherein The second-order value gradient loss function formula is as follows: Among them, H(x) is the Hessian matrix of the value function, s is the state vector, and V(s) represents the value function. denotes the first-order gradient of the value function, the reference point of point s′ near state s, f θ (s i , a i ) represents the environmental model, and θ is the environmental model parameter.

5. The industrial process control method based on second-order value gradient model reinforcement learning according to claim 1, wherein, The training termination condition is that the root mean square error between the numerical value of the virtual control data output by the model and the actual numerical value is within the threshold range.

6. The industrial process control method based on second-order value gradient model reinforcement learning according to claim 1, wherein The state sampling strategy uses the Boltzmann probability distribution for sampling, and the formula is as follows: p(s) ∝ e βV(s) Among them, β is a hyperparameter that controls the proportion of high-value states, V(s) is the estimated value of the state value through the value network of the agent, and the current state quantity s in all actual control data t is substituted into the above formula to calculate their probabilities, and then sampled according to the probability calculation results to obtain an initial state s t .

7. An industrial process control method based on reinforcement learning of a second-order value gradient model according to claim 1, characterized in that, Updating the parameters of the agent, including the parameters of the dual value network module, the dual target network module, and the policy network module, includes the following steps: The parameters of the dual-value network module are ω1 and ω2 respectively, and they correspond to the dual-objective network module respectively; among them, for the two value networks, one with a smaller value is selected each time; the update method of the two objective networks is Taking the virtual control data as the input, training and updating the parameters of the policy π network module according to the Soft Actor-Critic method to make the loss function converge, and obtaining an updated agent.

Citation Information

Patent Citations

  • Reinforcement learning exploration method and device based on generative adversarial mechanism

    CN112052936A

  • Industrial control method, device and system based on reinforcement learning and electronic equipment

    CN114721345A

  • Continuous body mechanical arm motion control method based on deep reinforcement learning

    CN116038691A

  • Heat exchange station control method based on reinforcement learning

    CN116430732A

  • Path planning method fusing deep neural network and reinforcement learning method

    CN116448117A