A process parameter optimization method based on reinforcement learning
Through the method of combining deep reinforcement learning and generative adversarial networks, process parameters are optimized, and the problems of unreliability and high trial and error cost of traditional manual adjustment process parameters are solved, and the reliability and stability of process parameter optimization are achieved.
Patent Information
- Application Number
- CN202310162833.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-24
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2043-02-24
AI Technical Summary
Traditional manual adjustment process parameters are unreliable during the production process, and the trial and error cost is too high, making it difficult to guarantee product quality.
DQN algorithm based on deep reinforcement learning is adopted to train uncertainty quantization models by generating adversarial networks, generate low-uncertainty data sets, and use this data set to train action-value functions to optimize process parameters.
It reduces trial and error costs, improves the reliability and stability of process parameter optimization, reduces the risk of process defects, and achieves dynamic optimization.
Smart Images

Figure CN116048028B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of industrial design and provides a process parameter optimization method based on reinforcement learning. Background Art
[0002] In the process of process production, the quality of the product mainly depends on the content of the raw materials, the status of the equipment operation and the process parameters, as well as the operating procedures for human intervention. The composition of the raw materials and the management of human behavior are generally fixed, so high-quality products depend on the appropriate process parameter settings. The optimization of process parameters is an important part of industrial production. However, in the process of process parameter optimization in actual production practice, it is generally adjusted by experts based on professional knowledge and experience, and there is almost no effective application of machine learning methods. However, due to the coupling characteristics between different process parameter variables, the relationship between product quality and process parameters is unclear, and the manual analysis method is difficult to practice and unreliable. The manual parameter adjustment method generally requires repeated trial and error to determine the final result, and since each adjustment needs to be actually applied to the production line before the effect can be seen, the trial and error cost is too high.
[0003] The research on process parameter optimization can be divided into two categories: the first category is static process parameter optimization, which obtains global optimization results based on the modeling model. For example, Tong Xi et al. used genetic algorithms to optimize process parameters after building an environmental model in the article "Optimization of TWIP steel heat treatment process parameters based on BP neural network and genetic algorithm. Hot processing technology" in "Hot processing technology"; the second category is dynamic optimization based on knowledge or historical cases, which gradually achieves optimization results through interaction. For example, Huang Bin et al. used fuzzy reasoning to control the heating furnace pressure parameters in the article "Application of heating furnace pressure control system based on hybrid fuzzy PID" in "Metallurgical Automation". Studying the dynamic optimization of process parameters can avoid the limitations of static methods, but the existing process parameter optimization methods using fuzzy reasoning are affected by knowledge, rules or case databases, and this information is usually difficult to obtain. The dynamic optimization program is a sequential decision problem that can be modeled as a Markov decision process, which can be solved by reinforcement learning. Reinforcement learning can learn a strategy to determine how to adjust process parameters based on the current product status. This method is obviously more reliable. The present invention utilizes the DQN algorithm of deep reinforcement learning to train and generate corresponding strategies. By combining deep learning with reinforcement learning, the action-value function can be better fitted than the traditional reinforcement learning algorithm, and the strategy for optimizing parameters can be directly obtained according to the action-value function. Summary of the invention
[0004] The present invention mainly solves the problem of unreliability of traditional manual adjustment of process parameters in the production process, which is accompanied by excessive trial and error costs. A process parameter optimization strategy is obtained by training with a reinforcement learning algorithm. After multiple iterations of the optimization strategy, the recommended process parameters are directly obtained. This process parameter can be directly used in the production line, which fully reduces the trial and error costs. At the same time, the uncertainty of the data is measured using an uncertainty quantification algorithm, and low-uncertainty data is used for reinforcement learning and fitting of product defect-process parameter training data, so that a relatively stable model can be obtained through limited data. When using the strategy to generate recommended new process parameters, the recommended process parameters are limited to low-uncertainty data, which reduces the risk of actually causing larger process defects when the new parameters are used in the production line.
[0005] In order to achieve the above object, the technical solution adopted by the present invention is:
[0006] A process parameter optimization method based on reinforcement learning includes the following steps:
[0007] In the first step, a data set containing process parameters and corresponding product status is input as the original sample data.
[0008] The second step is to use the generative adversarial network to train the uncertainty quantization model F GAN , where model F GAN Including generator and discriminator; based on F GAN The samples generated by the generator determine the uncertainty confidence interval [q 1 ,q 2 ], this method can be used to determine the uncertainty quantification value of the data. The uncertainty quantification model F based on the generative adversarial network is constructed GAN The specific steps are as follows: first, in order to construct the mapping distribution from the noise space to the original data space, the original data is input to reconstruct the distribution of the data space through the generative adversarial network; secondly, through F GAN The generator generates the reconstructed generated samples, and the generated samples are extracted using Markov Monte Carlo sampling to obtain new samples based on the original data distribution; then, 95% of the samples in the middle of the generated samples are selected as low uncertainty data, and the uncertainty confidence interval of the samples is defined based on the uncertainty quantification function Griewank [q 1 ,q 2 ]; Finally, using the confidence interval [q 1 ,q 2 ] Uncertainty is estimated for each sample in the original data, and the data with uncertainty quantification values within the confidence interval are screened out as low uncertainty data sets.
[0009] Further, the Griewank function is a confidence interval selection function f used to construct the confidence interval of the original data, and its formula is as follows:
[0010]
[0011] where d represents the dimension of the vector; x (i) represents the i-th component of the sample x.
[0012] Thirdly, train with the low-uncertainty data set, and use the reinforcement learning algorithm based on the DQN algorithm to obtain the action-value function as the Q function. The specific steps are as follows:
[0013] 3.1) The actions and updates in the reinforcement learning modeling are set as follows. Let x i ={x 1 ,…,x m ,x m+1 ,…,x M} be a sample of the currently executed optimization strategy. The form of this sample is the same as the original data (during training, this sample comes from the low-uncertainty data set in the second step, and during testing, the sample comes from the test set), with a total of M process parameters. Denote the number of process parameters that need to be adjusted during optimization as m (m < M). In the present invention, m process parameters are selected for change in each action, and their update directions are determined, while other parameters remain unchanged. The variation range of the m process parameters to be updated is within ±0.1. During the reinforcement learning training, an ∈-greedy random policy π θ is used to generate the discrete action a, where θ is the parameter corresponding to the policy. This policy selects a random action with a probability of ∈ and selects the action with the largest action-value function with a probability of 1 - ∈.
[0014] The discrete action a determines the process parameter components to be updated and the update direction of each component. Then, the process parameters need to be updated. For x 1 ,…,x m ,x m+1 ,…,x M selected from, for the selected x 1 ,…,x m , each process parameter is updated according to the update direction of each component, and the rule is as follows:
[0015]
[0016] where r is a random number generated from a [0, 1] distribution, and x old represents the value of the process parameter component before update, x newIt indicates the updated process parameter component value, direction=+1 means that the component update direction is positive, and direction=-1 means that the update direction is negative.
[0017] 3.2) Determine the calculation formula of the reinforcement learning reward value. First, use the low uncertainty data set and use the deep neural network to build the quality evaluation model F. Here, the neural network is the same as F in the second step. GAN The network structure of the discriminator is the same as that of the training set, and the training set is a low uncertainty data set. In the learning process of reinforcement learning, the defect description item is predicted by model F to calculate the reward value. Set the state at time t to the current process parameter value x t After executing an action, the process parameter value is updated to x t+1 ; In order to get closer to the optimal state at each step, the reward function is related to the defect prediction value F(x t+1 ) and the uncertainty quantification value f(x t+1 ), the reward function is defined as:
[0018]
[0019] in λ 1 ,λ 2 ,λ 3 is the coefficient.
[0020] 3.3) Determine the parameters required for reinforcement learning training: reinforcement learning reward discount factor γ; sample buffer pool capacity N; delay step size c; target defect value reduction ratio R.
[0021] 3.4) Finally, the low uncertainty dataset is used as reinforcement learning training samples, and reinforcement learning training is performed based on the DQN algorithm.
[0022] In each training, a sample is randomly selected from the reinforcement learning training samples for a round of training, with a total of 1000 rounds of training.
[0023] For each round of training, each iteration is defined as time t. The parameters θ of the initial Q function are inherited from the parameters obtained in the previous round of training (the parameters of the first round of Q function are randomly selected). According to the initial Q function of this round, actions are selected based on the ∈-greddy strategy, and the reward value is calculated by combining the process parameter values and actions before and after each action with the empirical sample {x t ,a t ,r t ,x t+1} in the data buffer pool; when the action reaches 50 times, the Q function training begins: after each action, the experience sample is stored in the data buffer pool, and then 20 samples are randomly sampled from the data buffer pool to train the Q function. The structure of the Q function is a neural network, and two networks are used during training: the prediction network Q θ (s,a) and target network Use the latest prediction network with parameters θ to calculate the Q value Q corresponding to the current state-action pair θ (x,a), using the parameters θ from a long time ago - The target network is used to predict the corresponding Q value after the next action According to the mean square error MSE of the two Q networks = (qQ θ (x,a)) 2 The gradient direction updates the parameter θ, and updates the target network every c steps The termination condition for each round of training is: the updated process parameter x terminal The corresponding uncertainty quantification value f(x terminal ) in [q 1 ,q 2 ], and the predicted value of the defect F(x terminal ) is lower than (1-r)F(x 0 ), where x 0 is the initial process parameter value of the sample. After completing 1000 rounds of training, the final Q function Q Θ (x,a), Θ is the parameter.
[0024] The fourth step is to use the Q function Q obtained in the third step Θ (x, a) launches the process parameter optimization strategy. In order to increase the feasibility of the optimization process, the present invention uses a random strategy. When selecting actions, the probability Randomly select, where Q Θ (x t ,a) indicates that when the process parameter value is x t When action a is selected, the corresponding cumulative reward expectation is Indicates Q Θ (x t ,a) achieve the minimum cumulative reward expectation corresponding to action a. According to the process of strategy optimization: first input the test sample x 0 In each iteration, action a is selected according to p(a), and then the process parameters are updated according to action a. After multiple iterations, if the final updated process parameters x terminal The corresponding uncertainty quantification value f(x terminal ) in [q 1 ,q 2 ], and the predicted value of the defect F(x terminal) is lower than (1-R)F(x 0 ), where x 0 is the initial process parameter value of the sample.
[0025] Step 5: When testing the use of this strategy, input the sample to be optimized (the sample comes from the test set), and output the recommended new parameters after iterative optimization according to the strategy in step 4.
[0026] The beneficial effects of the present invention are:
[0027] The present invention provides a method for optimizing process parameters based on reinforcement learning. The method uses defect description items and uncertainty quantification values to construct a reward function, learns a strategy that can achieve the maximum cumulative reward through the DQN algorithm, and then performs multi-step optimization through the strategy to obtain the value of the corrected process parameter. Such a method reduces the debugging cost. The optimization method proposed in the present invention belongs to dynamic optimization, which is reliable and more stable, and has strong feasibility. The process of training the model is very convenient and easy to integrate into the industrial production process. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 The figure is a flowchart for implementing a process parameter optimization method based on reinforcement learning. DETAILED DESCRIPTION
[0029] In order to make the method problems solved by the present invention, the method solutions adopted and the method effects achieved clearer, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It is understood that the specific embodiments described herein are only used to explain the present invention, rather than to limit the present invention. It is also necessary to explain that, for the convenience of description, only the parts related to the present invention are shown in the accompanying drawings, rather than all the contents.
[0030] Figure 1 The present invention provides a flowchart of a process parameter optimization method based on reinforcement learning for an implementation case of the present invention. The present invention provides an implementation case: the correlation data between the slab process parameters and the incidence of slab scale defects in a hot rolling production line of a steel plant is used as a modeling target. Figure 1 As shown, a Gaussian process robust optimization method based on uncertain quantization provided by an embodiment of the present invention includes:
[0031] Step 1: Input a data set containing process parameters and corresponding product status as the original sample data. The process parameters are 169-dimensional high-dimensional data, with a total of 479 data sets. These data are normalized to eliminate the dimensional differences between each process parameter.
[0032] The second step is to use the original sample data to train the uncertainty quantification model F based on the generative adversarial network. GAN, and output the uncertainty confidence interval [q 1 ,q 2 ] and selected low-uncertainty datasets.
[0033] The network structure and parameter settings in the generative adversarial network: The dimension of the noise vector of the generator is 20. The structure of the generator is a neural network with 4 fully connected layers. The activation function of the middle layer is the ReLU function, the number of neurons is 100, and the activation function of the output layer is the Sigmoid function; the structure of the discriminator is a neural network with 4 fully connected layers. The activation function of the middle layer is the ReLU function, the number of neurons is 50, and the activation function of the output layer is the Sigmoid function. The solvers for weight optimization in the generator and discriminator are Adam with learning rates of 0.001 and 0.002 respectively. The training batch of the generative adversarial network is 128, the number of iterations is 2000, and the generator and discriminator are trained alternately. After the generative adversarial network is trained, the number of MCMC samples using the generator based on the noise space is 20,000. Then, the samples are input into the Griewank function, and the middle 95% of the data are taken as high-certainty data to obtain the uncertainty confidence interval [0.592557, 0.734635]. Then, the original data set is screened, and the obtained low-uncertainty sample set contains 200 groups of data.
[0034] The third step is to use the low-uncertainty sample set training and the reinforcement learning algorithm based on the DQN algorithm to obtain the action-value function as the Q function.
[0035] 3.1) Data processing. Since the dimension of process parameters is too high, in order to prevent overfitting of noise data from affecting the performance of the model and the high-dimensional action space from causing the training to converge too slowly, the weights of the process parameters corresponding to the product defects generated by the defect tracing method are used, and only the six process parameter components with the highest weights are selected for optimization, and one process parameter component is optimized at a time.
[0036] 3.2) Train and fit the quality evaluation model of defect description-parameters. The training data set uses a low uncertainty sample set. The structure of the quality evaluation model network is a neural network with 4 fully connected layers. The activation function of the middle layer is the ReLU function, the number of neurons is 50, the activation function of the output layer is the Sigmoid function, the weight optimization solver is 0.001 Adam, the training batch is 32, and the number of iterations is 3000. The quality evaluation model F is constructed using Tensorflow. Substituting the quality evaluation model into formula (3) can obtain the reward function.
[0037] 3.3) Training reinforcement learning model. Set the Q network to contain 4 fully connected layers, the activation function of the middle layer is the ReLU function, the number of neurons is 100, the output layer directly outputs the sampled weight value for each action, the network weight optimization uses the Adam learner with a learning rate of 0.01, and the training batch is 32. For the target network, update the parameters to the same as the prediction network every 20 iterations; use the ∈-greddy strategy with a coefficient of 0.1 during training, and the strategy during training is a deterministic strategy, that is, directly select the action with the largest corresponding Q function value; set the decision to use a random strategy, that is, use the normalized result of the Q function value of each action as the probability of being selected; set the action to select 1 parameter each time to adjust according to the direction, so the action space size is 12; set the buffer pool size to 2000, store samples in the form of an array, and the appropriate buffer pool size is conducive to the use of samples during training. Start training the Q function after the experience buffer samples reach 50. In order to prevent the old samples from affecting the training, replace the old samples in order to store new samples after the buffer pool is full. Set the target defect value reduction ratio R = 90%.
[0038] After 1000 rounds of training, we get the DQN model and the corresponding action-value function Q function Q Θ (x,a).
[0039] The fourth step is to use the Q function Q according to reinforcement learning Θ (x, a) directly derives the strategy. In order to increase the feasibility of the optimization process, the present invention uses a random strategy. When selecting actions, the probability Random selection.
[0040] In the fifth step, the test data set is input and the process parameters are iteratively optimized using the strategy in the fourth step. Table 1 shows the final recommended results based on the optimization strategy.
[0041] Table 1 Optimization results of process parameter optimization method based on reinforcement learning
[0042]
[0043] Finally, in order to verify that the strategy as a dynamic optimization method can maintain good results even when the environment changes slightly, noise is added to the original data to simulate the data collected in the new environment, and the neural network is used to fit the quality evaluation model F new . Use the original strategy to test in the new environment (that is, set the quality evaluation model to F new , Q function Q Θ (x, a) remain unchanged), which proves that the strategy trained by reinforcement learning can better overcome the impact of slight changes in the environment. Table 2 shows the final recommended results based on the optimized strategy in the new environment. The experimental results prove the stability of the strategy.
[0044] Table 2 Optimization results of the process parameter optimization method based on reinforcement learning in a new environment with noise
[0045]
[0046]
[0047] Finally, it should be noted that the above embodiments are only used to illustrate the method scheme of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, ordinary method personnel in the field should understand that modifying the method schemes recorded in the aforementioned embodiments, or equivalently replacing some or all of the method features therein, does not cause the essence of the corresponding method scheme to deviate from the scope of the method schemes of the embodiments of the present invention.
Claims
1. A process parameter optimization method based on reinforcement learning, It is characterized in that It includes the following steps: In the first step, a data set containing process parameters and corresponding product status is input as the original sample data; The second step is to use the generative adversarial network to train the uncertainty quantization model F GAN , where model F GAN Including generator and discriminator; based on F GAN The samples generated by the generator determine the uncertainty confidence interval [q 1 ,q 2 ], determine the uncertainty quantification value of the data, and then obtain a low uncertainty data set; The third step is to use the low uncertainty data set for training and adopt the reinforcement learning algorithm based on the DQN algorithm to obtain the action-value function as the Q function; The fourth step is to use the Q function Q obtained in the third step Θ (x,a) Launch process parameter optimization strategy; Use a random strategy to select actions based on probability. Randomly select, where Q Θ (x t ,a) indicates that when the process parameter value is x t When action a is selected, the corresponding cumulative reward expectation is Indicates Q Θ (x t ,a) achieve the minimum cumulative reward expectation corresponding to action a; according to the strategy optimization process: first input the test sample x 0 In each iteration, action a is selected according to p(a), and then the process parameters are updated according to action a. After multiple iterations, the final updated process parameters x are satisfied. terminal The corresponding uncertainty quantification value f(x terminal ) in [q 1 ,q 2 ], and the predicted value of the defect F(x terminal ) is lower than (1-R)F(x 0 ), where x 0 is the initial process parameter value of the sample; Step 5: When testing the use of this strategy, input the sample to be optimized, and output the recommended new parameters after iterative optimization according to the strategy in step 4.
2. According to the process parameter optimization method based on reinforcement learning according to claim 1, It is characterized in that In the second step, the uncertainty quantification model F is constructed based on the generative adversarial network. GAN The specific steps are as follows: first, in order to construct the mapping distribution from the noise space to the original data space, the original data is input to reconstruct the distribution of the data space through the generative adversarial network; secondly, through F GAN The generator generates the reconstructed generated samples, and the generated samples are extracted using Markov Monte Carlo sampling to obtain new samples based on the original data distribution; then, 95% of the samples in the middle of the generated samples are selected as low uncertainty data, and the uncertainty confidence interval of the samples is defined based on the uncertainty quantification function Griewank [q 1 ,q 2 ]; Finally, using the confidence interval [q 1 ,q 2 ] Uncertainty is estimated for each sample in the original data, and the data with uncertainty quantification values within the confidence interval are screened out as low uncertainty data sets.
3. According to the process parameter optimization method based on reinforcement learning according to claim 2, It is characterized in that The Griewank function is a confidence interval selection function f used to construct the confidence interval of the original data, and its formula is as follows: Where d represents the dimension of the vector; x( i ) represents the i-th component of sample x.
4. According to the process parameter optimization method based on reinforcement learning according to claim 1, It is characterized in that The specific steps of the third step are: 3.1) The action and update settings in reinforcement learning modeling are as follows; let x i ={x 1 ,…,x m ,x m+1 ,…,x M } is a sample of the currently executed optimization strategy. The form of this sample is the same as the original data, with a total of M process parameters. The number of process parameters that need to be adjusted during optimization is recorded as m (m<M). In each action, m process parameters are selected for change and their update direction is determined. Other parameters remain unchanged. The range of change of the m process parameters to be updated is within ±0.
1. The random strategy π of ∈-greddy is used in reinforcement learning training. θ Generate a discrete action a, where θ is the parameter corresponding to the strategy, which selects a random action with probability ∈ and selects the action with the largest action-value function with probability 1-∈; After the discrete action a determines the process parameter components that need to be updated and the update direction of each component, the process parameters are updated. 1 ,...,x m ,x m+1 ,...,x M The x selected from 1 ,…,x m , update each process parameter according to the update direction of each component, the rules are as follows: Among them, r is a random number generated from a [0,1] distribution, x old represents the process parameter component value before updating, x new Indicates the updated process parameter component value, direction = +1 means that the component is updated in the positive direction, and direction = -1 means that the update direction is negative; 3.2) Determine the calculation formula of the reinforcement learning reward value; First, use the low uncertainty data set and use the deep neural network to build the quality evaluation model F. Here, the neural network is the same as F in the second step. GAN The network structure of the discriminator is the same as that of the training set, and the training set is a low uncertainty data set; in the learning process of reinforcement learning, the defect description item is predicted by model F to calculate the reward value; the state at time t is set to the current process parameter value x t , after executing an action, the process parameter value is updated to x t+1 ; In order to get closer to the optimal state at each step, the reward function is related to the defect prediction value F(x t+1 ) and the uncertainty quantification value f(x t+1 ), the reward function is defined as: in λ 1 ,λ 2 ,λ 3 is the coefficient; 3.3) Determine the parameters required for reinforcement learning training: reinforcement learning reward discount factor γ; sample buffer pool capacity N; delay step size c; target defect value reduction ratio R; 3.4) Using the low uncertainty data set as the reinforcement learning training sample, the reinforcement learning training is performed based on the DQN algorithm to obtain the final Q function Q Θ (x,a), Θ is the parameter.
5. According to the process parameter optimization method based on reinforcement learning according to claim 4, It is characterized in that The specific steps of the third step 3.4) are: In each training, a sample is randomly selected from the reinforcement learning training samples for a round of training, and training is performed for N rounds; For each round of training, each iteration is defined as time t. The parameters θ of the initial Q function are inherited from the parameters obtained in the previous round of training, and the parameters of the first round of Q function are randomly selected. According to the initial Q function of this round, the action is selected based on the ∈-greddy strategy, and the reward value is calculated by combining the process parameter values and actions before and after each action with the empirical sample {x t ,a t ,r t ,x t+1 } in the data cache pool; when the action reaches 50 times, the Q function training begins: after each action, the experience sample is stored in the data cache pool, and then 20 samples are randomly sampled from the data cache pool to train the Q function; the structure of the Q function is a neural network, and two networks are used during training - the prediction network Q θ (s,a) and target network Use the latest prediction network with parameters θ to calculate the Q value Q corresponding to the current state-action pair θ (x,a), using the parameter θ - The target network is used to predict the corresponding Q value after the next action According to the mean square error MSE of the two Q networks = (qQ θ (x,a)) 2 The gradient direction updates the parameter θ, and updates the target network every c steps The termination condition for each round of training is: the updated process parameter x terminal The corresponding uncertainty quantification value f(x terminal ) in [q 1 ,q 2 ], and the predicted value of the defect F(x terminal ) is lower than (1-R)F(x 0 ), where x 0 is the initial process parameter value of the sample; after completing N rounds of training, the final Q function Q Θ (x,a), Θ is the parameter.
Citation Information
Patent Citations
Intrusion detection method based on noise network and reinforcement learning
CN113392878A
Production process parameter-oriented Gaussian process robust optimization method
CN114742289A