A Sickle Curve Control Policy Generation Method Based on Deep Reinforcement Learning
By employing deep reinforcement learning methods and utilizing XGBoost and DDPG algorithms to generate sickle bend control strategies, and automatically adjusting process parameters, the accuracy and lag issues of traditional control methods in hot rolling mills are resolved, thereby improving the automation rate and quality of strip steel production.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- PANZHIHUA IRON & STEEL RES INST OF PANGANG GROUP
- Filing Date
- 2023-08-29
- Publication Date
- 2026-05-26
AI Technical Summary
Traditional methods for controlling the camber curve are difficult to meet production requirements in hot rolling mills due to insufficient accuracy and lag, leading to a decline in strip steel product quality and the occurrence of production accidents.
A deep reinforcement learning-based approach is adopted, using the XGBoost algorithm to establish a sickle bend prediction model, and combining it with the DDPG algorithm to generate a sickle bend control strategy, automatically calculating process parameter adjustment values to control the straightness of the strip.
It improves the automation rate of hot rolling production and the quality of strip steel, reduces manual intervention, lowers labor intensity, and effectively avoids the adverse effects of sickle bend on production.
Smart Images

Figure CN117075557B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of hot strip rolling production, specifically to a method for generating sickle bend control strategies based on deep reinforcement learning. Background Technology
[0002] With the rapid development of the national economy, especially the booming industries of automobile manufacturing, aerospace, home appliances, and precision instruments, the demand for strip steel is increasing, and the requirements for its product quality are also becoming more stringent. The sickle-shaped bend phenomenon is a difficult problem in strip steel production, easily causing wedge-shaped export strip steel, and even leading to accidents such as tailing and steel piling during production. This seriously affects product quality and normal production, resulting in huge economic losses and resource waste.
[0003] Due to the nonlinear and time-varying characteristics of hot roughing and the diversity of production data, various factors do not exist in isolation but often interact, resulting in a highly complex mechanism for camber defects. Traditional control methods are insufficient to meet the production requirements of hot rolling mills. Typically, camber control relies on operators manually adjusting the strip production input characteristics, which is not only lacking in accuracy but also exhibits a certain degree of lag. Therefore, providing a camber control method adapted to the strip production environment is crucial for improving strip production quality. Summary of the Invention
[0004] The purpose of this invention is to provide a method for generating sickle bend control strategies based on deep reinforcement learning. By using deep learning and reinforcement learning to model the generation and control process of sickle bend in hot roughing strip, the method automatically calculates the process parameter settings for strip production in subsequent passes, ensuring the straightness of intermediate billets in strip production and improving the automation rate and quality of hot roughing production.
[0005] To achieve the above objectives, this invention proposes a method for generating sickle-shaped control policies based on deep reinforcement learning, comprising the following steps:
[0006] S1) Collect process parameters during the hot roughing strip rolling process as intermediate billet input data, and quantitative data of intermediate billet bending state as intermediate billet response data. Preprocess the intermediate billet input data and intermediate billet response data. The preprocessing is divided into five parts in sequence: abnormal data processing, standardization processing, centerline offset smoothing processing, calculation of intermediate billet sickle bending amount and data merging.
[0007] S2) Establish and train a sickle bending amount prediction model based on the XGBoost algorithm. The model outputs the predicted sickle bending amount based on the characteristics of the input rolling process parameters. The sickle bending amount prediction model based on the XGBoost algorithm predicts the sickle bending amount at the end of a certain rolling pass based on the input data of a certain rolling pass, i.e. the rolling process parameters, thereby simulating the sickle bending state. The trained sickle bending amount prediction model based on the XGBoost algorithm is obtained through training.
[0008] S3) Establish a sickle bend control strategy generation model based on a deep deterministic strategy gradient algorithm. This control strategy generation model uses the sickle bend bending amount prediction model to simulate the sickle bend bending amount to build an intelligent agent environment. Based on the input data and the trained sickle bend bending amount prediction model based on the XGBoost algorithm, the sickle bend control strategy generation model calculates the adjustment value of the process parameters to be adjusted.
[0009] S4) Train the sickle bend control strategy generation model based on the deep deterministic policy gradient algorithm, and save the model parameters to obtain the trained sickle bend control strategy generation model based on the deep deterministic policy gradient algorithm.
[0010] S5) Input the input data of the intermediate billet of hot roughing strip into the trained sickle bend control strategy generation model based on the deep deterministic strategy gradient algorithm, and output the sickle bend control strategy of the intermediate billet.
[0011] This invention analyzes a large amount of historical data and combines it with actual business analysis to identify key process parameters affecting camber, determine the set of process parameters to be adjusted, and improve the camber state by controlling the values of these parameters. This ensures the straightness of the slab in the length direction after hot roughing rolling, reducing the adverse effects of camber on strip rolling. This technical solution does not use a mechanistic model; instead, it models the production control process of camber in hot roughing rolling and constructs a camber amount prediction model based on the XGBoost model and a camber control strategy generation model based on DDPG. Utilizing automatic control based on reward feedback in reinforcement learning, it can effectively model the nonlinear coupling relationship of multiple process parameters in hot roughing rolling, thereby finding optimized values of the process parameters to be adjusted. The implementation of this technical solution, through automatic control, can effectively avoid control errors caused by current reliance on human experience, reduce manual intervention, and lower labor intensity. Simultaneously, controlling the straightness of the roughing slab provides a strong guarantee for the stability of finishing rolling production. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below.
[0013] Figure 1 A flowchart of a method for generating sickle-shaped control strategies based on deep reinforcement learning;
[0014] Figure 2 The image shows the prediction results of the sickle bend amount prediction model based on the XGBoost algorithm.
[0015] Figure 3 The figure shows the feature importance analysis results of the sickle bend amount prediction model based on the XGBoost algorithm.
[0016] Figure 4 A model structure diagram for generating a sickle bend control strategy based on a deep deterministic policy gradient algorithm;
[0017] Figure 5 A flowchart illustrating the practical application of a deep reinforcement learning-based sickle curve control strategy generation method.
[0018] Figure 6 Generate the model test results for all samples of the sickle bend control strategy based on the deep deterministic policy gradient algorithm;
[0019] Figure 7 The result of testing the sickle bend control strategy generation model based on the deep deterministic policy gradient algorithm for samples with a sickle bend bending amount greater than 5.
[0020] Figure 8 The image shows the test results of a sickle bend control strategy generation model based on a deep deterministic policy gradient algorithm, with samples having a sickle bend curvature greater than 10. Detailed Implementation
[0021] The methods described in this invention will be further explained below with reference to the accompanying drawings and specific embodiments.
[0022] Figure 1 This is a flowchart of a sickle-shaped control policy generation method based on deep reinforcement learning proposed in this invention. The method specifically includes the following steps:
[0023] S1) Collect process parameters (such as rolling force, rolling speed, etc.) during the hot roughing and rolling of strip steel as input data for intermediate billets, and quantify the bending state of intermediate billets as response data for intermediate billets. Preprocess the input and response data of the intermediate billets. The preprocessing is divided into five parts in sequence: abnormal data processing, standardization processing, centerline offset smoothing processing, calculation of the bending amount of the intermediate billet's sickle bend, and data merging. The specific steps are as follows:
[0024] S11) Remove data samples from the intermediate blank input data that have missing fields, duplicate content, or values outside the normal range, and use the remaining fields as input feature information for the model. This mainly consists of the following steps:
[0025] (S111) During the hot-rolled strip steel production process, the production line's sensing equipment can collect intermediate billet input data describing equivalent rolling process parameters and response data representing the bending state of the intermediate billet. The intermediate billet input data consists of m samples with n-dimensional features (rolling force, rolling speed, etc.), and the intermediate billet input dataset F constitutes the intermediate billet input dataset. m,n It can be represented as a matrix:
[0026]
[0027] Among them, f xy (x = 1, 2, ..., m, y = 1, 2, ..., n) represents a feature value of a sample, where x represents the sample number and y represents the feature number;
[0028] S112) If f xy (Where x∈[1,m], y∈[1,n]) The value is empty (data missing) or the value is abnormal (e.g., the characteristic corresponding to y is rolling force but f is not). xy If the value is negative, then remove the sample with number x (remove the dataset F). m,n [f] x1 ,…,f xn ]);
[0029] S113) If there exist sample data with completely identical n-dimensional feature values, i.e., [f x1 ,…,f xn ] and [f x′1 …f x′n If the input dataset F is the same, where x = 1, 2, ..., m and x′ = 1, 2, ..., m (x ≠ x′), then remove the sample numbered x′ (remove the intermediate blank input dataset F). m,n [f] x1 ,…,f xn Because the intermediate blank input dataset F after steps S112) and S113) m,n Some samples may be removed, so the processed intermediate blank input dataset is F. m′,n The number of samples is m′ (m′≤m);
[0030] S12) The intermediate blank input dataset processed in step S11) is standardized using the z-score standardization method, which mainly consists of the following steps:
[0031] S121) Calculate the mean and standard deviation of the data characteristics for the intermediate billet input dataset F. m′,n The mean and standard deviation of the y-th feature are:
[0032] mean
[0033]
[0034] Standard deviation
[0035]
[0036] Among them, f iy Let y be the feature value of the i-th sample in dimension y, where i = 1, 2, ..., m′;
[0037] S122) Input dataset F for intermediate billet m′,n Perform z-socre standardization, and input the intermediate billet dataset F m′,n Any sample f iy After z-socre standardization, it becomes:
[0038]
[0039] The intermediate billet input dataset after processing (S123) can be represented as follows:
[0040]
[0041] S13) The Savitzky-Golay filter is used to smooth the feature of the intermediate billet centerline offset in the intermediate billet response data. This mainly consists of the following steps:
[0042] S131) The centerline offset of the intermediate billet is the content of the "DATACONTENT" field in the intermediate billet response dataset, which records the centerline offset of the intermediate billet. Its form is S = [s1, s2, ..., s L ],s1,s2,...,s L s1 and s are obtained by sampling at equal intervals along the rolling direction. L s represents the sampling start point and sampling end point, respectively. l The value of (l=1,2,...,L) is the l-th offset between the center line of the intermediate billet and the center line of the rolling mill;
[0043] S132) The Savitzky-Golay filter is used to smooth the offset of the center line of the intermediate billet. The formula is as follows:
[0044]
[0045] Let s[ω] be a set of data within a sliding window of length 2m+1, where ω = -m, -m+1, ..., 0, ..., m-1, m, and ω takes 2m+1 consecutive integer values. Now, we construct an H-order (H≤2m+1) polynomial f(ω) to fit the data s[ω]:
[0046]
[0047] Among them is S is the filtering result of the l-th data point; m is the order of the Savitzky-Golay filter, i.e., the length of the window, used to specify the size of the sliding window, and must be a positive odd number, here the value is 5; s l+ω Let ω be the ω-th data point within a sliding window of length 2m+1 centered on the l-th data point in S. This data point is obtained by fitting a polynomial f(ω) to the data point (b0, b1, ..., bj). H (These are polynomial coefficients); c ω These are the filter coefficients, obtained by fitting the polynomial using the least squares method.
[0048] The offset of the centerline of the intermediate billet is S = [s1, s2, ..., s L The centerline offset after smoothing is expressed as...
[0049] S14) Calculate the camber of the intermediate billet using the smoothed centerline offset, and add it as a new feature to the intermediate billet response dataset. This mainly consists of the following steps:
[0050] S141) Defines the method for calculating the bending amount of a sickle:
[0051]
[0052] Where, bending refers to the degree of bending of the sickle. for The maximum value in the middle; and These represent the centerline offsets of the sampling start and end points of the intermediate billet after smoothing, respectively.
[0053] S142) The intermediate billet response data contains p samples with q-dimensional features (width, thickness, centerline offset, etc.), m = p, which constitute the intermediate billet response dataset G. p,q The data matrix can be used to represent:
[0054]
[0055] Among them, g x″y″Let x″ = 1, 2, ..., p, y″ = 1, 2, ..., q represent a feature value of a sample, x″ represent the sample number, and y″ represent the feature number. Adding a bending feature makes G... p,q The feature dimension is expanded to q+1 dimensions, resulting in G. p,q+1 :
[0056]
[0057] g i′(q+1) i′=1,2,…,p is offset by g from the center line of the intermediate billet. i′o (i′ is the sample number, o is the feature number, g) i′o The offset of the center line of the intermediate billet of sample i′ is calculated after smoothing, i.e., g i′o =S=[s1,s2,...,s L The smoothing process is performed using the method given in step S132) to obtain... Furthermore, the bending amount of the sickle is obtained using the method given in step S141). in express The maximum value in, These represent the centerline offsets of the sampling start and end points of the intermediate billet after smoothing, respectively.
[0058] S15) Using the common field in the intermediate billet input data and response data—the intermediate billet plate number—process the input data obtained in step S12). The response data result G obtained from step S14) p,q+1 The process of concatenating the data to obtain the dataset mainly consists of the following steps:
[0059] The intermediate billet input dataset obtained after preprocessing (S151) is:
[0060]
[0061] The intermediate billet response dataset obtained after preprocessing (S152) is as follows:
[0062]
[0063] S153) The processed intermediate billet input dataset and the processed intermediate billet response dataset are left-joined based on the common feature—the intermediate billet plate number (a unique number for each billet in strip rolling). The processed intermediate billet response dataset only includes the bending amount of the sickle bend, resulting in the dataset:
[0064]
[0065] This invention uses the bending amount of the sickle bend as the basis for generating the sickle bend control strategy. There is a one-to-one correspondence between the samples in the intermediate billet input data and the response data. Input data F m,n After preprocessing, the result is Where m′≤m and m=p, m′ is With G p,q+1 The number of samples after left join concatenation. Because left join concatenation was used, the number of samples in the dataset is less than the number of samples in the intermediate input dataset. The sample size is the same, m′ samples.
[0066] S2) Establish and train a sickle bending amount prediction model based on the XGBoost algorithm. This model can output the predicted sickle bending amount based on input rolling process parameters and other features. In this invention, the sickle bending amount prediction model can predict the sickle bending amount at the end of a certain rolling pass based on the input features (process parameters during rolling, etc.), thereby simulating the sickle bending state and building an intelligent agent environment for the sickle bending control strategy generation model. This control strategy generation model will calculate the adjustment value of the process parameters to be adjusted based on the input feature data and the sickle bending amount prediction model. The establishment and training steps of the sickle bending amount prediction model include dividing the training set and test set, and setting and optimizing the XGBoost algorithm parameters. The specific steps are as follows:
[0067] S21) Divide the dataset after step S153) into a training set and a test set. Randomly select 80% of the data in the dataset as the training set trainset1, and the remaining data as the test set testset1.
[0068] S22) The sickle bending amount prediction model is trained using the default parameters of the XGBoost algorithm. Based on the model performance metrics, namely the mean absolute error (MAE) and mean square error (MSE), the parameters of the sickle bending amount prediction model are adjusted. The model parameters that meet the prediction effect requirements (MAE≤4 and MSE≤30) are saved to obtain the trained sickle bending amount prediction model.
[0069] Step S22) mainly includes the following sub-steps:
[0070] S221) The bending amount of the sickle in trainingset1 is used as the output Y. train The remaining features in trainset1 (strip width, rolling mill force, etc.) are used as input X. trainA sickle bend amount prediction model was established based on the XGBoost algorithm and trained using the default parameters of the XGBoost algorithm to obtain the trained sickle bend amount prediction model. In this embodiment, the prediction effect of the trained sickle bend amount prediction model based on the XGBoost algorithm is as follows: Figure 2 As shown, the vertical axis represents the bending amount of the sickle curve, and the horizontal axis represents the sample (each sample has a bending_predict value and a bending_true value, the bending_predict value represents the bending predicted value, and the bending_true value represents the bending actual value).
[0071] S222) The bending amount of the sickle in testset1 is used as the true value Y of the bending amount of the sickle. test (This value is calculated from the centerline offset data of the intermediate billet collected during actual production.) The remaining features in testset1 serve as the input X to the trained sickle bend prediction model. test The predicted value Y of the sickle's bending amount is obtained. predict ;
[0072] S223) Using Y predict With Y test The mean squared error (MSE) and mean absolute error (MAE) are used to evaluate the prediction performance of the trained sickle bend prediction model. The model's prediction performance is optimized by adjusting its parameters, and model parameters that meet the prediction performance requirements (MAE ≤ 4 and MSE ≤ 30) are saved. The formulas for mean absolute error and mean squared error are:
[0073] Mean absolute error:
[0074]
[0075] Mean square error:
[0076]
[0077] Where m′ is the sample size, |·| denotes the absolute value, and y i To calculate the actual value Y of the sickle bend bending amount based on the offset of the sickle bend centerline in the actual production records. test , The predicted value Y output by the sickle bending amount prediction model predict .
[0078] S3) Establish a sickle bend control policy generation model based on the DDPG algorithm (Deep Deterministic Policy Gradient Algorithm). This model uses the sickle bend curvature prediction model trained in step S22) to simulate the sickle bend curvature and build an agent environment. The specific steps are as follows:
[0079] S31) The dataset after step S153) is re-divided into training set and test set. 80% of the data in the dataset is randomly selected as training set trainset2 and the remaining data is used as test set testset2.
[0080] S32) In this invention, the sickle bend control strategy generation model uses the intermediate billet as an intelligent agent. The agent's state consists of the intermediate billet input features and the sickle bend amount. The agent's actions are the adjustment values of process parameters. The agent's environment refers to the state of the agent that can be obtained from its current state and actions. The sickle bend control strategy generation model uses the input features and the sickle bend amount to calculate the adjustment values of the process parameters to be adjusted. It mainly consists of the following steps:
[0081] S321) The state of the intermediate billet is composed of the input features of the intermediate billet and the sickle bend amount. These features are all the features of a sample in the training set trainset2, used to construct the state of that sample. As the input to the sickle bend control strategy generation model, the state of the i-th sample can be represented as:
[0082]
[0083] It is the feature value of feature number j corresponding to sample number i, g i(q+1) The bending amount of the sickle for sample number i;
[0084] (S322) The agent's action is specifically the adjustment value of the process parameters to be adjusted (i.e., some process parameters), which serves as the output of the sickle bend control strategy generation model. The partial process parameters are selected from the intermediate billet rolling process parameters (partial features of the intermediate billet input data) to identify the process parameters that have a key impact on the sickle bend. The action of the i-th sample... i It can be represented as:
[0085]
[0086] Where 1 ≤ k < n (k is the number of process parameters to be adjusted, which are some features of the intermediate billet data; generally, k is much smaller than n), and n is the number of features in the training set trainset2. This indicates the value that feature j′ of sample number i should be adjusted to;
[0087] S323) The environment refers to the state of the intermediate billet, which can be obtained from the current state and the action taken. Since the state is composed of all the input features of the intermediate billet and the bending amount of the sickle, and the action is the adjustment value of the process parameter to be adjusted, all the input features of the next state state' can be obtained from the state and the action. The trained sickle bending amount prediction model based on the XGBoost algorithm can obtain the corresponding sickle bending amount from all the input features of the next state state'. Therefore, the intermediate billet can obtain the next state state' (the intermediate billet input features and the sickle bending amount of the next state state') from the current state and the action taken.
[0088] The reward for the S324) sickle bend control strategy generation model is an indicator of the effectiveness of the current control strategy. A larger reward value indicates a better control strategy. The reward is related to the change in the sickle bend amount of the intermediate billet and the adjustment of some process parameters. The reward is defined as:
[0089] reward = reciprocal + punish
[0090] Among them, reciprocal and punish together constitute the reward. Reciprocal is related to the change in the bending amount of the sickle, while punish is related to some process parameters.
[0091] The reciprocal in the reward is related to the bending amount of the sickle when obtaining the next state's state' from the current state and the selected action, and is specifically defined as:
[0092]
[0093] Where |·| represents taking the absolute value;
[0094] The adjustment penalty for input feature j′ j′ Defined as:
[0095] like
[0096]
[0097] like
[0098]
[0099] j′ represents the feature number of the action, j′=1,2,…,k.
[0100] The deep reinforcement learning process is divided into several rounds. Assume that we are currently in any round t of the learning process. This represents the feature value of feature j′ at round t. This represents the adjustment value of the input feature j′ to the action taken in round t. and These represent the lower and upper bounds of the feature values of the input feature j′, respectively.
[0101] Definition of punish:
[0102]
[0103] Where punish represents the total penalty at round t, and k represents the number of process parameters to be adjusted.
[0104] In this embodiment, the state of the intermediate billet is a 58-dimensional feature, including 57-dimensional input features (rolling force, rolling speed, temperature, etc.) and a 1-dimensional bending value. The action of the intermediate billet is a 5-dimensional feature (only features with a significant impact on bending are selected for adjustment). The feature importance analysis of the bending value prediction model based on the XGBoost algorithm is performed and then filtered to obtain the desired result. The feature importance analysis results are as follows: Figure 3 As shown, in this embodiment, the standardized upper and lower bounds of the 5-dimensional feature are 1 and -1, respectively.
[0105] S33) Define the Actor and Critic network structures of the Deep Deterministic Policy Gradient Algorithm (DDPG), and set parameters such as the learning rate, state dimension, action dimension, and memory bank capacity of the neural network for the sickle-shaped control policy generation model based on the DDPG algorithm. This mainly consists of the following steps:
[0106] S331) The Actor network consists of three linear layers: The first linear layer of the Actor network takes the state as its input and outputs a 256-dimensional matrix with the ReLU activation function; the second linear layer of the Actor network takes the output of the first linear layer as its input and outputs a 256-dimensional matrix with the ReLU activation function; the third linear layer of the Actor network takes the output of the second linear layer as its input and outputs an action matrix with the tanh activation function.
[0107] The S332) Critic network consists of three linear layers: the first linear layer takes the state as input and outputs a 128-dimensional matrix; the second linear layer takes the action as input and outputs a 128-dimensional matrix; the third linear layer takes the sum of the output matrices of the first and second linear layers (element-wise addition) as input and outputs a 1-dimensional matrix, with ReLU as the activation function.
[0108] S333) Set parameters such as the neural network learning rate, state dimension, action dimension, and memory bank capacity of the sickle curve control strategy generation model based on the DDPG algorithm.
[0109] S4) Training and saving of the model for the sickle-shaped bend control strategy generation model based on the DDPG algorithm: The specific steps are as follows:
[0110] S41) The DDPG algorithm has four neural networks: the real Actor network, the real Critic network, the target Actor network, and the target Critic network. These four neural networks are established and their parameters are randomly initialized (the structures of the real Actor network and the target Actor network are the Actor network structures described in step S331, and the structures of the real Critic network and the target Critic network are the Critic network structures described in S332).
[0111] S42) Training begins. Deep reinforcement learning training is divided into several rounds. Every certain number of rounds (usually 100) the state is reset. The reset state is obtained by randomly selecting a sample from the training set trainset2.
[0112] S43) The execution process for each step in each round is as follows;
[0113] S431) For any round t, obtain the state. t And calculate the action. t ;
[0114] S432) Execute the action t And obtain reward value t and the new state value t+1 ;
[0115] S433) will (state) t action t reward t ,state t+1 Stored in the experience storage area;
[0116] S434) Perform random batch sampling from the experience storage area;
[0117] S435) Calculate the loss function of the Critic network:
[0118] y t =reward t +γQ′(state t+1 ,μ′(state t+1 |θ μ′ )|θ Q′ )
[0119] y t Let reward be the expected value of the target Critic network in round t. t The state in round t is... t Select action below t The reward obtained at each round, where γ represents the reward decay coefficient, t represents the t-th round, t = 1, 2, ..., N, and N represents the total number of rounds (deep reinforcement learning training consists of many rounds, and the number of rounds is usually very large, depending on the training task; this project involves millions of rounds). μ′(state) t+1 |θ μ′ ) represents the network parameters of the target Actor network μ′ as θ. μ′ The target Actor network μ′ input state t+1 Get the action t+1 (action t+1 =μ′(state t+1 |θ μ′ )). Q′(state) t+1 ,μ′(state t+1 |θ μ′ )|θ Q′ ) represents the target Critic network Q′ (the network parameters of the target Critic network Q′ are θ). Q′ ) for state t+1 The following uses action t+1 The Critic network score is used to evaluate the merits of choosing an action in the current state; a higher score indicates that the action is more suitable.
[0120]
[0121] L represents the loss function of the real-world Critic network, which is used to update the Critic network. t represents the t-th round, t=1,2,...,N, where N represents the total number of rounds. Q(state)t action t |θ Q ) represents the real-world Critic network Q (the network parameters of the real-world Critic network Q are θ). Q ) for state t Take action t The score.
[0122] S436) Update the real Critic network based on single-step gradient decay;
[0123] S437) Update the real Actor network based on the single-step gradient gain:
[0124]
[0125] The Monte Carlo method can serve as an unbiased estimate; therefore, the above formula can be used to calculate the expected return of the real Actor network with respect to θ. μ The gradient of J. J represents the expected reward of the real-world Actor network. Express the expected return of a real-world Actor network in terms of θ. θ The gradient (the network parameters of the real Actor network μ are θ) μ ). t represents round t, t=1,2,...,N, where N represents the total number of rounds. Q(state,action|θ) Q ) represents the real-world Critic network Q (the network parameters of the real-world Critic network Q are θ). Q The score is given for the action taken in the state.
[0126] The state is represented by the symbol "state". t The action is called "action". t (action t =μ(state t ), μ(state) t ) represents the real-world Actor network μ consisting of state. t Get the action t ) under Q(state) t action t |θ Q Find the gradient with respect to action. The state is represented by the symbol "state". t Below μ(state|θ) μ Find information about θ μ The gradient, μ(state|θ) μ) represents the real-world Actor network μ (the network parameters of the real-world Actor network μ are θ). μ ) by state t Get action
[0127] S438) Update the target Actor network and the target Critic network;
[0128] S44) The training of the sickle curve control strategy generation model is completed, and the trained sickle curve control strategy generation model is obtained. In this embodiment, the model structure of the sickle curve control strategy generation model based on the DDPG algorithm is as follows: Figure 4 As shown.
[0129] S5) Input the intermediate billet data (process parameters during rolling, such as rolling force and rolling speed) into the trained sickle bend control strategy generation model, and output the intermediate billet sickle bend control strategy. The specific steps are as follows:
[0130] S51) Perform data preprocessing on the real-time acquired input data of hot roughing strip steel intermediate billet to obtain the processed input features;
[0131] S52) Input the processed input features into the trained sickle bending amount prediction model based on the XGBoost algorithm to obtain the sickle bending amount;
[0132] S53) The processed input features and the bending amount of the sickle constitute the state, which is input into the trained sickle bending control strategy generation model and outputs the action.
[0133] S54) By performing de-standardization on the action, the adjustment value of the process parameter that needs to be adjusted corresponding to the action can be obtained, which is the output of the sickle bend control strategy generation model based on the DDPG algorithm.
[0134] In this embodiment, a flowchart of the practical application of a sickle curve control strategy generation method based on deep reinforcement learning is shown below. Figure 5 As shown. The validation of the sickle-shaped control policy generation model based on the deep deterministic policy gradient algorithm using the test set test2 is as follows. The test results for all samples in the test set test2 are as follows. Figure 6 As shown, the test results for samples in test set test2 where the sickle's bending amount is greater than 5 are as follows. Figure 7 As shown, the test results for samples in test set test2 where the sickle's bending amount is greater than 10 are as follows. Figure 8 As shown in the test results, `bending_initial` represents the bending amount of the sickle before leveling, and `bending_final` represents the bending amount of the sickle after leveling. Figure 6, Figure 7 , Figure 8 As can be seen, the positive optimization rate is defined as the proportion of samples where the adjusted sickle bend is less than the unadjusted sickle bend. Therefore, for all input samples, the positive optimization rate is 63.7%, for samples with a sickle bend greater than 5, the positive optimization rate is 74%, and for samples with a sickle bend greater than 10, the positive optimization rate is 100%.
[0135] The embodiments described above are merely some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
Claims
1. A method for generating sickle-shaped control strategies based on deep reinforcement learning, characterized in that, The method includes the following steps: S1) Collect process parameters during the hot roughing strip rolling process as intermediate billet input data, and quantitative data of intermediate billet bending state as intermediate billet response data. Preprocess the intermediate billet input data and intermediate billet response data. The preprocessing is divided into five parts in sequence: abnormal data processing, standardization processing, centerline offset smoothing processing, calculation of intermediate billet sickle bending amount and data merging. S2) Establish and train a sickle bending amount prediction model based on the XGBoost algorithm. The model outputs the predicted sickle bending amount based on the characteristics of the input rolling process parameters. The sickle bending amount prediction model based on the XGBoost algorithm predicts the sickle bending amount at the end of a certain rolling pass based on the input data of a certain rolling pass, i.e. the rolling process parameters, thereby simulating the sickle bending state. The trained sickle bending amount prediction model based on the XGBoost algorithm is obtained through training. S3) Establish a sickle bend control strategy generation model based on a deep deterministic strategy gradient algorithm. This control strategy generation model uses the sickle bend bending amount prediction model to simulate the sickle bend bending amount to build an intelligent agent environment. Based on the input data and the trained sickle bend bending amount prediction model based on the XGBoost algorithm, the sickle bend control strategy generation model calculates the adjustment value of the process parameters to be adjusted. S4) Train the sickle control strategy generation model based on the deep deterministic policy gradient algorithm and save the model parameters to obtain the trained sickle control strategy generation model based on the deep deterministic policy gradient algorithm. S5) Input the input data of the intermediate billet of hot roughing strip into the trained sickle bend control strategy generation model based on the deep deterministic strategy gradient algorithm, and output the sickle bend control strategy of the intermediate billet. Step S2) specifically includes: S21) Divide the dataset into a training set and a test set. Randomly select 80% of the data in the dataset as the training set trainset1, and the remaining data as the test set testset1. S22) Train the sickle bend amount prediction model based on the XGBoost algorithm using the default parameters. Adjust the parameters of the sickle bend amount prediction model according to the model performance metrics: Mean Absolute Error (MAE) and Mean Squared Error (MSE). Save the model parameters that meet the prediction performance requirements to obtain the trained sickle bend amount prediction model based on the XGBoost algorithm. Step S22) includes the following sub-steps: S221) Adjust the bending amount of the sickle in trainingset1. Features as output The remaining features in trainset1 are used as input. A sickle bend amount prediction model based on the XGBoost algorithm was established and trained using the default parameters of the XGBoost algorithm to obtain the trained sickle bend amount prediction model based on the XGBoost algorithm. S222) The bending amount of the sickle in testset1 The feature serves as the true value of the sickle's bending amount. , The remaining features in testset1 are used as inputs to the trained sickle bend prediction model based on the XGBoost algorithm. The offset data of the intermediate billet centerline collected during the actual production process are used to calculate the offset. The predicted value of the sickle's bending amount is obtained. ; S223) Use and The mean squared error and mean absolute error are used to evaluate the prediction performance of the trained sickle bend amount prediction model based on the XGBoost algorithm. The prediction performance of the model is optimized by adjusting the model parameters of the sickle bend amount prediction model, and the model parameters that meet the prediction performance requirements are saved.
2. The method for generating sickle-shaped control strategies based on deep reinforcement learning according to claim 1, characterized in that, Step S1) specifically includes: S11) Removing data samples from the intermediate blank input data that have missing field content, duplicate content, or values exceeding the normal range mainly consists of the following steps: S111) During the hot roughing and rolling of strip steel, the sensing equipment on the production line collects intermediate billet input data and intermediate billet response data. The intermediate billet input data consists of m samples with n-dimensional features. The intermediate billet input data constitutes the intermediate billet input dataset. Represented in matrix form as follows: in, Indicates the first The first sample 3D eigenvalues Indicates the sample number. Indicates the feature number, , ; S112) If If the value is empty or the value behaves abnormally, then remove the number. The sample, i.e., the dataset removed. In ; S113) If there exist sample data with completely identical n-dimensional feature values, i.e. and Same, among which , and Then remove the number. The sample, i.e., the input dataset after removing intermediate blanks. In ; Due to the intermediate billet input dataset processed in steps S112)-S113) Some samples will be removed, so the intermediate blank input dataset after step S11) is: The sample size is , ; S12) Standardize the intermediate blank input dataset processed in step S11) using the z-score standardization method, including the following steps: S121) Calculate the mean and standard deviation of the data characteristics for the intermediate billet input dataset. No. The mean and standard deviation of the 3D features are as follows: mean Standard deviation in, Represents any number of The first sample 3D eigenvalues ; S122) Input dataset for intermediate billet Perform z-socre standardization. any sample After z-socre standardization, it becomes: S123) The intermediate billet input dataset after processing in step S122) is represented as follows: S13) The intermediate billet centerline offset feature in the intermediate billet response data is smoothed using the Savitzky-Golay filter, including the following steps: S131) The centerline offset of the intermediate billet is the content of the DATACONTENT field in the intermediate billet response dataset, which consists of the intermediate billet response data. It records the centerline offset of the intermediate billet, and the centerline offset is expressed as... , Obtained by sampling at equal intervals along the rolling direction. and These represent the sampling start point and sampling end point, respectively. , The value is the first of the intermediate billet centerline and the rolling mill centerline. One offset; S132) The Savitzky-Golay filter is used to smooth the offset of the center line of the intermediate billet. The formula is as follows: Suppose a set of data is contained within a sliding window of length 2m+1. , , The value of is 2m+1 consecutive integer values. Now, construct a... Step, polynomial To fit the data : in, yes No. Data points The filtering result; m is the order of the Savitzky-Golay filter, which is also the length of the window, used to specify the size of the sliding window, and must be a positive odd number; Therefore The Middle Data points The first sliding window centered on the first and with a length of 2m+1. There are 10 data points, which are represented by a polynomial. The fitting yielded the following results: Represents the polynomial coefficients. ; These are the Savitzky-Golay filter coefficients, obtained by fitting the polynomial using the least squares method; The offset of the centerline of the intermediate billet after smoothing is expressed as: ; S14) Calculate the camber amount of the intermediate billet using the smoothed centerline offset, and add it as a new feature to the intermediate billet response dataset, including the following steps: S141) Defines the method for calculating the bending amount of a sickle: in, The bending distance of the sickle, for The maximum value in the middle; and These represent the centerline offsets of the sampling start and end points of the intermediate billet after smoothing, respectively. S142) The intermediate billet response data consists of p samples with q-dimensional features, forming the intermediate billet response dataset. Represented using a data matrix: in, , , It represents a certain feature value of a certain sample. Represents the sample number. Represents feature number, newly added Features make The feature dimension is expanded to q+1 dimensions, resulting in : The bending amount of the sickle is obtained using the method given in step S141). , ; S15) Using the common field in the intermediate billet input data and intermediate billet response data—the intermediate billet plate number—process the intermediate billet input dataset obtained in step S12). The intermediate billet response dataset obtained from step S14) The data is then concatenated to obtain the dataset, specifically: The processed intermediate billet input dataset and the processed intermediate billet response dataset are left-joined based on a common feature—the intermediate billet plate number. The processed intermediate billet response dataset only includes the sickle bend measurement. To obtain the dataset: Bend the sickle a certain amount As the basis for generating the sickle-shaped bend control strategy, there is a one-to-one correspondence between the samples in the intermediate billet input data and the intermediate billet response data. The dataset has the following sample size: .
3. The method for generating sickle-shaped control strategies based on deep reinforcement learning according to claim 2, characterized in that, The formulas for the mean absolute error and the mean square error are as follows: Mean absolute error: Mean square error: in, For the sample size, This indicates finding the absolute value. To calculate the actual value of the sickle bend's bending amount based on the offset of the sickle bend's center line in the actual production records. , The predicted value output by the sickle bending amount prediction model .
4. The method for generating sickle-shaped control strategies based on deep reinforcement learning according to claim 3, characterized in that, Step S3) specifically includes: S31) The dataset is re-divided into training set and test set. 80% of the data in the dataset is randomly selected as training set trainset2 and the remaining data is used as test set testset2. S32) The sickle bend control strategy generation model based on the deep deterministic strategy gradient algorithm uses the intermediate billet as an agent. The agent's state consists of the intermediate billet input features and the sickle bend amount. The agent's action is the adjustment value of the process parameters. The agent's environment refers to the next state of the agent obtained from the current agent state and action. This sickle bend control strategy generation model uses the intermediate billet input features and the sickle bend amount to calculate the adjustment value of the process parameters to be adjusted, including the following steps: S321) The agent's state is composed of intermediate blank input features and sickle bend amount, serving as the input to the sickle bend control policy generation model based on the deep deterministic policy gradient algorithm. The state of the i-th sample is represented as: It is the feature value of feature number j corresponding to sample number i. , The bending amount of the sickle for sample number i; (S322) The agent's action is specifically the adjustment value of some process parameters, which serves as the output of the sickle bend control strategy generation model. These process parameters are selected from the intermediate billet rolling process parameters and are the key process parameters that have a significant impact on the sickle bend. The action of the i-th sample... Represented as: in, k is the number of process parameters to be adjusted, and n is the number of features in the training set trainset2. The features of sample number i The value that should be adjusted is... ; S323) Environment refers to the state of the intermediate billet that can be obtained based on its current state and the action taken. Because the state is composed of the intermediate billet input features and the bending amount of the sickle, and the action is the adjustment value of some process parameters, the next state is obtained from the state and the action. The intermediate blank input features, the trained sickle bend bending amount prediction model based on the XGBoost algorithm, are then processed by the next state. The intermediate blank's input features yield the corresponding sickle bend amount. Then, the intermediate blank's next state is obtained from the current state and the action taken. ; (S324) The reward of the sickle bend control strategy generation model is an indicator of the effectiveness of the current control strategy. The larger the reward value, the better the control strategy. The reward is related to the change in the sickle bend amount of the intermediate billet and the adjustment of some process parameters. The reward is defined as: in, , Together they form the reward. It is related to the change in the bending amount of the sickle, and It is related to some process parameters; The reward function, along with the current state and the selected action, determines the next state. The bending amount of the sickle Relevant, specifically defined as: in, Indicates taking the absolute value; For input features Adjustment of punishment Defined as: like : like : The deep reinforcement learning process is divided into several rounds. Assume that we are currently in any round t of the learning process. Representing features at round t eigenvalues, This represents the action taken in round t based on the input features. The adjustment value, and Representing the input features respectively The lower and upper bounds of the eigenvalues; Definition of punish: Where punish represents the total penalty at round t; S33) Define the Actor network and Critic network structures in the deep deterministic policy gradient algorithm, and set the parameters of the sickle-shaped control policy generation model based on the deep deterministic policy gradient algorithm, including the following steps: S331) The Actor network consists of three linear layers: The first linear layer of the Actor network takes the state as its input and outputs a 256-dimensional matrix with the ReLU activation function; the second linear layer of the Actor network takes the output of the first linear layer as its input and outputs a 256-dimensional matrix with the ReLU activation function; the third linear layer of the Actor network takes the output of the second linear layer as its input and outputs an action matrix with the tanh activation function. (S332) The Critic network consists of three linear layers: the first linear layer of the Critic network takes the state as input and outputs a 128-dimensional matrix; the second linear layer of the Critic network takes the action as input and outputs a 128-dimensional matrix; the third linear layer of the Critic network takes the sum of the output matrices of the first and second linear layers of the Critic network as input and outputs a 1-dimensional matrix, with ReLU as the activation function. S333) Set the parameters of the sickle curve control policy generation model based on the deep deterministic policy gradient algorithm, including the neural network learning rate, state dimension, action dimension, and memory bank capacity.
5. The method for generating sickle-shaped control strategies based on deep reinforcement learning according to claim 4, characterized in that, Step S4) specifically includes: S41) The deep deterministic policy gradient algorithm has four neural networks: the real Actor network, the real Critic network, the target Actor network, and the target Critic network. These four neural networks are established and their parameters are randomly initialized. The structure of the real Actor network and the target Actor network is the same as the Actor network described in step S331), and the structure of the real Critic network and the target Critic network is the Critic network described in step S332. S42) Start training the sickle control policy generation model based on the deep deterministic policy gradient algorithm. The deep reinforcement learning training is divided into several rounds. The state is reset every certain number of rounds. The reset state is obtained by randomly selecting a sample from the training set trainset2. S43) The execution process for each step in each round is as follows: S431) For any... Round, gain status And calculate the action ; S432) Execute the action And obtain reward points and new state value ; S433) will Stored in the experience storage area; S434) Perform random batch sampling from the experience storage area; S435) Calculate the loss function of the target Critic network: Let be the expected value of the target Critic network in round t. State in round t Select action The reward received at that time This represents the reward decay coefficient, where t represents round t. N represents the total number of rounds. Represents the target Actor network The network parameters are Target Actor Network Input status Get action , , Represents the target Critic network State The following actions are adopted The scoring, target Critic network The network parameters are The Critic network score is used to evaluate the merits of choosing an action in the current state. A higher score indicates that the action is more suitable to be chosen. L represents the loss function of the real-world Critic network, which is used to update the real-world Critic network. Represents the Q-pair state of a real-world Critic network. Take action below The scoring, the network parameters of the real Critic network Q are: ; S436) Update the real Critic network based on single-step gradient decay; S437) Update the real Actor network based on the single-step gradient gain: Using the Monte Carlo method as an unbiased estimate, the expected return of the real Actor network is calculated using the above formula. gradient, This represents the expected return of a real-world Actor network. The expected return of a real-world Actor network is expressed as... The gradient, in real-world Actor networks The network parameters are , Represents the Q-pair state of a real-world Critic network. Take action below The score; Representing a real-world Actor network From state Get action , Representing a real-world Actor network From state Get action Below Find the gradient with respect to action. Indicates the state as Below Seeking information about gradient, Representing a real-world Actor network From state Get action , ; S438) Update the target Actor network and the target Critic network; S44) The training of the sickle-shaped control policy generation model based on the deep deterministic policy gradient algorithm is completed, and the trained sickle-shaped control policy generation model based on the deep deterministic policy gradient algorithm is obtained.
6. The method for generating sickle-shaped control strategies based on deep reinforcement learning according to claim 5, characterized in that, Step S5) specifically includes: S51) Perform data preprocessing on the real-time acquired input data of hot roughing strip steel intermediate billet to obtain the processed input features; S52) Input the processed input features into the trained sickle bending amount prediction model based on the XGBoost algorithm to obtain the sickle bending amount; S53) The processed input features and the bending amount of the sickle constitute the state, which is input into the trained sickle bending control strategy generation model and outputs the action. S54) By de-standardizing the action, the adjustment value of the process parameter that needs to be adjusted corresponding to the action can be obtained, which is the output of the sickle bend control strategy generation model based on the deep deterministic strategy gradient algorithm.
7. The method for generating sickle-shaped control strategies based on deep reinforcement learning according to claim 6, characterized in that, The value of m in step S13) is 5.
8. The method for generating sickle-shaped control strategies based on deep reinforcement learning according to claim 7, characterized in that, The requirement of meeting the predicted effect means and .
9. The method for generating sickle-shaped control strategies based on deep reinforcement learning according to claim 8, characterized in that, The process parameters include rolling force and rolling speed.