Hot rolling production scheduling method and system based on dual DQN network and electronic equipment

By combining dual DQN reinforcement learning and NSGAIII, the model construction and multi-objective optimization problems in the small and medium-sized sample scenario of hot rolling production scheduling are solved, and the rapid and high-quality production scheduling scheme generation is achieved, and the quality of solution efficiency and reconciliation is improved.

CN119940887AActive Publication Date: 2025-05-06BEIJING METALS TECHNOLOGY LTD CO

Patent Information

Application Number
CN202510436616.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-05-06
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

Hot rolling production scheduling is a complex NP difficult combination optimization problem. The existing technology is difficult to effectively solve how to build and train models in small sample scenarios, and how to integrate multi-objective optimization strategies to achieve global optimal solutions.

Method used

Combining dual DQN reinforcement learning and NSGAIII, a hot rolled production scheduling method based on dual DQN network is proposed. By building dual DQN network, designing comprehensive reward functions, and using training sets and test sets to train and verify models to generate the optimal production scheduling solution.

Benefits of technology

Quickly obtain high-quality hot rolling production scheduling schemes under small sample data, improve the quality of solution efficiency and solutions, and can quickly solve and handle trade-offs between multiple optimization goals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940887A_ABST
    Figure CN119940887A_ABST
Patent Text Reader

Abstract

The invention provides a hot rolling production scheduling method and system based on double DQNs and electronic equipment, and relates to the technical field of production scheduling. The method comprises the following steps that a historical hot rolling production data set is collected and preprocessed; dividing the preprocessed historical hot rolling production data set into a training set and a test set according to a preset proportion; a dual DQN network is constructed, and a comprehensive reward function is designed by combining three targets of the minimum rolling unit number, the minimum specification hopping cost and the minimum unarranged slab punishment; respectively training and verifying the dual DQN network by using the training set and the test set to obtain a scheduling optimization model; inputting the slab specification information, the production demand information and the constraint condition information of the current production batch into the scheduling optimization model, and outputting an optimal action strategy by the scheduling optimization model; and solving the optimal action strategy to obtain an optimal production scheduling scheme. According to the method, a high-quality hot rolling production scheduling scheme can be quickly obtained under the condition of small sample data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of production scheduling, and in particular to a hot rolling production scheduling method, system and electronic equipment based on a dual DQN network. Background Art

[0002] As a key link in the steel production process, the main research direction of hot rolling production scheduling is to achieve intelligent and green, which has a direct impact on the quality of the final product, production efficiency and resource consumption. In the context of intensified market competition, customer needs tend to be diversified, prompting product production decisions to consider multiple goals comprehensively. In essence, hot rolling production scheduling is to determine the rolling unit to which the candidate slab belongs and to determine the order of the slabs in the unit (see Figure 1 ), is a complex NP-hard combinatorial optimization problem involving multiple conflicting objectives.

[0003] Traditional experience-based or rule-based methods mainly rely on the experience and intuition of production managers, lacking systematicity and scientificity; mathematical modeling methods are relatively complex and difficult to adapt to the real-time needs; heuristic algorithms such as genetic algorithms and simulated annealing often rely on the selection of initial solutions, and the convergence speed and stability of the algorithms are difficult to guarantee. In recent years, Deep Reinforcement Learning (DRL) has shown great potential in complex optimization problems due to its powerful decision-making ability and adaptive learning characteristics. However, directly applying DRL to hot rolling scheduling problems still faces challenges, especially how to effectively build and train models in small sample scenarios, and how to integrate multi-objective optimization strategies to achieve the global optimal solution. Summary of the invention

[0004] Purpose of the invention: The present invention combines dual DQN reinforcement learning and NSGAIII to propose a hot rolling production scheduling method, system and electronic equipment based on dual DQN network to solve the above-mentioned problems existing in the prior art.

[0005] In a first aspect of the present invention, a hot rolling production scheduling method based on a dual DQN network is proposed, and the steps are as follows:

[0006] Collect historical hot rolling production data sets, including slab specification information, production demand information, and constraint information;

[0007] Preprocessing the historical hot rolling production data set so that the slab specification information, production demand information, and constraint condition information are on the same scale;

[0008] Dividing the preprocessed historical hot rolling production data set into a training set and a test set according to a predetermined ratio;

[0009] A dual DQN network was constructed, and a comprehensive reward function was designed by combining the three objectives of minimizing the number of rolling units, minimizing the cost of specification jumps, and minimizing the penalty for not being included in the slab.

[0010] Using the training set and the test set to train and verify the dual DQN network respectively, to obtain a scheduling optimization model;

[0011] Inputting slab specification information, production demand information, and constraint information of the current production batch into the scheduling optimization model, and outputting the optimal action strategy by the scheduling optimization model;

[0012] Solve the optimal action strategy to obtain the optimal production scheduling solution.

[0013] In a further embodiment of the first aspect, the slab specification information records the width, thickness, and hardness of the slab;

[0014] The production demand information records the number of rolling batches , the number of rolling units in each rolling batch ;

[0015] The constraint condition information records the maximum and minimum specification differences between slabs.

[0016] In a further embodiment of the first aspect, the dual DQN network consists of a Q network and a target network;

[0017] The Q network and the target network both include an input layer, a hidden layer and an output layer; wherein the input layer receives the state vector of the current rolling batch ; n is the number of features; Indicates the status characteristics of the current i-th sample, including the number of rolling units, the specifications of the slabs that have been discharged, and the remaining slabs;

[0018] The hidden layer is composed of a multi-layer fully connected neural network, which is used to extract state features and calculate the expected probability Q of taking each possible action in the current state. The output layer outputs the action space A, which has actions, the output layer is neurons, output vector .

[0019] In a further embodiment of the first aspect, the expression of the comprehensive reward function R is as follows:

[0020]

[0021] In the formula, is the number of rolling units; is the number of slabs that have been discharged; is the weight coefficient, non-negative and ;

[0022] is the cost of specification jump between the i-th and i+1-th slabs:

[0023]

[0024] In the formula, and , and , and are the width, thickness and hardness jump costs between the i-th and i+1-th slabs respectively; is the weight of the three specification attributes of width, thickness and hardness, which is non-negative and ;

[0025] The total penalty for not entering the slab is:

[0026]

[0027] In the formula, is the number of slabs not discharged; is the penalty value for a single slab not included in the slab.

[0028] In a further embodiment of the first aspect, the weight parameters of the Q network and the target network are trained using the weighted TD error, and the calculation formula is as follows:

[0029]

[0030]

[0031] In the formula, is the weighted TD error of the sample; , Represent the weight parameters of the Q network and the target network respectively; s, are the current state and the next state respectively; a, are the current action and the next possible action respectively; r is the reward obtained after executing action a from state s; A is the action space; is the discount factor; D is the experience replay pool; is the predicted value of the Q network; is the predicted value of the target network; is a hyperparameter used to amplify the weight of the state-action pair that receives the reward.

[0032] In a further embodiment of the first aspect, the loss function of the dual DQN network is as follows:

[0033]

[0034] In the formula, is the number of samples for each training batch; is the current action of sample i; is the next possible action of sample i; , Represent the weight parameters of the Q network and the target network respectively; is the predicted value of Q network for sample i; For sample i in action The maximum predicted value at ; From the state characteristics Execute an action Rewards received after is the importance sampling weight of sample i; A is the action space; is the discount factor.

[0035] In a further embodiment of the first aspect, solving the optimal action strategy to obtain an optimal production scheduling solution specifically includes:

[0036] Using the trained dual DQN network, the action a is selected according to the current state s, and the rolling batch is gradually constructed until all slabs are arranged or the preset termination condition is reached. The generated strategy is used as the initial population of NSGA-III; each solution represents a rolling batch plan, which contains a series of slab arrangement orders and the number of rolling units;

[0037] The optimization objectives are to minimize the number of rolling units, minimize the cost of specification jumps, and minimize the penalty for not arranging slabs. The comprehensive reward function R is used as the fitness function. According to the standard deviation of the distribution in each dimension of the target space, Dynamically adjust the distribution of reference points.

[0038] In a further embodiment of the first aspect, the adaptive function is designed Dynamically adjust the distribution of reference points, adaptive function The expression is as follows:

[0039]

[0040]

[0041] In the formula, is the standard deviation of the population distribution in the i-th target dimension; is the value of the jth individual on the i-th target dimension; is the mean value on the i-th target dimension; N is the number of individuals in the population; , , are the mean, maximum and minimum values ​​of the standard deviation of the distribution on all target dimensions respectively; k is the adjustment coefficient.

[0042] A second aspect of the present invention provides a hot rolling production scheduling optimization system, the system comprising:

[0043] The acquisition module is used to collect historical hot rolling production data sets, including slab specification information, production demand information, and constraint information;

[0044] A first data processing module is used to pre-process the historical hot rolling production data set so that the slab specification information, production demand information, and constraint information are on the same scale;

[0045] A second data processing module is used to divide the preprocessed historical hot rolling production data set into a training set and a test set according to a predetermined ratio;

[0046] The first building block is used to construct a dual DQN network, combining the three objectives of minimizing the number of rolling units, minimizing the cost of specification jumps, and minimizing the penalty for not being included in the slab, to design a comprehensive reward function;

[0047] The second building module is used to respectively train and verify the dual DQN network using a training set and a test set to obtain a scheduling optimization model;

[0048] A first output module is used to input the slab specification information, production demand information, and constraint condition information of the current production batch into the scheduling optimization model, and the scheduling optimization model outputs the optimal action strategy;

[0049] The second output module is used to solve the optimal action strategy and obtain the optimal production scheduling solution.

[0050] According to a third aspect of the present invention, an electronic device is proposed, comprising: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the hot rolling production scheduling method based on a dual DQN network as described in the first aspect and its further embodiments is implemented.

[0051] Compared with the prior art, the present invention has at least the following beneficial effects:

[0052] (1) Through reinforcement learning, dual DQN can learn better strategies to generate initial production batch solutions with limited training samples, providing a good starting point for NSGAIII optimization solution.

[0053] (2) Based on the initial solution generated by dual DQN, NSGAIII can further search for a better solution, thereby improving the solution efficiency and solution quality.

[0054] (3) It combines the advantages of reinforcement learning and multi-objective optimization, and can both quickly solve the problem and handle the trade-offs between multiple optimization objectives.

[0055] (4) It can provide a novel solution method for other similar production scheduling problems and has broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work, among which:

[0057] Figure 1 This is a schematic diagram of hot rolling production scheduling.

[0058] Figure 2 The flowchart of the algorithm proposed in the present invention. DETAILED DESCRIPTION

[0059] In the following description, a large number of specific details are provided to provide a more thorough understanding of the present invention. However, it is apparent to those skilled in the art that the present invention can be implemented without one or more of these details. In other examples, in order to avoid confusion with the present invention, some technical features known in the art are not described.

[0060] The embodiment of the present invention combines dual DQN reinforcement learning and NSGAIII to propose a novel hot rolling production scheduling method based on dual DQN network, so that a high-quality hot rolling production scheduling plan can be quickly obtained under small sample data conditions.

[0061] Firstly, a five-layer dual DQN network is constructed, and the state space is defined as a vector space containing the current number of rolling units, the specifications (width, thickness, hardness) of the slabs that have been arranged and the remaining slabs, and the action space is a discrete space of the expected probability of taking each possible action in the current state; secondly, with the goal of minimizing the number of rolling units, minimizing the cost of specification jumps, and minimizing the penalty for not arranging slabs, a comprehensive reward function for dual DQN training is designed in a weighted manner; then, the dual DQN network is trained and an adaptive -A greedy strategy is used to select actions. A weighted TD error (Temporal Difference Error) is designed based on the introduction of importance sampling weights based on reward sparsity. Samples are extracted from the experience replay pool according to sample priority. After calculating the target Q value, a new loss function is designed in combination with the importance sampling weights. The network weights are optimized and updated. After training, a strategy for selecting the optimal action (i.e., production scheduling decision) under different states is obtained. Finally, the trained DQN network is used to generate the initial solution for hot rolling production scheduling. An adaptive function is designed to dynamically adjust the reference point to improve NSGAIII and iteratively optimize the initial solution to obtain the optimal production scheduling solution.

[0062] For detailed process, see Figure 2 As shown, the steps are as follows:

[0063] (1) Step 1: Collect historical hot rolling production data, including slab specifications (width, thickness, hardness), production demand information (number of rolling batches (denoted as ), the number of rolling units in each batch (denoted as ), constraint information (maximum and minimum specification differences between slabs), the number of rolling units (valued as ), the specifications (width, thickness, hardness) data of the slabs that have been arranged and the slabs that have not been arranged are taken as the data set.

[0064] (2) Step 2: Data preprocessing of the collected slab specification information, production demand information and constraint information. ( is the value of the jth feature of the i-th sample after filling, is the value of the jth characteristic data of the kth sample, N is the number of non-missing samples) for mean filling; the slab specification information is filled using ( is the value of the jth feature data of the i-th sample, , are the mean and variance of the jth feature respectively) to remove outliers; for all features, ( After standardization The value of , Represent the maximum and minimum values ​​of the j-th feature respectively) to standardize the data of different dimensions so that they are on the same scale.

[0065] (3) Step 3: Divide the processed dataset into a training set and a test set in a ratio of 80% and 20%, respectively, for training and verifying the generalization performance of the dual DQN.

[0066] (4) Step 4: Construct a dual DQN based on the input feature dimension and the output action space size. The dual DQN consists of a Q network and a target network, which have the same structure but different weights. The Q network is used to estimate the value of the current state-action pair, and the target network is used to stabilize the training process and regularly copy parameters from the Q network.

[0067] First, both the Q network and the target network consist of an input layer, three hidden layers, and an output layer. The input layer receives the state vector of the current rolling batch. , where n is the number of features, Indicates the number of rolling units, the size of the slabs that have been placed, and the remaining slabs contained in the current i-th sample. The hidden layer consists of a multi-layer fully connected neural network, which is used to extract state features and calculate the Q value. Design three hidden layers, each with m neurons, then the outputs of the first, second, and third hidden layers are expressed as , and ,in is the ReLU activation function, and , and , and are the weights and biases of the first, second, and third hidden layers, respectively. The output layer activation function is softmax, and the output is an action value vector, which represents the expected probability (Q value) of taking each possible action (corresponding to different possible slab selections or ending the current rolling unit) in the current state. Action space Assume that the discrete space is actions, the output layer is neurons, output vector .

[0068] Secondly, the DQN hyperparameters are initialized and configured as shown in Table 1. Each of the three hidden layers contains 256 neurons and uses the ReLU activation function.

[0069] Table 1 DQN parameter initialization configuration

[0070] (5) Step 5: Design a comprehensive reward function by combining the three objectives of minimizing the number of rolling units, minimizing the cost of specification jumps, and minimizing the penalty for not being included in the slab. Specifically, a negative reward is given for each additional rolling unit, and rewards or penalties are given according to the specification differences between adjacent slabs. Negative rewards are given for slabs that are not included in any rolling unit. The overall reward function R is designed as follows:

[0071]

[0072]

[0073]

[0074] in, is the number of rolling units; is the number of slabs that have been discharged; is the total penalty for not being discharged into the slab; is the number of slabs not discharged; is the penalty value for a single slab not discharged; is the weight coefficient; non-negative and ; is the specification jump cost between the i-th and i+1-th slabs; and , and , and are the width, thickness and hardness jump costs between the i-th and i+1-th slabs, is the weight of the three specification attributes of width, thickness and hardness, which is non-negative and .

[0075] (6) Step 6: Train the dual DQN network to obtain the strategy for selecting the optimal action under different states.

[0076] Specifically, first, the He initialization method is used to change the weight values ​​of the Q network and the target network from the normal distribution where n is the number of neurons in the previous layer, Initialized to 10000, Initialized to 10000.

[0077] At the same time, in order to balance exploration and exploitation, adaptive - Greedy strategy to choose actions:

[0078]

[0079] in, is the probability of exploring the strategy, , They are The initial maximum and minimum values ​​of are 1.0 and 0.1 respectively; is the decay rate, control The speed at which the value decreases is taken as 0.1; t is the current time step; T is the total number of training rounds.

[0080] Importance sampling weights are introduced to improve experience replay to correct sampling bias, and a target network mechanism is used to stabilize network training. A weight factor is introduced based on reward sparsity Design a weighted TD error method to give higher weights to state-action pairs that receive rewards in an environment with sparse rewards, so as to accelerate the learning process and encourage the exploration of state-action pairs that may lead to high rewards.

[0081] The weighted TD error is calculated as follows:

[0082]

[0083]

[0084] in, is the weighted TD error of the sample; , Represent the weight parameters of the Q network and the target network respectively; s, are the current state and the next state respectively; a, are the current action and the next possible action respectively; r is the reward obtained after executing action a from state s; A is the action space; is the discount factor; D is the experience replay pool; is the predicted value of the Q network; is the predicted value of the target network; is a hyperparameter used to amplify the weight of the state-action pair that receives the reward. , , is the total number of state-action pairs, for The number of state-action pairs that receive non-zero rewards in , k is a tuning parameter used to control Sensitivity with changes in S, where k is 1.5.

[0085] Then, according to Update the priority of each sample ( is a small positive number used to avoid the situation where the priority is 0, here it is taken as 0.001), select a larger p value and extract a batch of samples from the experience playback buffer according to the sample importance each time (here it is taken as ), periodically (taken as every 5000 steps) copy parameters from the Q network to the target function as a parameter The value of .

[0086] The importance of the i-th sample is calculated as:

[0087]

[0088] in, is the importance sampling weight of sample i; is the experience replay pool size; is an adjustment parameter used to control the weight. When , all samples have the same weight; when , high-priority samples will get greater weights, where Take it as 0.1.

[0089] DQN usually uses mean square error as the loss function. This paper designs a new loss function by combining importance sampling weights:

[0090]

[0091] in, is the number of samples for each training batch; is the current action of sample i; is the next possible action of sample i; , Represent the weight parameters of the Q network and the target network respectively; is the predicted value of Q network for sample i; For sample i in action The maximum predicted value at ; From the state characteristics Execute an action Rewards received after is the importance sampling weight of sample i; A is the action space; is the discount factor.

[0092] Finally, set the maximum number of training rounds or observe the changing trend of the reward function. When the preset maximum number of rounds is reached or the loss function tends to be stable, the model converges, stops training, and learns the strategy of selecting the optimal action (i.e., production scheduling decision) under different states.

[0093] (7) Step 7: Based on step 6, improve NSGAIII to optimize and solve the optimal rolling batch plan.

[0094] First, using the trained dual DQN network, the action a is selected according to the current state s, and the rolling batch is gradually constructed until all slabs are arranged or a certain termination condition (such as the maximum number of rolling units) is reached. The generated strategy is used as the initial population of NSGA-III. Each solution represents a rolling batch plan, which contains a series of slab arrangement orders and the number of rolling units.

[0095] Secondly, the parameters of NSGAIII are initialized as shown in Table 2.

[0096] Table 2 NSGAIII parameter initialization configuration

[0097] Then, the NSGAIII optimization is improved to solve the optimal production scheduling solution. The optimization objectives are to minimize the number of rolling units, minimize the cost of specification jumps, and minimize the penalty for not arranging slabs, and the total reward function R in step 5 is used as the fitness function. According to the distribution standard deviation of each dimension of the target space To adjust the distribution of reference points to guide the population to explore the Pareto frontier more effectively.

[0098] Designing adaptive functions Dynamically adjust the reference point:

[0099]

[0100]

[0101] in, is the standard deviation of the population distribution in the i-th target dimension; is the value of the jth individual on the i-th target dimension; is the mean value on the i-th target dimension; N is the number of individuals in the population; , , are the mean, maximum, and minimum values ​​of the standard deviation of the distribution in all target dimensions, respectively; k is the adjustment coefficient, which is used to control the scaling and direction of the adaptive function output. k usually takes a positive value to indicate that more reference points are added when the standard deviation is large. Here it is taken as 1.5.

[0102] At the same time, non-dominated sorting is performed based on the dominance relationship between individuals. Assume that the population P has each individual , for each individual p, initialize its dominance set , dominated set ; For each pair of individuals p, q, if p dominates q, then , if q dominates p, then , finally according to The values ​​are sorted hierarchically.

[0103] The individual crowding degree is calculated as:

[0104]

[0105] in, is the crowding degree of the ith individual, and are the values ​​of individuals i+1 and i-1 on the jth target, and are the maximum and minimum values ​​on the j-th target.

[0106] Next, the present invention uses real number coding to perform crossover and mutation operations on the individuals in the population to obtain the next generation population. For parent individuals and Perform crossover operation to obtain individual , , the operation is as follows:

[0107]

[0108]

[0109] The probability of mutation Perform mutation operation on the parent individual Parent to obtain the individual Child as follows:

[0110]

[0111] in , are the upper and lower bounds of the gene respectively.

[0112] After executing crossover and mutation, the next generation population is obtained.

[0113] When the termination condition is reached, a set of Pareto optimal solutions is output, that is, the optimal rolling batch configuration, including the number of rolling units, slab composition and sorting.

[0114] Based on the technical process of the hot rolling production scheduling method based on the dual DQN network disclosed in the above embodiment, a hot rolling production scheduling optimization system can be developed. The system can be composed of an acquisition module, a first data processing module, a second data processing module, a first building module, a second building module, a first output module, and a second output module.

[0115] The acquisition module is used to collect historical hot rolling production data sets, including slab specification information, production demand information, and constraint information. The first data processing module is used to preprocess the historical hot rolling production data sets so that the slab specification information, production demand information, and constraint information are on the same scale. The second data processing module is used to divide the preprocessed historical hot rolling production data sets into training sets and test sets according to a predetermined ratio. The first construction module is used to construct a dual DQN network, and design a comprehensive reward function by combining the three objectives of minimizing the number of rolling units, minimizing the cost of specification jumps, and minimizing the penalty for not including slabs. The second construction module is used to train and verify the dual DQN network using the training set and the test set, respectively, to obtain a scheduling optimization model. The first output module is used to input the slab specification information, production demand information, and constraint information under the current production batch into the scheduling optimization model, and the scheduling optimization model outputs the optimal action strategy. The second output module is used to solve the optimal action strategy to obtain the optimal production scheduling plan.

[0116] The technical process of the hot rolling production scheduling method based on the dual DQN network disclosed in the above embodiment can be implemented in whole or in part through software, hardware, firmware or any other combination.

[0117] When implemented in hardware, the above embodiments can run all or part of the working logic and calculation process on an electronic device after being compiled by software. The electronic device includes a processor, a memory, a communication interface and a communication bus. The processor, the memory and the communication interface communicate with each other through the communication bus. The memory is used to store at least one executable instruction, and the executable instruction enables the processor to execute the technical process of the hot rolling production scheduling method based on the dual DQN network disclosed in the above embodiments.

[0118] When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. If the above method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application can be essentially or partly embodied in the form of a software product that contributes to the relevant technology. The software product is stored in a storage medium, including several instructions to enable an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a disk or an optical disk. In this way, the embodiment of the present application is not limited to any specific hardware, software or firmware, or any combination of hardware, software, and firmware.

[0119] As described above, although the present invention has been shown and described with reference to specific preferred embodiments, it should not be construed as limiting the present invention itself. Various changes in form and details may be made without departing from the spirit and scope of the present invention as defined in the appended claims.

Claims

1. A hot rolling production scheduling method based on a dual DQN network, characterized in that: The steps include: Collect historical hot rolling production data sets, including slab specification information, production demand information, and constraint information; Preprocessing the historical hot rolling production data set so that the slab specification information, production demand information, and constraint condition information are on the same scale; Dividing the preprocessed historical hot rolling production data set into a training set and a test set according to a predetermined ratio; A dual DQN network was constructed, and a comprehensive reward function was designed by combining the three objectives of minimizing the number of rolling units, minimizing the cost of specification jumps, and minimizing the penalty for not being included in the slab. Using the training set and the test set to train and verify the dual DQN network respectively, to obtain a scheduling optimization model; Inputting slab specification information, production demand information, and constraint information of the current production batch into the scheduling optimization model, and outputting the optimal action strategy by the scheduling optimization model; Solve the optimal action strategy to obtain the optimal production scheduling solution.

2. The hot rolling production scheduling method based on the dual DQN network according to claim 1 is characterized in that: The slab specification information records the width, thickness and hardness of the slab; The production demand information records the number of rolling batches , the number of rolling units in each rolling batch ; The constraint condition information records the maximum and minimum specification differences between slabs.

3. The hot rolling production scheduling method based on the dual DQN network according to claim 1 is characterized in that: The dual DQN network consists of a Q network and a target network; The Q network and the target network both include an input layer, a hidden layer and an output layer; wherein the input layer receives the state vector of the current rolling batch ; n is the number of features; Indicates the status characteristics of the current i-th sample, including the number of rolling units, the specifications of the slabs that have been discharged, and the remaining slabs; The hidden layer is composed of a multi-layer fully connected neural network, which is used to extract state features and calculate the expected probability Q of taking each possible action in the current state. The output layer outputs the action space A, which has actions, the output layer is neurons, output vector .

4. The hot rolling production scheduling method based on the dual DQN network according to claim 1 is characterized in that: The expression of the comprehensive reward function R is as follows: ; In the formula, is the number of rolling units; is the number of slabs that have been discharged; is the weight coefficient, non-negative and ; is the cost of specification jump between the i-th and i+1-th slabs: ; In the formula, and , and , and are the width, thickness and hardness jump costs between the i-th and i+1-th slabs respectively; is the weight of the three specification attributes of width, thickness and hardness, which is non-negative and ; The total penalty for not entering the slab is: ; In the formula, is the number of slabs not discharged; is the penalty value for a single slab not included in the slab.

5. The hot rolling production scheduling method based on the dual DQN network according to claim 3 is characterized in that: The weighted TD error is used to train the weight parameters of the Q network and the target network. The calculation formula is as follows: ; ; In the formula, is the weighted TD error of the sample; , Represent the weight parameters of the Q network and the target network respectively; s, are the current state and the next state respectively; a, are the current action and the next possible action respectively; r is the reward obtained after executing action a from state s; A is the action space; is the discount factor; D is the experience replay pool; is the predicted value of the Q network; is the predicted value of the target network; is a hyperparameter used to amplify the weight of the state-action pair that receives the reward.

6. The hot rolling production scheduling method based on the dual DQN network according to claim 5 is characterized in that: The loss function of the dual DQN network is as follows: ; In the formula, is the number of samples for each training batch; is the current action of sample i; is the next possible action of sample i; , Represent the weight parameters of the Q network and the target network respectively; is the predicted value of Q network for sample i; For sample i in action The maximum predicted value at ; From the state characteristics Execute an action Rewards received after is the importance sampling weight of sample i; A is the action space; is the discount factor.

7. The hot rolling production scheduling method based on the dual DQN network according to claim 4 is characterized in that: Solving the optimal action strategy to obtain the optimal production scheduling solution specifically includes: Using the trained dual DQN network, the action a is selected according to the current state s, and the rolling batch is gradually constructed until all slabs are arranged or the preset termination condition is reached. The generated strategy is used as the initial population of NSGA-III; each solution represents a rolling batch plan, which contains a series of slab arrangement orders and the number of rolling units; The optimization objectives are to minimize the number of rolling units, minimize the cost of specification jumps, and minimize the penalty for not arranging slabs. The comprehensive reward function R is used as the fitness function. According to the standard deviation of the distribution in each dimension of the target space, Dynamically adjust the distribution of reference points.

8. The hot rolling production scheduling method based on the dual DQN network according to claim 7 is characterized in that: Designing adaptive functions Dynamically adjust the distribution of reference points, adaptive function The expression is as follows: ; ; In the formula, is the standard deviation of the population distribution in the i-th target dimension; is the value of the jth individual on the i-th target dimension; is the mean value on the i-th target dimension; N is the number of individuals in the population; , , are the mean, maximum and minimum values ​​of the standard deviation of the distribution on all target dimensions respectively; k is the adjustment coefficient.

9. A hot rolling production scheduling optimization system, characterized in that: include: The acquisition module is used to collect historical hot rolling production data sets, including slab specification information, production demand information, and constraint information; A first data processing module is used to pre-process the historical hot rolling production data set so that the slab specification information, production demand information, and constraint information are on the same scale; A second data processing module is used to divide the preprocessed historical hot rolling production data set into a training set and a test set according to a predetermined ratio; The first building block is used to construct a dual DQN network, combining the three objectives of minimizing the number of rolling units, minimizing the cost of specification jumps, and minimizing the penalty for not being included in the slab, to design a comprehensive reward function; The second building module is used to respectively train and verify the dual DQN network using a training set and a test set to obtain a scheduling optimization model; A first output module is used to input the slab specification information, production demand information, and constraint condition information of the current production batch into the scheduling optimization model, and the scheduling optimization model outputs the optimal action strategy; The second output module is used to solve the optimal action strategy and obtain the optimal production scheduling solution.

10. An electronic device, characterized in that: The device comprises: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the hot rolling production scheduling method based on the dual DQN network as described in any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Multi-target dynamic flexible job shop scheduling method about energy consumption based on deep reinforcement learning

    CN116644902A

  • Steel rolling process power demand response joint optimization method and system

    CN116914729A

  • Wide and thick plate integrated feeding plan optimization method and system based on NSGA-III

    CN117540636A

  • Internet of Things acquisition platform computing resource scheduling method based on reinforcement learning

    CN117687791A

  • Fixed-point area path optimization method and system for response type bus service

    CN119090108A

Cited By

  • Process optimization method based on MES data analysis

    CN121526040A

  • A process optimization method based on MES data analysis

    CN121526040B