Vacuum induction furnace production plan scheduling method and device based on two-stage optimization

By using a two-stage optimization model and a Q-learning-driven differential evolution algorithm, the low solution efficiency and local optimum traps of large-scale, multi-constraint, and nonlinear problems in vacuum induction furnace scheduling are solved, achieving efficient and stable production planning and scheduling, and improving smelting efficiency and production quality.

CN121745535APending Publication Date: 2026-03-27AUTOMATION RES & DESIGN INST OF METALLURGICAL IND +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-18
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing vacuum induction furnace scheduling methods are unable to solve large-scale, multi-constraint, nonlinear complex problems within an effective timeframe. They lack adaptive mechanisms, have low solution efficiency, and are prone to getting trapped in local optima, affecting smelting efficiency and product quality.

Method used

A two-stage optimization approach is adopted to construct a first allocation model that associates production tasks with furnace numbers, and a second allocation model that associates furnace numbers with slots. The Q-learning-driven differential evolution algorithm is used to solve the problem, and global optimization is achieved through a conflict task identification and reallocation mechanism.

Benefits of technology

It improves the efficiency and quality of solving production plans for vacuum induction furnaces, generates feasible and optimal global production solutions, adapts to diverse production orders, improves smelting efficiency, and reduces costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121745535A_ABST
    Figure CN121745535A_ABST
Patent Text Reader

Abstract

The invention relates to a vacuum induction furnace production plan scheduling method and device based on two-stage optimization, and the method comprises the steps: taking the load balance of multiple induction furnaces and the centralized smelting of tasks of the same steel grade as targets, and constructing a first distribution model related to production tasks and furnace numbers; constructing a second distribution model associated with the furnace number and the slot position by taking minimization of the furnace washing operation times and the furnace age deviation degree as targets; based on a Q learning driven differential evolution algorithm, solving the first distribution model and obtaining a first distribution scheme, solving the second distribution model and obtaining a second distribution scheme, and obtaining a preliminary scheduling scheme of correlation among production tasks, furnace numbers and slot positions; and determining a final scheduling scheme after global optimization through global termination judgment and local optimal detection. The method is suitable for diversified production orders and dynamic working conditions, the robustness and adaptability of the induction furnace scheduling system are improved, the smelting efficiency is remarkably improved, and the production cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of vacuum induction furnace planning and scheduling technology, and in particular to a method and apparatus for vacuum induction furnace production planning and scheduling based on two-stage optimization. Background Technology

[0002] Vacuum induction furnaces, as core equipment in special metallurgy and the preparation of high-end alloy materials, play an irreplaceable role in key fields such as aerospace, new energy, and defense technology. Through induction heating and precise temperature control in a vacuum environment, they can effectively melt reactive metals, ultra-pure alloys, and high-performance steels, making them a crucial industrial device for achieving precise control of material microstructure and efficient purification of composition. Optimizing the scheduling of the vacuum induction furnace production process directly affects melting efficiency, energy consumption levels, and final product quality, and is of great significance for improving enterprise production efficiency and market competitiveness.

[0003] However, solving the scheduling problem of vacuum induction furnaces faces numerous challenges. Traditional single-stage optimization methods struggle to reconcile the requirements of global optimization with local refinement. Optimization methods in related technologies can be broadly categorized into two types: exact algorithms and heuristic algorithms. While exact algorithms can guarantee optimality, they struggle to obtain feasible solutions within a reasonable timeframe for large-scale, multi-constraint, and nonlinear problems. Heuristic algorithms, when dealing with multi-stage, multi-objective problems, often suffer from premature convergence, parameter sensitivity, and difficulty in balancing exploration and utilization. This is particularly true for scenarios like vacuum induction furnace scheduling, where decision variables are discrete, constraints are complex, and objectives conflict. Traditional heuristic algorithms typically rely heavily on manual parameter tuning and lack adaptive mechanisms. Furthermore, existing two-stage optimization frameworks often employ sequential solutions or simple iterations, failing to achieve deep information interaction and guidance between stages. This results in low efficiency and poor stability in solving NP-hard problems, severely impacting the quality and efficiency of scheduling results. Summary of the Invention

[0004] Based on the above analysis, the present invention aims to provide a method and apparatus for scheduling production plans of vacuum induction furnaces based on two-stage optimization, in order to solve the problems that existing models are unable to handle large-scale, multi-constraint, nonlinear complex problems, lack adaptive mechanisms, and have low solution efficiency.

[0005] On one hand, embodiments of the present invention provide a two-stage optimized production planning and scheduling method for vacuum induction furnaces, including:

[0006] With the goal of balancing the load of multiple induction furnaces and concentrating the smelting of the same steel grade, a first allocation model is constructed that associates production tasks with furnace numbers.

[0007] To minimize the number of furnace cleaning operations and the deviation in furnace age, a second allocation model is constructed that associates furnace number with furnace slot.

[0008] Based on the Q-learning-driven differential evolution algorithm, the first allocation model is solved to obtain the first allocation scheme, the second allocation model is solved to obtain the second allocation scheme, and the first allocation scheme and the second allocation scheme are combined to obtain the preliminary scheduling scheme in which production tasks, furnace numbers and slots are interconnected.

[0009] Determine whether the preliminary scheduling scheme meets the global termination condition: if it does, then the preliminary scheduling scheme is used as the final scheduling scheme and output; otherwise, perform local optimum detection and determine the final scheduling scheme based on the detection result.

[0010] Furthermore, the objective function F1 of the first allocation model is obtained according to the following formula:

[0011]

[0012] in, The load balancing function represents minimizing the deviation between the operating time of each induction furnace and the average value; Let be the total dispersion function, representing the minimization of the dispersion of the same steel grade across different induction furnaces;

[0013] For the first allocation model, represents the state of whether task i is assigned to induction furnace j for execution; I is the total number of induction furnaces; J is the total number of production tasks; S is the total number of steel types.

[0014] T represents the average operating time of each induction furnace. i Standard operating time required to complete task i;

[0015] c1 is the first penalty coefficient; c2 is the second penalty coefficient; This serves as an indicator variable for steel grade affinity.

[0016] Furthermore, the solution to the first allocation model must at least satisfy the following constraints:

[0017]

[0018] in, This indicates the first task allocation constraint; This indicates that the value of task i belongs to the set of task integers.

[0019] This indicates furnace capacity limitations; Indicator coefficient; This indicates that the value of induction furnace j belongs to the set of integers for induction furnaces.

[0020] This indicates a task affinity constraint for the same steel grade; This indicates that the value of any steel grade s belongs to the set of integers for steel grades. This indicates that task i belongs to the s-th steel type.

[0021] Furthermore, the objective function F2 of the second allocation model is obtained according to the following formula:

[0022]

[0023] in, The function is the number of furnace washes. C is the furnace age function; C3 is the third penalty coefficient; C4 is the fourth penalty coefficient; K j This represents the total number of slots required for the induction furnace.

[0024] Let $\frac{ ... The first furnace age penalty variable; This is the second furnace age penalty variable.

[0025] Furthermore, the solution to the second allocation model must at least satisfy the following constraints:

[0026]

[0027] in, This indicates slot allocation constraints; This indicates the second task allocation constraint; For the second allocation model, represents the state of whether task i is assigned to slot k; This indicates that the value of any slot k belongs to the set of integer slots.

[0028] This indicates penalties and constraints for furnace cleaning operations; This indicates whether task q is assigned to slot k+1; This represents the set of tasks q that need to trigger the furnace cleaning operation after task i; This indicates that task q belongs to the set. This indicates that task q does not belong to the set.

[0029] This indicates the optimal furnace age range constraint. This indicates that task i is located at slot k. This is the lower limit of the optimal furnace life; To achieve the optimal furnace lifespan, The first furnace age penalty variable; This is the second furnace age penalty variable.

[0030] Further, solving the first allocation model to obtain the first allocation scheme includes:

[0031] Generate initial population This makes the initial population Each individual in the scheme corresponds to a candidate scheme in the first allocation scheme;

[0032] According to the predefined load balancing function and total dispersion function The initial population Individuals are classified into different state categories S(t);

[0033] Construct the first action set based on the predefined mutation operation operator, and construct the initial first state action value matrix;

[0034] Based on the state category S(t), a corresponding mutation strategy is selected for each individual from the first action set, and a test individual U is generated through a crossover operation. n (t+1) to update the initial population

[0035] According to the test individual U n The reward value is calculated based on the fitness improvement at (t+1), and the action value matrix of the first state is updated based on the reward value.

[0036] Based on the updated population The first allocation model is iteratively solved using the first state action value matrix until the first termination condition is met, at which point the first allocation scheme is output.

[0037] Further, solving the second allocation model to obtain the second allocation scheme includes:

[0038] Select the induction furnaces to be scheduled and generate the initial population for the second allocation model. This makes the initial population Each individual in the scheme corresponds to one of the candidate schemes in the second allocation scheme;

[0039] Based on a predefined furnace washing frequency function and furnace age function The initial population Individuals are classified into different state categories S(h);

[0040] Construct a second set of actions based on predefined mutation operators, and construct an initialized second-state action value matrix;

[0041] Based on the state category S(h), a corresponding mutation strategy is selected for each individual from the second action set, and a test individual U is generated through a crossover operation. n (h+1) to update the initial population

[0042] According to the test individual U n Calculate the reward value for the quality improvement of (h+1), and update the second state action value matrix based on the reward value;

[0043] Based on the updated population The second allocation model is iteratively solved using the second state action value matrix until the second termination condition is met, at which point the second allocation scheme is output.

[0044] Furthermore, the execution of local optimal detection includes:

[0045] Calculate the absolute deviation of the objective function value of the second allocation model between the current iteration and the previous iteration, and determine the detection result of the local optimal detection;

[0046] The first allocation model and the second allocation model are iteratively modified based on the detection results to obtain the final scheduling scheme.

[0047] Further, the iterative correction of the first allocation model and the second allocation model based on the detection results includes:

[0048] When the system gets stuck in a local optimum, a random perturbation is applied to the preliminary scheduling scheme to clear the allocation information, and the first allocation model and the second allocation model are iteratively corrected.

[0049] Otherwise, based on the preliminary scheduling scheme, the induction furnace with the worst performance is selected, and the conflicting task with the lowest compatibility is identified from the task sequence of the induction furnace; and after the redistribution of the conflicting task is completed, the first allocation model and the second allocation model are iteratively corrected.

[0050] On the other hand, embodiments of the present invention provide a vacuum induction furnace production planning and scheduling device based on two-stage optimization, comprising:

[0051] The first construction module is used to construct a first allocation model that associates production tasks with furnace numbers, with the goal of balancing the load of multiple induction furnaces and concentrating the smelting of the same steel grade.

[0052] The second construction module is used to construct a second allocation model that associates furnace number and slot with the goal of minimizing the number of furnace cleaning operations and furnace age deviation.

[0053] The calculation and solution module is used to solve the first allocation model and obtain the first allocation scheme based on the Q-learning driven differential evolution algorithm, solve the second allocation model and obtain the second allocation scheme, and combine the first allocation scheme and the second allocation scheme to obtain a preliminary scheduling scheme in which the production tasks, furnace numbers and slots are interconnected.

[0054] The global optimization module is used to determine whether the preliminary scheduling scheme meets the global termination condition: if it does, the preliminary scheduling scheme is used as the final scheduling scheme and output; otherwise, local optimum detection is performed, and the final scheduling scheme is determined based on the detection result.

[0055] Compared with the prior art, the present invention can achieve at least one of the following beneficial effects:

[0056] First, unlike related technologies where models struggle to handle large-scale, multi-constraint, and nonlinear complex problems, this invention aims to balance the load of multiple induction furnaces and concentrate the melting of the same steel grade, constructing a first allocation model that associates production tasks with furnace numbers; and aims to minimize the number of furnace cleaning operations and furnace age deviation, constructing a second allocation model that associates furnace numbers with slots. Thus, by integrating a two-stage optimization model of "task-furnace number-slot" with multi-objective coupled constraints, and a model construction method that satisfies physical rule constraints, the scheduling scheme is highly adapted to the actual production scenario.

[0057] Second, unlike related technologies which suffer from low solution efficiency, lack of adaptive mechanisms, and poor coordination, this invention designs a differential evolution algorithm based on Q-learning and solves a two-stage assignment model using this algorithm. This embeds the intelligent decision-making mechanism of reinforcement learning into the global search framework of differential evolution technology, improving the solution efficiency and quality of the two-stage model and giving the algorithm stronger adaptive optimization capabilities. This provides an efficient and stable intelligent solution for solving complex NP-hard problems.

[0058] Third, unlike related technologies that lack global optimization capabilities and are prone to falling into local optima traps, this invention designs a conflict task identification and dynamic reallocation mechanism to perform global termination judgment and local optimum detection on the preliminary scheduling scheme, and obtains a globally optimized final scheduling scheme. This constructs a negative feedback optimization closed loop between stages, achieving deep synergy between macro-allocation and micro-ranking, and ensuring global optimality. The scheduling method proposed in this invention can generate feasible and optimal global production schemes, significantly improving smelting efficiency and reducing production costs. It is particularly suitable for diversified production orders and dynamic operating conditions, improving the robustness and adaptability of the scheduling system.

[0059] In this invention, the above-described technical solutions can be combined with each other to achieve more preferred combinations. Other features and advantages of this invention will be set forth in the following description, and some advantages may become apparent from the description or be learned by practicing the invention. The objects and other advantages of this invention can be realized and obtained from what is particularly pointed out in the description and drawings. Attached Figure Description

[0060] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Throughout the drawings, the same reference numerals denote the same parts.

[0061] Figure 1 This is a flowchart of a two-stage optimized production planning and scheduling method for vacuum induction furnaces according to an embodiment of the present invention;

[0062] Figure 2 This is a schematic diagram of the technical route of the vacuum induction furnace production planning and scheduling method based on two-stage optimization according to an embodiment of the present invention.

[0063] Figure 3 This is a schematic diagram of the first-stage solution process according to an embodiment of the present invention;

[0064] Figure 4 This is a schematic diagram of the second-stage solution process in an embodiment of the present invention;

[0065] Figure 5 This is a schematic diagram of the main modules of the vacuum induction furnace production planning and scheduling device based on two-stage optimization according to an embodiment of the present invention. Detailed Implementation

[0066] Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, which form part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not intended to limit the scope of the present invention.

[0067] A specific embodiment of the present invention discloses a production planning and scheduling method for vacuum induction furnaces based on two-stage optimization, such as... Figure 1 As shown, the steps S1 to S4 are as follows:

[0068] Step S1: With the goal of load balancing of multiple induction furnaces and concentrated smelting of the same steel grade, a first allocation model is constructed that associates production tasks with furnace numbers.

[0069] The first allocation model of this invention, also known as the first-stage "task-furnace number" optimization model, aims to generate the current optimal "task-furnace number" allocation scheme. The first allocation model takes the balance of the total working time of multiple induction furnaces as the core optimization objective. At the same time, it prioritizes the allocation of production tasks of the same steel grade to the same induction furnace for smelting, so as to improve production efficiency and energy economy.

[0070] Specifically, refer to Figure 2 As shown, a two-stage optimization model needs to be constructed based on the production task list and induction furnace equipment parameters. The first allocation model aims to balance the load of multiple induction furnaces and concentrate the melting of tasks of the same steel grade. The second allocation model aims to minimize the number of furnace cleaning operations and the deviation of furnace age. The production task list includes at least information such as steel grade, weight, and delivery date. The induction furnace equipment parameters include at least parameters such as furnace number, furnace capacity, and furnace age. Then, the two-stage optimization models are solved by a differential evolution algorithm driven by Q-learning. The two-stage optimization models are iteratively corrected by a conflict task identification and reassignment mechanism to output a feasible and globally optimal final scheduling scheme.

[0071] To accurately describe the optimization problem in the first stage, we first define the parameters and decision variables required for the mathematical model:

[0072] make Indicates the task index. Indicates the index of the induction furnace. This indicates the steel grade category index.

[0073] Define binary decision variables This is used to characterize when task i is assigned to induction furnace j for execution. Otherwise, it is 0.

[0074] Define parameter T i : Indicates the standard operation time required to complete task i.

[0075] Define a set Let represent the set of all tasks belonging to steel type s. All I tasks can be categorized into S different steel types based on their metallurgical properties. If task i belongs to the s-th steel type, it is denoted as .

[0076] Define indicator coefficient This is used to characterize whether induction furnace j is qualified to produce task i. If task i is in the qualification set of induction furnace j... Inside, then Otherwise, it is 0.

[0077] In some implementations, the constraints of the first allocation model include at least the first task allocation constraint, the furnace capacity limitation constraint, and the task affinity constraint for the same steel grade.

[0078] The first task allocation constraint means that each production task must be assigned to one induction furnace for execution. The specific formula includes:

[0079]

[0080] Furnace capacity limitation: This refers to the fact that tasks can only be assigned to induction furnaces with production qualifications, reflecting the capacity limitations of the furnace. Specifically, this includes:

[0081] make This represents the set of tasks for which induction furnace j has production qualifications, and defines the indicator coefficient:

[0082]

[0083] For any induction furnace With any task Apply furnace volume limitation constraints:

[0084]

[0085] This constraint indicates that if If induction furnace j can be used for production task i, then this constraint is a relaxed constraint; if Then forced

[0086] Steel Grade Affinity Constraints: These constraints aim to concentrate the production of the same steel grade in the same induction furnace as much as possible. Specifically, they include:

[0087] For any induction furnace and steel grades Define the affinity indicator variable:

[0088]

[0089] If at least one task of steel grade s is assigned to induction furnace j, then Otherwise, the value is 0. To establish With decision variables The relationship between them, for any satisfying The task, steel grade, and furnace number combination are subject to the same steel grade task affinity constraint:

[0090]

[0091] Ideally, all tasks for the same steel grade s should be assigned to the same induction furnace as much as possible, i.e., it is desirable Minimize as much as possible, with an optimal value of 1. This is equivalent to maximizing the concentration of tasks related to the same type of steel.

[0092] To simultaneously optimize the workload balance of multiple induction furnaces and the concentration of steel production, the objective function of the first stage integrates two sub-objectives: workload balance and steel production concentration. The objective function of the first allocation model is set as follows:

[0093]

[0094] in, The load balancing function represents minimizing the deviation between the operating time of each induction furnace and the average value; Let be the total dispersion function, representing the minimization of the dispersion of the same steel grade across different induction furnaces;

[0095] For the first allocation model, represents the state of whether task i is assigned to induction furnace j for execution; I is the total number of induction furnaces; J is the total number of production tasks; S is the total number of steel types.

[0096] T represents the average operating time of each induction furnace. i Standard operating time required to complete task i;

[0097] c1 is the first penalty coefficient; c2 is the second penalty coefficient; Let c1∈R(0,∞) and c2∈R(0,∞) be the indicator variables for steel grade affinity. R[a,b] and R[a,b] represent the set of integers and the set of real numbers with upper and lower limits of b and a, respectively.

[0098] As can be seen from the above, the constraints of the first allocation model can be summarized in the following form:

[0099]

[0100] in, This indicates the first task allocation constraint; This indicates that the value of task i belongs to the set of task integers.

[0101] This indicates furnace capacity limitations; Indicator coefficient; This indicates that the value of induction furnace j belongs to the set of integers for induction furnaces.

[0102] This indicates a task affinity constraint for the same steel grade; This indicates that the value of any steel grade s belongs to the set of integers for steel grades. This indicates that task i belongs to the s-th steel type.

[0103] In summary, the allocation model in the first stage boils down to solving for the decision variables under the constraints described above. To minimize the objective function F1, the optimal "task-furnace number" allocation scheme can be obtained by combining formulas (6) and (7) above. That is, the first allocation scheme.

[0104] Step S2: With the goal of minimizing the number of furnace cleaning operations and the deviation of furnace age, a second allocation model is constructed to associate furnace number and slot.

[0105] The second allocation model of this invention, namely the second-stage "furnace number-slot" sorting model, aims to further optimize and generate a refined "task-furnace number-slot" scheduling scheme based on a given "task-furnace number" allocation scheme. In other words, the second-stage allocation model is a "furnace number-slot" sorting model, based on the allocation scheme output in the first stage. The task queue within each induction furnace J is optimized by sorting the slots. The core objective is to minimize the number of furnace cleanings during the production process and make the actual production furnace age of each task as close as possible to its optimal furnace age range.

[0106] During implementation, the parameters and decision variables required for the mathematical model should be defined as follows:

[0107] For induction furnace j: the total number of tasks it is assigned, i.e., the number of slots required, is denoted as K. j For each induction furnace Based on known decision variables The details of the assigned tasks can be determined.

[0108] Define binary decision variables If task i is scheduled to be executed in slot k, then Otherwise, it is 0.

[0109] Define a set If task q is scheduled after task i and requires triggering the furnace cleaning operation, then

[0110] Define parameters and and represent the upper and lower limits of the optimal furnace age range for task i, respectively.

[0111] Define auxiliary variables This indicates that if tasks i and q are adjacent and The penalty for washing the furnace that occurs at that time.

[0112] Define auxiliary variables and These represent the deviation penalty for task i being earlier or later than the optimal interval, respectively.

[0113] In some implementations, the constraints of the second allocation model include at least slot allocation constraints, second task allocation constraints, furnace washing operation penalty constraints, and optimal furnace age range constraints.

[0114] Slot allocation constraint: This means that each task must be assigned to one and only one slot. The specific formula includes:

[0115]

[0116] The second task allocation constraint means that each slot must be assigned one and only one task.

[0117]

[0118] Furnace cleaning operation penalty constraint: refers to the furnace cleaning cost constraint caused by switching between adjacent slots, specifically including:

[0119] set up Let i be the set of tasks that, if scheduled after task i, require triggering the furnace cleaning operation. For any task... and slots Define non-negative integer penalty variables The furnace washing cost is characterized when task i and task q are arranged adjacently, with the following specific constraints:

[0120]

[0121] Optimal furnace age range constraint: Define the slot position of task i as... Specifically, it includes:

[0122] For any task Define its slot position as set up Define a non-negative integer furnace age penalty variable for the optimal furnace age range of task i. and Furnace age deviation is calculated using the following constraints:

[0123]

[0124] Ideally, it should meet the following requirements: This means that the actual furnace slot is within the optimal furnace age range.

[0125] In implementation, the objective function F2 of the second stage consists of two parts: furnace cleaning penalty and furnace age deviation penalty. Its aim is to minimize unnecessary furnace cleaning times and schedule tasks within the optimal furnace age range as much as possible. In other words, it aims to minimize the total furnace cleaning cost and the total furnace age deviation cost, specifically obtained according to the following formula:

[0126]

[0127] in, The function is the number of furnace washes. c3 is the furnace age function; c4 is the third penalty coefficient; c3∈R(0,∞), c4∈R(0,∞); K j This represents the total number of slots required for the induction furnace. Let $\frac{ ... The first furnace age penalty variable; This is the second furnace age penalty variable.

[0128] As can be seen from the above, the second allocation model can be summarized in the following form:

[0129]

[0130] in, This indicates slot allocation constraints; This indicates the second task allocation constraint; For the second allocation model, represents the state of whether task i is assigned to slot k; This indicates that the value of any slot k belongs to the set of integer slots.

[0131] This indicates penalties and constraints for furnace cleaning operations; This indicates whether task q is assigned to slot k+1; This represents the set of tasks q that need to trigger the furnace cleaning operation after task i; This indicates that task q belongs to the set. This indicates that task q does not belong to the set.

[0132] This indicates the optimal furnace age range constraint. This indicates that task i is located at slot k. This is the lower limit of the optimal furnace life; To achieve the optimal furnace lifespan, The first furnace age penalty variable; This is the second furnace age penalty variable.

[0133] Step S3: Based on the Q-learning driven differential evolution algorithm, solve the first allocation model to obtain the first allocation scheme, solve the second allocation model to obtain the second allocation scheme, and combine the first allocation scheme and the second allocation scheme to obtain a preliminary scheduling scheme that correlates production tasks, furnace numbers and slots.

[0134] Each stage of the aforementioned two-stage optimization problem is an integer combinatorial optimization problem, exhibiting typical NP-hard characteristics. It generally suffers from inherent difficulties such as high solution space dimensionality, slow convergence speed, and susceptibility to local optima. Especially when the two stages are coupled, the complexity of the iterative search process increases significantly, leading to a decrease in co-optimization efficiency and severely limiting the algorithm's performance in practical steelmaking scenarios. The objective function of this problem consists of two sub-objectives. To effectively coordinate the trade-offs between these sub-objectives, this invention proposes a Q-learning-driven differential evolution algorithm, also known as the Q-DE algorithm. By introducing a reinforcement learning mechanism, it adaptively guides the population evolution direction, thereby achieving efficient and stable solutions to the two-stage co-optimization problem.

[0135] In practice, by combining formulas (12) and (13) above, the model is solved independently for each induction furnace j to obtain the optimal slot arrangement scheme. This is the second allocation scheme, which, based on the output results of the two stages, forms a complete "task-furnace number-slot" scheduling scheme. This is a preliminary scheduling plan, which provides precise guidance for the efficient and lean production of vacuum induction furnaces.

[0136] Solving the first allocation model yields the first allocation scheme, including:

[0137] Generate initial population This makes the initial population Each individual in the scheme corresponds to a candidate scheme in the first allocation scheme;

[0138] According to the predefined load balancing function and total dispersion function The initial population Individuals are classified into different state categories S(t);

[0139] Construct the first action set based on the predefined mutation operation operator, and construct the initial first state action value matrix;

[0140] Based on the state category S(t), a corresponding mutation strategy is selected for each individual from the first action set, and a test individual U is generated through a crossover operation. n (t+1) to update the initial population

[0141] According to the test individual U n The reward value is calculated based on the fitness improvement at (t+1), and the action value matrix of the first state is updated based on the reward value.

[0142] Based on the updated population The first allocation model is iteratively solved using the first state action value matrix until the first termination condition is met, at which point the first allocation scheme is output.

[0143] The Q-learning-driven differential evolution algorithm in this embodiment of the invention integrates the powerful global search capability of differential evolution algorithm with the intelligent decision-making mechanism of Q-learning. Through its inherent "perception-decision-learning" closed loop, it can achieve adaptive selection of optimization strategy, which significantly improves the convergence speed, stability and final solution quality when solving complex coupled optimization problems.

[0144] Preferably, combined with Figure 3 As shown, the Q-DE algorithm in this invention treats the solution of the first-stage optimization model as a sequential decision-making process, for decision variables being... The first stage optimization problem, the solution steps of which specifically include:

[0145] (1) Population initialization: Let t be the current iteration step, t max Let NP be the maximum number of generations, and let NP be the population size. NP binary decision vectors X1(t)∈{0,1} are randomly generated. IJ ,...,X NP (t)∈0,1} IJ , where each vector X n (t) represents a complete "task-furnace number" allocation scheme, which forms the initial population.

[0146] (2) Q-learning agent parameter settings: Initialize the first state action value matrix, i.e., the Q table, where Q∈R 3×4 All elements are initialized to zero, and reinforcement learning hyperparameters are configured for them, including the learning rate γ1∈R[0,0.2], the discount factor γ2∈R[0.7,0.9], the initial exploration rate ε∈R[0.8,1], and the exploration rate decay coefficient ε. delay ∈R[0.98,1].

[0147] (3) Individual state classification: Based on the performance of each individual in the population on the two sub-objective functions, state perception and classification are performed, and X is calculated for each individual. n The load balance function value of (t) and total dispersion function value Based on their numerical performance, the individual states S(t) are divided into the following three categories:

[0148] State S1(t): Characterization The value is relatively small (good load balance) and Individuals with relatively large values ​​(high steel grade dispersion).

[0149] State S2(t): Characterization Larger values ​​(poor load balance) Individuals with smaller values ​​(high concentration of steel grades).

[0150] State S3(t): Characterization and The values ​​are relatively balanced, meaning that the individuals do not belong to S1(t) or S2(t).

[0151] (4) Action space definition: Define a set of heterogeneous mutation strategies to form an action set. The decisions that an agent can choose specifically include the following actions:

[0152] a1: Perform the following standard difference mutation operation to explore by perturbing the difference vectors of three random individuals:

[0153]

[0154] Where μ1∈R(0,2) is the mutation factor, and this action is achieved through the difference combination of three individuals.

[0155] a2: Perform the following multi-difference mutation operation and introduce additional difference terms to enhance global exploration capabilities:

[0156]

[0157] Where μ2∈R(0,2), this action is an extension of a1, which enhances the exploration capability by introducing an additional difference term.

[0158] a3: Perform the following elite-oriented mutation operation, which drives the search towards known high-performance regions:

[0159]

[0160] in This action guides the search toward the neighborhood of the current best solution (elite solution).

[0161] a4: Perform the following hybrid mutation operation, which introduces random perturbations on the basis of elitism, balancing the learning from elite individuals with the ability to explore random perturbations:

[0162]

[0163] (5) Fitness assessment and state update: For each individual in the current population, calculate the overall objective function value F1(X). n (t) and its corresponding sub-objective function value, and update its state S(t) according to the aforementioned individual state classification rules.

[0164] (6) Mutation operation based on ε-greedy strategy: For each individual, with probability ε, from the action set In the process, a mutation strategy is randomly selected for exploration, or the action with the highest Q value corresponding to the current state is selected with probability 1-ε. The selected action is then executed to generate a mutation vector V. n (t+1), and then experimental individuals U are generated through crossover operations. n (t+1).

[0165] Calculate F1(U) n (t+1)) and perform greedy selection: if U n (t+1) is better than X n (t), then let X n (t+1)=U n (t+1), otherwise retain the original individual; that is, through a greedy selection mechanism: if F1(U n (t+1)) <F1(X n If (t) is true, then the original individual is replaced with the experimental individual; otherwise, the original individual is retained.

[0166] (7) Reward Calculation and Online Update of Q-Table: Design a reward function oriented towards optimization effect, based on the experimental individual U n (t+1) Quality Calculation Instant Reward r: If U n (t+1) is also superior to X n (t) and V n (t+1), then r = +10; if U n If (t+1) is better than one of them, then r = +5; otherwise, r = -2.

[0167] The update formula is learned based on the next state S(t+1) and Q:

[0168] Q(S(t),a(t))←Q(S(t),a(t))+γ1(r+γ2·maxQ(S(t+1),a(t))-Q(S(t),a(t)))(18)

[0169] The Q-table is updated online to learn the value of the optimal action under different states. After all individuals have completed the update, the value is calculated as ε(t+1) = ε(t)·ε. delay Decay exploration rate.

[0170] (8) Termination judgment: Determine whether the current algebraic t has reached t. max If the termination condition of the first stage (maximum number of generations) is met, the algorithm outputs the current optimal solution. The optimal "task-furnace number" allocation scheme is determined by the first allocation model; otherwise, t←t+1 is set and the fitness evaluation and state update operations are performed again to continue iterative optimization.

[0171] Therefore, the Q-learning-driven differential evolution algorithm is used to solve the first-stage optimization model, which effectively solves the inherent NP-hard characteristics, easy getting trapped in local optima and slow convergence of large-scale integer combinatorial optimization problems, and improves the adaptive ability of the solution process.

[0172] Solving the second allocation model to obtain the second allocation scheme includes:

[0173] Select the induction furnaces to be scheduled and generate the initial population for the second allocation model. This makes the initial population Each individual in the scheme corresponds to one of the candidate schemes in the second allocation scheme;

[0174] Based on a predefined furnace washing frequency function and furnace age function The initial population Individuals are classified into different state categories S(h);

[0175] Construct a second set of actions based on predefined mutation operators, and construct an initialized second-state action value matrix;

[0176] Based on the state category S(h), a corresponding mutation strategy is selected for each individual from the second action set, and a test individual U is generated through a crossover operation. n (h+1) to update the initial population

[0177] According to the test individual U n Calculate the reward value for the quality improvement of (h+1), and update the second state action value matrix based on the reward value;

[0178] Based on the updated population The second allocation model is iteratively solved using the second state action value matrix until the second termination condition is met, at which point the second allocation scheme is output.

[0179] Specifically, after obtaining the optimal "task-furnace number" allocation scheme in the first stage, the allocation model for the second stage needs to be solved to obtain the optimal "furnace number-slot" sorting scheme, which is the second allocation scheme. It should be noted that this stage also uses a differential evolution algorithm based on Q-learning to solve the problem, but the decision variables, optimization objectives, and population individual encoding methods are different from those in the first stage and need to be processed independently.

[0180] In some implementations, the second allocation model is designed for each induction furnace. Solve an independent subproblem of slot sorting, with the following decision variables: This indicates whether task i is assigned to slot k. The optimization objective of the second stage is to minimize the total number of furnace cleaning cycles. Deviation from total furnace life

[0181] Combination Figure 4 As shown, the specific process for sorting the j-th induction furnace is as follows:

[0182] (1) Induction furnace traversal and loop control: Initialize the induction furnace index j=1, and optimize each induction furnace independently in turn. If j>J, it means that all induction furnaces have been processed and the process ends; otherwise, continue to optimize the slot scheduling of the current j-th induction furnace.

[0183] (2) Population initialization: For the j-th induction furnace, let h be the iteration step number, h max This represents the maximum number of iterations. MP decision vectors are randomly generated. Constructing the initial population Each vector represents a slot sorting scheme.

[0184] (3) Individual State Perception and Classification: Based on the two sub-objective functions of the second stage, the state of each individual in the population is perceived. The number of furnace washing cycles is calculated. Furnace age function and its overall objective function F2(Y) n Based on their numerical performance, individual states S(h) are divided into the following three categories:

[0185] State S1(h): Characterization The value is relatively small (fewer furnace cleaning cycles) and Individuals with larger values ​​(high deviation in furnace age).

[0186] State S2(h): Characterization The value is relatively large (more furnace cleaning times) and Individuals with smaller values ​​(high furnace age matching degree).

[0187] State S3(h): Characterization and The values ​​are relatively balanced, meaning that the individuals do not belong to S1(h) or S2(h).

[0188] (4) Action space definition: Define the action set, and the mutation operation corresponding to each action is the same as the definition of formula (14) to formula (17) in the first stage, which will not be repeated here.

[0189] (5) Q-learning parameter settings: Initialize the second-state action value matrix, i.e., the Q-table Q∈R 3×4All elements are set to zero, and reinforcement learning hyperparameters are configured for them, including learning rate γ1∈R[0,0.2], discount factor γ2∈R[0.7,0.9], initial exploration rate ε∈R[0.8,1], and exploration rate decay coefficient ε. delay ∈R[0.98,1].

[0190] (6) Objective function evaluation: For each individual in the population Calculate its furnace washing frequency function Furnace age function and its overall objective function F2(Y) n (h)), and divide its state S(h)∈{S1(h),S2(h),S3(h)} according to the numerical value.

[0191] (7) Individual Update Operations and Population Evolution: This step uses the same decision-making framework as the first stage, but operates on a different solution space. Specifically, it includes:

[0192] Action execution: Based on an ε-greedy strategy, an action is randomly selected from AC with probability ε, or the action a(h) with the highest Q value in the current state is selected with probability 1-ε. That is, for each individual, a mutation strategy (action) is selected from the action set AC = {a1, a2, a3, a4} corresponding to its state. Based on the selected action, the parent individual Y... n (h) Generate the mutation vector V n (h+1);

[0193] Crossover and selection: Perform a crossover operation on the mutation vector to generate experimental individuals U. n (h+1), calculate its overall objective function value F2(U n (h+1)) and decides whether to replace the parent individual using a greedy selection mechanism, i.e., if U n (h+1) is better than Y n (h), then let Y n (h+1)=U n (h+1), otherwise retain the original individual.

[0194] (8) Reward Calculation and Knowledge Learning: Calculate the reward value r based on the quality improvement of the experimental individuals. If U n (h+1) is also superior to Y n (h) and V n (h+1), then r=+10; if U n If (h+1) is better than one of them, then r = +5; otherwise, r = -2. The update formula is learned based on the next state S(h+1) and Q:

[0195] Q(S(h),a(h))←Q(S(h),a(h))+γ1(r+γ2maxQ(S(h+1),a(h))-Q(S(h),a(h)))(19)

[0196] After updating the Q-table and processing all individuals, the decay rate of exploration is ε(h+1) = ε(h)·ε delay .

[0197] (9) Termination judgment and output: Determine whether the current iteration step h has been reached. max .

[0198] If the termination condition of the second stage (maximum number of generations) is met, then output the current optimal solution. The optimal slot arrangement for the j-th induction furnace is determined. Then, j ← j + 1, and the process returns to the induction furnace traversal step to process the next furnace. If the termination condition is not met, h ← h + 1 and the process returns to the individual state perception and classification step to continue iteratively optimizing the current induction furnace population.

[0199] Therefore, the Q-learning-driven differential evolution algorithm in this embodiment can autonomously perceive and learn optimization strategies applicable to different parent individuals based on the performance status of individuals in the population, achieving complete self-adaptation of strategy selection and parameter adjustment, fundamentally avoiding the tedious and experience-dependent manual parameter tuning process; its built-in reward mechanism is guided by the improvement of the objective function, making the search process have clear optimization goals and directions, effectively avoiding blind search.

[0200] Furthermore, by introducing a multi-state classification mechanism, it is possible to distinguish parent individuals with different performance characteristics and implement differentiated search strategies accordingly. This design intelligently allocates computing resources, concentrating the main computing power on search paths with higher potential, thereby significantly improving the convergence speed and solution efficiency of the algorithm while maintaining population diversity.

[0201] As can be seen from the above, this embodiment solves the two-stage model by using a differential evolution algorithm driven by Q-learning, which realizes the adaptive search of each induction furnace, achieves the best balance between reducing the number of furnace cleanings and ensuring the optimal furnace age, and then calculates a refined production scheduling sequence and determines the second allocation scheme; then, the allocation scheme of the two-stage model can be combined to output a complete global "task-furnace number-slot" scheduling scheme.

[0202] Step S4: Determine whether the preliminary scheduling scheme meets the global termination condition. If it does, the preliminary scheduling scheme is used as the final scheduling scheme and output. Otherwise, perform local optimum detection and determine the final scheduling scheme based on the detection result.

[0203] Performing local optimum detection includes: calculating the absolute deviation of the objective function value of the second allocation model between the current iteration and the previous iteration, and determining the detection result of the local optimum detection;

[0204] The first allocation model and the second allocation model are iteratively modified based on the detection results to obtain the final scheduling scheme.

[0205] When the system gets stuck in a local optimum, a random perturbation is applied to the preliminary scheduling scheme to clear the allocation information, and the first allocation model and the second allocation model are iteratively corrected.

[0206] Otherwise, based on the preliminary scheduling scheme, the induction furnace with the worst performance is selected, and the conflicting task with the lowest compatibility is identified from the task sequence of the induction furnace; and after the redistribution of the conflicting task is completed, the first allocation model and the second allocation model are iteratively corrected.

[0207] Preferably, the iterative solution of the two-stage optimization model in this embodiment of the invention includes an iterative solution framework comprising conflict identification, intelligent reallocation, and a feedback mechanism. This framework uses an outer loop to globally coordinate and correct the two-stage optimization results, ultimately outputting a feasible and high-performance global scheduling solution. The specific implementation scheme is as follows:

[0208] (1) Initialize the iteration environment: Set the global iteration counter l = 0, and initialize the historical objective function values ​​F1(0) = F2(0) = 0; define the maximum number of global iterations l. max And the minimum positive threshold φ used to determine convergence.

[0209] (2) Two-stage sequential solution and global evaluation: In each global iteration l, the following core optimization process is executed sequentially:

[0210] The first stage of solution: The Q-learning-driven differential evolution algorithm is used to solve the first allocation model and obtain the current optimal allocation scheme. The first allocation scheme can obtain the induction furnace j to which each task belongs.

[0211] The second stage of solution: Using the first stage scheme as a fixed input, for each induction furnace j, a differential evolution algorithm based on Q-learning is used to independently solve its internal "furnace number-slot" sorting model to obtain the optimal second allocation scheme for the current slot sequence. This, combined with the first allocation scheme, forms a complete global scheduling scheme. This serves as a preliminary scheduling plan.

[0212] (3) Global Termination Judgment and Local Optimality Detection: Determine whether the global iteration satisfies the global termination condition: If the maximum number of iterations l is reached... maxIf the current global preliminary scheduling scheme is output, the process ends; if it does not terminate, local optimum detection is performed, and the absolute deviation of the objective function value of the second stage of the current iteration and the previous iteration is calculated as |F2(l)-F2(l-1)|. If the deviation is less than or equal to the threshold φ, the optimization process is determined to be trapped in a local optimum; otherwise, the optimization is considered to be still progressing effectively.

[0213] (4) Dynamic perturbation and intelligent redistribution of conflicting tasks: If a local optimum is detected, a perturbation operation is initiated, and N tasks are randomly selected. err A production task, for example, its calculation formula includes Clear the furnace number and slot information currently assigned to the generation task. Then, let l = l + 1 and return to the two-stage sequential solution steps for re-iteration. This introduces randomness to help the algorithm escape the local optimum trap.

[0214] (5) Conflict task identification and extraction: If it is detected that the furnace is not trapped in a local optimum, the conflict task identification mechanism is activated to evaluate the performance of all induction furnaces after the second stage optimization and select the induction furnace j with the largest objective function value F2. old This refers to the furnace with the worst performance and the most problems; from the task sequence of this furnace, accurately identify and extract N err The conflicting task with the lowest fit, such as the task with the highest furnace cleaning cost or the most serious deviation in furnace age, is cleared from its current assignment information. err The value decreases as the iteration progresses, thus enabling adaptive adjustment from global exploration to local fine-grained search.

[0215] (6) Intelligent reallocation of conflict tasks: For the extracted set of conflict tasks with low compatibility, based on the principles of load balancing and steel grade affinity, the intelligent reallocation of each conflict task i in different candidate induction furnaces j is calculated. can On the receiving weight The specific formula is as follows:

[0216]

[0217] in, It is a candidate induction furnace. can Current cumulative working hours, For average working hours, To characterize the steel grade affinity coefficient between the candidate induction furnace and conflicting task i (for example, if there is already a task of the same steel grade in the furnace, the steel grade affinity coefficient is higher), α1 and α2 are weighting coefficients; according to the calculated receiving weight, each conflicting task is forcibly reassigned to its optimal candidate induction furnace, that is, let

[0218] (7) Iterative Feedback and Global Optimization: After the redistribution of conflicting tasks, the iteration counter l = l + 1 is updated, and the new "task-furnace number" correspondence is fed back to re-execute the two-stage iterative solution. In this way, bottlenecks and conflicts in the global scheme are gradually eliminated through multiple iterations, and the global optimal final scheduling scheme is finally approached and locked.

[0219] It is evident that by iteratively revising the two-stage scheme through conflict task identification and intelligent allocation mechanisms, deep information interaction and global optimization between stages are achieved, ultimately outputting a global scheduling scheme that meets engineering requirements.

[0220] The iterative solution mechanism in this invention breaks through the limitations of the traditional two-stage model's sequential solution and lack of feedback, significantly improving the overall performance and robustness of the scheduling system, and enhancing the quality and efficiency of the scheduling scheme.

[0221] Therefore, the embodiments of the present invention firstly construct a two-stage optimization model covering "task-furnace number" allocation and "furnace number-slot" sorting, establishing a multi-objective optimization framework that takes into account equipment capacity constraints, production process constraints, and quality constraints; secondly, it innovatively proposes a differential evolution algorithm based on Q-learning, which adaptively optimizes the algorithm search strategy through a reinforcement learning mechanism, achieving a leapfrog breakthrough in the solution process from experience dependence to intelligent autonomy; furthermore, by designing an iterative optimization mechanism based on conflict detection, a collaborative optimization paradigm with bidirectional feedback between stages is constructed, ensuring the feasibility and optimality of the global solution.

[0222] The vacuum induction furnace production planning and scheduling method provided by this invention can generate high-quality production scheduling schemes, significantly improve equipment utilization, reduce production energy consumption, and show excellent adaptability and stability to changing production environments and complex process requirements.

[0223] In another embodiment of the invention, a vacuum induction furnace production planning and scheduling device based on two-stage optimization is proposed, such as... Figure 5 As shown, it specifically includes the following modules:

[0224] The first construction module is used to construct a first allocation model that associates production tasks with furnace numbers, with the goal of balancing the load of multiple induction furnaces and concentrating the smelting of the same steel grade.

[0225] The second construction module is used to construct a second allocation model that associates furnace number and slot with the goal of minimizing the number of furnace cleaning operations and furnace age deviation.

[0226] The calculation and solution module is used to solve the first allocation model and obtain the first allocation scheme based on the Q-learning driven differential evolution algorithm, solve the second allocation model and obtain the second allocation scheme, and combine the first allocation scheme and the second allocation scheme to obtain a preliminary scheduling scheme in which the production tasks, furnace numbers and slots are interconnected.

[0227] The global optimization module is used to determine whether the preliminary scheduling scheme meets the global termination condition: if it does, the preliminary scheduling scheme is used as the final scheduling scheme and output; otherwise, local optimum detection is performed, and the final scheduling scheme is determined based on the detection result.

[0228] The above-described method and apparatus embodiments are based on the same principle, and their related aspects can be referenced from each other to achieve the same technical effect. For specific implementation processes, please refer to the foregoing embodiments, which will not be repeated here.

[0229] Those skilled in the art will understand that all or part of the processes of the methods described in the above embodiments can be implemented by a computer program instructing related hardware, and the program can be stored in a computer-readable storage medium. The computer-readable storage medium may be a disk, optical disk, read-only memory, or random access memory, etc.

[0230] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A two-stage optimization based vacuum induction furnace production planning and scheduling method, characterized in that, include: With the goal of balancing the load of multiple induction furnaces and concentrating the smelting of the same steel grade, a first allocation model is constructed that associates production tasks with furnace numbers. To minimize the number of furnace cleaning operations and the deviation in furnace age, a second allocation model is constructed that associates furnace number with furnace slot. Based on the Q-learning-driven differential evolution algorithm, the first allocation model is solved to obtain the first allocation scheme, the second allocation model is solved to obtain the second allocation scheme, and the first allocation scheme and the second allocation scheme are combined to obtain the preliminary scheduling scheme in which production tasks, furnace numbers and slots are interconnected. Determine whether the preliminary scheduling scheme meets the global termination condition: if it does, then the preliminary scheduling scheme is used as the final scheduling scheme and output; otherwise, perform local optimum detection and determine the final scheduling scheme based on the detection result.

2. The scheduling method of claim 1, wherein, The objective function F1 of the first allocation model is obtained according to the following formula: in, The load balancing function represents minimizing the deviation between the operating time of each induction furnace and the average value; Let be the total dispersion function, representing the minimization of the dispersion of the same steel grade across different induction furnaces; For the first allocation model, represents the state of whether task i is assigned to induction furnace j for execution; I is the total number of induction furnaces; J is the total number of production tasks; S is the total number of steel types. T represents the average operating time of each induction furnace. i Standard operating time required to complete task i; c1 is the first penalty coefficient; c2 is the second penalty coefficient; This serves as an indicator variable for steel grade affinity.

3. The scheduling method according to claim 2, characterized in that, The solution to the first allocation model must satisfy at least the following constraints: in, This indicates the first task allocation constraint; This indicates that the value of task i belongs to the set of task integers. This indicates furnace capacity limitations; Indicator coefficient; This indicates that the value of induction furnace j belongs to the set of integers for induction furnaces. This indicates a task affinity constraint for the same steel grade; This indicates that the value of any steel grade s belongs to the set of integers for steel grades. This indicates that task i belongs to the s-th steel type.

4. The scheduling method according to claim 1, characterized in that, The objective function F2 of the second allocation model is obtained according to the following formula: in, The function is the number of furnace washes. C is the furnace age function; C3 is the third penalty coefficient; C4 is the fourth penalty coefficient; K j This represents the total number of slots required for the induction furnace. Let $\frac{ ... The first furnace age penalty variable; This is the second furnace age penalty variable.

5. The scheduling method according to claim 4, characterized in that, The solution to the second allocation model must satisfy at least the following constraints: in, This indicates slot allocation constraints; This indicates the second task allocation constraint; For the second allocation model, represents the state of whether task i is assigned to slot k; This indicates that the value of any slot k belongs to the set of integer slots. This indicates penalties and constraints for furnace cleaning operations; This indicates whether task q is assigned to slot k+1; This represents the set of tasks q that need to trigger the furnace cleaning operation after task i; This indicates that task q belongs to the set. This indicates that task q does not belong to the set. This indicates the optimal furnace age range constraint. This indicates that task i is located at slot k. This is the lower limit of the optimal furnace life; To achieve the optimal furnace lifespan, The first furnace age penalty variable; This is the second furnace age penalty variable.

6. The scheduling method according to claim 3, characterized in that, Solving the first allocation model to obtain the first allocation scheme includes: Generate initial population This makes the initial population Each individual in the scheme corresponds to a candidate scheme in the first allocation scheme; According to the predefined load balancing function and total dispersion function The initial population Individuals are classified into different state categories S(t); Construct the first action set based on the predefined mutation operation operator, and construct the initial first state action value matrix; Based on the state category S(t), a corresponding mutation strategy is selected for each individual from the first action set, and a test individual U is generated through a crossover operation. n (t+1) to update the initial population According to the test individual U n The reward value is calculated according to the fitness improvement of (t+1), and the first state-action value matrix is updated according to the reward value; Based on the updated population The first allocation model is iteratively solved using the first state action value matrix until the first termination condition is met, at which point the first allocation scheme is output.

7. The scheduling method according to claim 5, characterized in that, Solving the second allocation model to obtain the second allocation scheme includes: Select the induction furnaces to be scheduled and generate the initial population for the second allocation model. This makes the initial population Each individual in the scheme corresponds to one of the candidate schemes in the second allocation scheme; Based on a predefined furnace washing frequency function and furnace age function The initial population Individuals are classified into different state categories S(h); Construct a second set of actions based on predefined mutation operators, and construct an initialized second-state action value matrix; Based on the state category S(h), a corresponding mutation strategy is selected for each individual from the second action set, and a test individual U is generated through a crossover operation. n (h+1) to update the initial population According to the test individual U n (h+1) and updating the second state-action value matrix according to the reward value. Based on the updated population The second allocation model is iteratively solved using the second state action value matrix until the second termination condition is met, at which point the second allocation scheme is output.

8. The scheduling method according to claim 1, characterized in that, The process of performing local optimal detection includes: Calculate the absolute deviation of the objective function value of the second allocation model between the current iteration and the previous iteration, and determine the detection result of the local optimal detection; The first allocation model and the second allocation model are iteratively modified based on the detection results to obtain the final scheduling scheme.

9. The scheduling method according to claim 8, characterized in that, The iterative correction of the first allocation model and the second allocation model based on the detection results includes: When the system gets stuck in a local optimum, a random perturbation is applied to the preliminary scheduling scheme to clear the allocation information, and the first allocation model and the second allocation model are iteratively corrected. Otherwise, based on the preliminary scheduling scheme, the induction furnace with the worst performance is selected, and the conflicting task with the lowest compatibility is identified from the task sequence of the induction furnace; and after the redistribution of the conflicting task is completed, the first allocation model and the second allocation model are iteratively corrected.

10. A production planning and scheduling device for vacuum induction furnaces based on two-stage optimization, characterized in that, include: The first construction module is used to construct a first allocation model that associates production tasks with furnace numbers, with the goal of balancing the load of multiple induction furnaces and concentrating the smelting of the same steel grade. The second construction module is used to construct a second allocation model that associates furnace number and slot with the goal of minimizing the number of furnace cleaning operations and furnace age deviation. The calculation and solution module is used to solve the first allocation model and obtain the first allocation scheme based on the Q-learning driven differential evolution algorithm, solve the second allocation model and obtain the second allocation scheme, and combine the first allocation scheme and the second allocation scheme to obtain a preliminary scheduling scheme in which the production tasks, furnace numbers and slots are interconnected. The global optimization module is used to determine whether the preliminary scheduling scheme meets the global termination condition: if it does, the preliminary scheduling scheme is used as the final scheduling scheme and output; otherwise, local optimum detection is performed, and the final scheduling scheme is determined based on the detection result.