Production planning and scheduling integrated model optimization method and system based on double-layer game model

Through the two-layer game model and the evolutionary algorithm of Markov chain learning, the optimization of production planning and scheduling is solved, and the balance of order drag and production costs is achieved, and the dual optimization of production efficiency and cost is achieved.

CN120373716APending Publication Date: 2025-07-25KUNMING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510411265.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively balance order delays and production costs, resulting in low production efficiency and increased costs, affecting the competitiveness of enterprises.

Method used

The production planning and scheduling integration model based on the two-layer game model is adopted, and the objective function is constructed to minimize the total order drag period and total production cost, and the two-layer evolution algorithm of Starkelberg game and Markov chain learning and heuristic rules are solved.

Benefits of technology

Generate high-quality solutions within a reasonable calculation time, reduce order delays and reduce production costs, provide effective decision-making support for enterprises, and improve production efficiency and customer satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373716A_ABST
    Figure CN120373716A_ABST
Patent Text Reader

Abstract

The invention discloses a production plan and scheduling integrated model optimization method and system based on a double-layer game model, and belongs to the technical field of production plan and scheduling integrated optimization. The method comprises the following steps: constructing upper and lower layer objective functions; forming a double-layer production plan and scheduling integrated model through upper and lower layer objective functions, and setting constraint conditions; and S2, constructing a double-layer evolutionary algorithm based on Markov chain learning and a fusion heuristic rule, wherein the double-layer evolutionary algorithm is used for solving the production plan and scheduling integrated model in the S1 and verifying the effectiveness of the method. The system comprises an establishment module and an optimization module. According to the integrated optimization method, a new solution thought is provided for the integrated optimization problem of production planning and scheduling through the production planning and scheduling integrated model, especially in the aspects of reducing the order tardiness and reducing the production cost, a group of solutions can be generated within reasonable calculation time, and the efficiency of the integrated optimization method is improved. Therefore, decision makers can select according to their own demands, and powerful support and guidance are provided for production plan decisions of enterprises.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of integrated optimization of production planning and scheduling, and particularly relates to an optimization method and system for an integrated model of production planning and scheduling based on a two-layer game model. Background Art

[0002] Order delay and production cost have always been important issues faced by manufacturing enterprises. Enterprises need to reduce order delay and production cost while ensuring product quality and production efficiency, so as to improve operational efficiency and reduce costs. Minimizing costs and increasing efficiency as much as possible to cope with the complex and rapidly changing global market environment has always been the core strategy for major enterprises to pursue sustainable development. With continuous technological innovation and increasing market competition, the problems of low production efficiency and high cost have become more prominent. The integrated decision-making of production planning and scheduling is crucial for the stable operation of enterprises. Once the balance between production planning and scheduling decisions is disrupted, with the plan lacking scheduling constraints and the scheduling unable to keep up with the production plan, it will not only affect production efficiency but also may lead to problems such as increased costs, delivery delays, and decreased customer satisfaction, ultimately weakening the competitiveness of the enterprise. Therefore, enterprises need to continuously strengthen the research on integrated decision-making to ensure its stability and reliability. Only by ensuring the continuous stability of integrated decision-making can enterprises reduce production costs, improve production efficiency, and thus achieve the goals of enhancing customer satisfaction and maximizing profits.

[0003] In modern manufacturing, as production planning management is increasingly affected by market dynamics, enterprises face many challenges in responding to changes and often struggle to provide timely feedback. The effectiveness of integrated decision-making in production planning and scheduling has become a key means to enhance the market competitiveness and sustainable development of enterprises. However, in the actual production process, the plan usually arranges tasks and schedules in advance within a specified period, while the scheduling assigns a set of processing tasks to available equipment within a certain period to achieve specific performance goals. In current decision-making models, a top-down structure is mostly adopted. However, some common measures to reduce costs and pursue profits often have an adverse impact on the stability of enterprise development. By comprehensively considering order delay and production cost and integrating the decision optimization of the two levels of planning and scheduling, not only can production costs be reduced, production efficiency be improved, but also customer satisfaction can be enhanced, promoting the sustainable development of enterprises. Therefore, for the integrated optimization problem of production planning and scheduling, it is crucial to comprehensively consider order delay and production cost.

[0004] Considering the structural characteristics of the problem, the upper layer is a 0-1 combinatorial problem, and the lower layer is a parallel machine scheduling problem. Therefore, it can be reasonably inferred that the integrated planning and scheduling problem considering order delay and production cost is essentially NP-hard. Generally, minimizing order lateness and minimizing production cost are two conflicting objectives. That is, minimizing order lateness often leads to an increase in production cost, while minimizing production cost may lead to order lateness. Therefore, through effective integrated decision-making, considering both order lateness and production cost can promote and restrict each other. Seeking a dynamic balance in this process of mutual game is of great significance. Summary of the Invention

[0005] Aiming at the deficiencies of the prior art, the present invention proposes an optimization method and system for an integrated model of production planning and scheduling based on a two-layer game model. This method can effectively obtain high-quality solutions and provide decision-making support for managers, thus solving the above technical problems.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] An optimization method for an integrated model of production planning and scheduling based on a two-layer game model, comprising the following steps:

[0008] S1. Construct an upper-layer objective function TD to minimize the total order lateness; at the same time, construct a lower-layer objective function TC to minimize the total production cost of orders; form a two-layer integrated model of production planning and scheduling through the upper and lower layer objective functions, and set constraint conditions;

[0009] The present invention uses the Stackelberg game to construct a two-layer integrated model of production planning and scheduling;

[0010] The expression of the upper-layer objective function TD to minimize the total order lateness is as follows:

[0011]

[0012] In the formula, w i represents the weight of order i; D i represents the due date of order i; C i represents the completion time of order i; N e represents the set of orders completed in advance, N d represents the set of orders completed with delay, N s represents the set of selected orders, N represents the set of orders to be processed;

[0013] The expression of the lower-layer objective function TC to minimize the total production cost of orders is as follows:

[0014]

[0015] In the formula, F represents the set of furnaces; ξ ij represents the number of times order i is processed on furnace j; t ij represents the production cost of processing one furnace of order i on furnace j; σ ij represents the silicon package maintenance cost required for processing order i on furnace j.

[0016] The said setting of constraint conditions includes:

[0017] N e ={i|i∈N s ; D i -C i ≥0};

[0018] N d =N s \N e ;

[0019] It also includes:

[0020] 1) Each order can only be assigned to one furnace and can only be assigned once. The expression is as follows:

[0021]

[0022] In the formula, λ ij represents a binary variable; among them, λ ij =1 indicates that order i is assigned to furnace j, and λ ij =0 indicates that order i is not assigned to furnace j;

[0023] 2) The completion time of the first order processed on furnace j is equal to the processing time. The expression is as follows:

[0024]

[0025] In the formula, represents the completion time of the first order on furnace j; represents the processing time of the first order on furnace j;

[0026] 3) Starting from the second order processed on furnace j, the completion time of the current order is equal to the completion time of the previous order plus the processing time of the current order. The expression is as follows:

[0027]

[0028] In the formula, π jRepresents the processing sequence of the order;

[0029] 4) The number of times order i is processed on furnace j is calculated as follows:

[0030]

[0031] In the formula, w i represents the weight of order i; V j represents the capacity of furnace j; u ij represents the weight of the last furnace of smelting of order i on furnace j; α represents the capacity coefficient of furnace j;

[0032] 5) The silicon package maintenance cost for order i processed on furnace j is expressed as follows:

[0033]

[0034] In the formula, Q represents the capacity of the container filled after order processing, called the silicon package; G i represents the grade of order i; h represents the cost of maintaining the silicon package once, indicating that the silicon package needs to be maintained after being filled 8 times, which conforms to the technical characteristics.

[0035] S2. Construct a two - layer evolutionary algorithm (MHBLEA) based on Markov chain learning and heuristic rules to solve the integrated production planning and scheduling model in S1; the upper layer of the two - layer evolutionary algorithm is a genetic algorithm, and the lower layer is a simulated annealing algorithm;

[0036] Markov chain learning is constructed based on the Markov model;

[0037] The specific steps are as follows:

[0038] S2.1. Set the solution vector of the two - layer evolutionary algorithm in a mixed - coding manner;

[0039] The solution vector includes: a binary vector x i and a real - number vector; select the processing sequence π j of the order as the real - number vector;

[0040] The solution vector is expressed as Φ = [x i |π j ; here, x i represents the selected order, where x i = 1 indicates that the i - th order is selected, and x i = 0 indicates that the i - th order is not selected.

[0041] S2.2. Set the key parameters in the double - layer evolutionary algorithm based on the Markov chain learning model and heuristic rules. The key parameters include the upper - layer population size, the number of optimization iterations, the crossover rate, the mutation rate, and the learning rate; the lower - layer initial temperature, the minimum temperature, the annealing rate, and the number of inner - loop iterations;

[0042] S2.3. Initialize the upper - layer population and sort the population according to the due date to form an initial processing sequence;

[0043] S2.4. Call the heuristic allocation rule to allocate orders to the furnaces, and calculate the un - optimized p ij 、C i 、TD and TC for the orders. Among them, p ij represents the processing time of order i on furnace j;

[0044] The heuristic allocation rule is as follows: For the selected orders, first allocate the orders to the furnaces according to the initial processing sequence; in each allocation process, first exclude the furnace with the largest number of currently allocated orders; next, the remaining furnaces will be considered as candidates, and the furnace with the smallest number of allocated orders will be preferentially selected; if the number of orders on multiple candidate furnaces is the same, calculate the total production cost on each candidate furnace and select the furnace with the lowest cost for allocation; this process will continue until all orders are allocated;

[0045] S2.5. In the upper - layer genetic algorithm, select dominant individuals through tournament selection, generate offspring by single - point crossover, and perform mutation by flipping bits, and update the population at the same time;

[0046] The genetic algorithm includes: selection operation, crossover operation, and mutation operation;

[0047] The criteria for selecting dominant individuals are as follows: First, the genes of each individual are represented as a binary vector, where 1 represents selecting the order and 0 represents not selecting; evaluate the quality of individuals by calculating the individual fitness, and the individual with the highest fitness value is preferentially selected and screened through the tournament selection mechanism; in each generation, individuals evolve through crossover and mutation operations to optimize the order - selection scheme, which is used to reduce the total tardiness and balance the number of selected orders;

[0048] S2.6. In the lower - layer simulated annealing algorithm, obtain the optimal processing sequence by updating the neighborhood solution;

[0049] The update criterion for the lower-layer individuals is as follows: In each iteration, two different order indices (idx1 and idx2) are randomly selected, and an insertion operation is performed. The first order is removed from its original position and inserted into the second position to generate a neighborhood solution. Then, by reallocating the orders to the furnaces and calculating the new cost, if the new cost is lower, the neighborhood solution is accepted, and the current solution is updated to the new neighborhood solution. At the same time, each time a new neighborhood solution is accepted, the current best solution is also checked and updated.

[0050] S2.7. Call the heuristic rules for decoding, allocate the orders to the furnaces, and calculate the optimized p ij , C i , TD, and TC of the orders, and update the population.

[0051] Calculate the optimized p ij , C i , TD, and TC of the orders, that is, calculate the processing time and completion time of the orders, and the objective function values of the upper and lower layers.

[0052] The calculation process is as follows: First, calculate the number of batches required for full-load smelting based on the weight of the orders and the capacity of the furnaces, and determine whether additional smelting batches are needed. If the remaining weight of the last furnace of the order exceeds 30% of the furnace capacity, add one furnace; otherwise, do not add. Then, calculate the processing time based on the total number of batches. The processing time for each batch is fixed at 3 hours. Next, determine whether there is already an order in the current furnace. If not, set the start time of the first order to the preset initial time. If there is already an order, use the completion time of the previous order as the start time of the current order. Finally, calculate the completion time of the order by adding the start time and the processing time, and update the start time and completion time of the order.

[0053] S2.8. Determine whether the learning condition is met. The learning condition is that the selected orders do not reach 80% of the total orders. If it is met, execute step S2.9. If not, jump back to execute step S2.5.

[0054] S2.9. Predict the new population by the Markov chain learning model. The steps are as follows:

[0055] S2.9.1. Extract the selection status of each order from the input order allocation situation and initialize the order selection status.

[0056] S2.9.2. Construct an expression for the transition probability matrix as follows:

[0057]

[0058] In the formula, Denote the one-step transition probability of order \(i\) from state \(z_1\) to \(z_2\); since the state space contains \(N\) sets of orders to be processed, there are \((N - 1)\) one-step transition probability matrices. Considering the memoryless property (memory length is 1) of the binary Markov chain, the dependency relationships between these matrices can be ignored; the OSTPM can be derived from the binary vector of the upper-layer solution (i.e., \(x\) i (gen));

[0059] S2.9.3. Normalize the transition matrix, that is, calculate the sum of each row and divide the elements of each row by the sum of that row to ensure that the elements of each row represent the probability of transitioning from one state to other states;

[0060] S2.9.4. According to the updated transition matrix, predict a new selected state for each order through a random mechanism; based on the selected state of the previous order and the probabilities in the transition matrix, decide whether to select the current order. The selection method is: if the random value is less than the corresponding transition probability, select the order, otherwise do not select;

[0061] S2.9.5. Update the state of the original order according to the new selected state;

[0062] S2.10. Determine whether the termination condition is reached. If it is reached, end and output the optimal processing sequence and the upper and lower layer objective function values. If the algorithm termination condition is not reached, return to S2.5 to continue iterative optimization.

[0063] The present invention also implements the above method through a system of an integrated model of production planning and scheduling based on a two-layer game model. The system includes: a building module and an optimization module;

[0064] The building module is used to perform the following steps:

[0065] S1. Construct the upper-layer objective function TD to minimize the total tardiness of orders; at the same time, construct the lower-layer objective function TC to minimize the total production cost of orders; form a two-layer integrated model of production planning and scheduling through the upper and lower layer objective functions, and set constraint conditions;

[0066] The optimization module is used to perform the following steps:

[0067] S2. Construct a two-layer evolutionary algorithm based on Markov chain learning and heuristic rules to solve the integrated model of production planning and scheduling in S1.

[0068] Advantages of the present invention:

[0069] By introducing an integrated model of production planning and scheduling with unit load balancing, the present invention provides a new solution idea for the integrated optimization problem of production planning and scheduling, especially in reducing order tardiness and production cost.

[0070] The present invention proposes a two - layer evolutionary algorithm based on Markov chain learning to solve this problem. The algorithm mainly includes: a problem - oriented learning model and a method of integrating heuristic rules. By using the transition probability matrix to predict potential excellent individuals, the search is guided into promising regions, stimulating the vitality of the search process and improving the overall performance of the algorithm. This integrated optimization method can generate a set of solutions within a reasonable computing time for decision - makers to select according to their own needs, providing strong support and guidance for the production planning decision - making of enterprises. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] Figure 1 is the flowchart of the steps of the present invention;

[0072] Figure 2 is the overall algorithm flowchart of the present invention;

[0073] Figure 3 is the coding method diagram of the present invention;

[0074] Figure 4 is the schematic diagram of the crossover operation of the present invention;

[0075] Figure 5 is the schematic diagram of the mutation operation of the present invention;

[0076] Figure 6 is the schematic diagram of the operation for generating domain solutions of the present invention;

[0077] Figure 7 is the schematic diagram of the Markov chain random search framework of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0078] The present invention will be further described in detail below with reference to specific embodiments.

[0079] As Figure 1 shown, an optimization method for an integrated model of production planning and scheduling based on a two - layer game model includes the following steps:

[0080] S1. Construct the upper - layer objective function TD to minimize the total order tardiness; at the same time, construct the lower - layer objective function TC to minimize the total production cost of the order; form a two - layer integrated model of production planning and scheduling through the upper - layer and lower - layer objective functions, and set constraint conditions;

[0081] The present invention uses the Stackelberg game to construct a two - layer integrated model of production planning and scheduling;

[0082] The expression of the upper - layer objective function TD to minimize the total order tardiness is as follows:

[0083]

[0084] where w i represents the weight of order i; D i represents the due date of order i; C i represents the completion time of order i; N e represents the set of orders completed in advance, N d represents the set of orders completed late, N s represents the set of selected orders, N represents the set of orders to be processed;

[0085] The expression of the lower-layer objective function TC to minimize the total production cost of orders is as follows:

[0086]

[0087] where F represents the set of furnaces; ξ ij represents the number of times order i is processed on furnace j; t ij represents the production cost of processing one furnace of order i on furnace j; σ ij represents the silicon package maintenance cost required for processing order i on furnace j.

[0088] The set constraints include:

[0089] N e ={i|i∈N s ; D i -C i ≥0};

[0090] N d =N s \N e ;

[0091] It also includes:

[0092] 1) Each order can only be assigned to one furnace and can only be assigned once. The expression is as follows:

[0093]

[0094] where λ ij represents a binary variable; among them, λ ij =1 means that order i is assigned to furnace j, and λ ij =0 means that order i is not assigned to furnace j;

[0095] 2) The completion time of the first order processed on furnace j is equal to the processing time. The expression is as follows:

[0096]

[0097] In the formula, represents the completion time of the first order on furnace j; represents the first order processing time on furnace j;

[0098] 3) Starting from the second order processed on furnace j, the completion time of the current order is equal to the completion time of the previous order plus the processing time of the current order. The expression is as follows:

[0099]

[0100] In the formula, π j represents the processing sequence of the order;

[0101] 4) The number of times order i is processed on furnace j is calculated as follows:

[0102]

[0103] In the formula, w i represents the weight of order i; V j represents the capacity of furnace j; u ij represents the weight of the last furnace of order i smelted on furnace j; α represents the capacity coefficient of furnace j;

[0104] 5) The silicon package maintenance cost for order i processed on furnace j is expressed as follows:

[0105]

[0106] In the formula, Q represents the capacity of the container filled after order processing, called the silicon package; G i represents the grade of order i; h represents the cost of maintaining the silicon package once, indicating that the silicon package needs to be maintained after being filled 8 times, which conforms to the technical characteristics.

[0107] S2. As Figure 2 shown, a two-layer evolutionary algorithm based on Markov chain learning and heuristic rules is constructed to solve the integrated production planning and scheduling model in S1; the upper layer of the two-layer evolutionary algorithm is a genetic algorithm, and the lower layer is a simulated annealing algorithm; the specific steps are as follows:

[0108] Markov chain learning is constructed based on the Markov model;

[0109] S2.1. Set the solution vector of the two-layer evolutionary algorithm in a mixed coding manner;

[0110] That is, convert the solution in the production scheduling problem into a representation form that can be processed by the algorithm, so as to provide an operable data structure for the subsequent evolutionary algorithm;

[0111] The solution vector includes: a binary vector x i and a real number vector; select the processing sequence π of the order j as a real number vector;

[0112] The solution vector is represented as Φ = [x i |π j ; where π j represents the selected order, where x i = 1 indicates that the i-th order is selected, and x i = 0 indicates that the i-th order is not selected;

[0113] Taking N = 10 and as an example, the solution vector is:

[0114] Φ = [1, 0, 1, 1, 0, 1, 0, 1, 1, 1|3, 0, 10, 8, 0, 1, 0, 2, 5, 4]

[0115] It means that orders 1, 3, 4, 6, 8, 9, 10 are selected, and the assigned order is: 3 → 10 → 8 → 1 → 2 → 5 → 4. The coding schematic diagram is as Figure 3 shown;

[0116] S2.2. Set the key parameters in the double-layer evolutionary algorithm based on the Markov chain learning model and heuristic rules. The key parameters include the upper population size, the number of optimization iterations, the crossover rate, the mutation rate, and the learning rate; the lower initial temperature, the minimum temperature, the annealing rate, and the number of inner loop iterations;

[0117] S2.3. Initialize the upper population and sort the population according to the due date to form the initial processing sequence;

[0118] S2.4. Call the heuristic assignment rule to assign the order to the furnace, and calculate the unoptimized p ij , C i , TD, and TC of the order; where p ij represents the processing time of order i on furnace j;

[0119] The heuristic assignment rule is: for the selected orders, first assign the orders to the furnaces according to the initial processing sequence; in each assignment process, first exclude the furnace with the most currently assigned orders; next, the remaining furnaces will be considered as candidates, and the furnace with the fewest assigned orders will be preferentially selected; if the number of orders of multiple candidate furnaces is the same, calculate the total production cost on each candidate furnace and select the furnace with the lowest cost for assignment; this process will continue until all orders are assigned;

[0120] S2.5. In the upper-layer genetic algorithm, dominant individuals are selected through tournament selection, offspring are generated by single-point crossover, and mutation is performed by flipping bits, while the population is updated.

[0121] The genetic algorithm includes: selection operation, crossover operation, and mutation operation; the crossover operation is as Figure 4 shown, and the mutation operation is as Figure 5 shown.

[0122] The criteria for selecting dominant individuals are as follows: First, the genes of each individual are represented as a binary vector, where 1 indicates selecting the order and 0 indicates not selecting; the fitness of the individual is calculated to evaluate the quality of the individual, and the individual with the highest fitness value is preferentially selected and screened through the tournament selection mechanism; in each generation, the individuals evolve through crossover and mutation operations to optimize the order selection scheme, which is used to reduce the total tardiness and balance the number of selected orders.

[0123] S2.6. As Figure 6 shown, in the lower-layer simulated annealing algorithm, the optimal processing sequence is obtained by updating the neighborhood solution.

[0124] The update criteria for the lower-layer individuals are as follows: In each iteration, two different order indices (idx1 and idx2) are randomly selected, and an insertion operation is performed. The first order is removed from its original position and inserted into the second position, thus generating a neighborhood solution; then, by reallocating the orders to the furnaces and calculating the new cost, if the new cost is lower, the neighborhood solution is accepted, and the current solution is updated to the new neighborhood solution; at the same time, each time a new neighborhood solution is accepted, the current best solution is also checked and updated.

[0125] S2.7. Call the heuristic rule for decoding, allocate the orders to the furnaces, calculate the optimized p ij , C i , TD, and TC of the orders, and update the population.

[0126] Calculating the optimized p ij , C i , TD, and TC of the orders means calculating the processing time and completion time of the orders.

[0127] The calculation process is as follows: First, calculate the number of batches required for full-load smelting based on the weight of the order and the capacity of the furnace, and determine whether additional smelting batches are needed. If the remaining weight of the last furnace of the order exceeds 30% of the furnace capacity, add one more furnace; otherwise, do not add. Then, calculate the processing time based on the total number of batches. The processing time for each batch is fixed at 3 hours. Next, determine whether there is already an order for the current furnace. If not, set the start time of the first order to the preset initial time. If there is already an order, use the completion time of the previous order as the start time of the current order. Finally, calculate the completion time of the order by adding the start time and the processing time, and update the start time and completion time of the order.

[0128] S2.8. Determine whether the learning condition is met. The learning condition is that the selected orders do not reach 80% of the total orders. If it is met, execute step S2.9. If not, jump back to execute step S2.5.

[0129] S2.9. As Figure 7 shown, predict the new total population by the Markov chain learning model. The steps are as follows:

[0130] S2.9.1. Extract the selection status of each order from the input order allocation situation and initialize the order selection status.

[0131] S2.9.2. Construct an expression for the transition probability matrix as follows:

[0132]

[0133] In the formula, represents the one-step transition probability of order i from state z1 to z2. Since the state space contains N processed order sets, there are (N - 1) one-step transition probability matrices. Considering the memoryless property (memory length is 1) of the binary Markov chain, the dependency between these matrices can be ignored. The OSTPM can be derived from the binary vector of the upper-layer solution (i.e., x i (gen)).

[0134] S2.9.3. Normalize the transition matrix, that is, calculate the sum of each row and divide each element in the row by the sum of the row to ensure that each element in the row represents the probability of transitioning from one state to other states.

[0135] S2.9.4. According to the updated transition matrix, predict the new selection status for each order through a random mechanism. Determine whether to select the current order based on the selection status of the previous order and the probability in the transition matrix. The selection method is: if the random value is less than the corresponding transition probability, select the order; otherwise, do not select.

[0136] S2.9.5. Update the status of the original order according to the new selection status;

[0137] S2.10. Determine whether the termination condition is reached. If it is reached, end and output the optimal processing sequence and the upper and lower layer objective function values. If the algorithm termination condition is not reached, return to S2.5 to continue iterative optimization.

[0138] To verify the effectiveness of the proposed two-layer evolutionary algorithm based on Markov chain learning and integrated heuristic rules (MHBLEA), the present invention was compared with two mainstream advanced algorithms, BLEAQ and NGA, and solved for different scales of test cases. The scales of the test cases are (processing order set N × furnace set F): 20×3_1, 20×3_2, 50×3_1, 50×3_2, 50×5_1, 50×5_2, 70×3_1, 70×3_2, 70×5_1, 70×5_2, 100×3_1, 100×3_2, 100×5_1, 100×5_2, 100×7_1, 100×7_2, 150×5_1, 150×5_2, 150×7_1, 150×7_2, a total of 20 test cases.

[0139] To ensure the fairness of the comparison, all algorithms were independently run 20 times on each test case, and 30×|N|×|N|×|F| (ms) was used as the algorithm termination condition.

[0140] All algorithms were run on an Intel I5 processor (1.8 GHz), 8 GB of memory, a Win10 operating system, and a Python 3.10 programming environment.

[0141] To evaluate the quality of the solutions obtained by the MHBLEA of the present invention through algorithm comparison, a well-known utility function was used as the performance metric; the evaluations of the objective functions at different levels were integrated to provide a comprehensive metric; the performance metric used is defined as follows:

[0142]

[0143] In the formula, λ1 represents the first coefficient, λ2 represents the second coefficient, and λ1 = λ2 = 0.5; represents the maximum upper layer objective function value in one run of the three algorithms; represents the minimum lower layer objective function value in one run of the three algorithms; and are the upper and lower layer objective function values of the current algorithm and are non-zero values; in MHBLEA, the smaller the calculated result of the given metric, the better the effect.

[0144] The obtained experimental results are shown in Table 1, which are the statistical results of the indicators obtained by the three algorithms of MHBLEA, BLEAQ, and NGA for solving 20 test cases. MHBLEA is significantly better than the two reference algorithms in all test cases and shows better performance in terms of WST, BST, and AVG values. This competitive performance can be attributed to the search component of MHBLEA, which discovers valuable knowledge and identifies high-quality solutions through learning methods. At the same time, the generation and update mechanism of the upper-layer high-quality population promotes the lower-layer local optimization by providing effective guidance, thus balancing the solution quality and diversity of the upper-layer population. These search components significantly enhance the global search ability of the MHBLEA algorithm. Therefore, the MHBLEA algorithm proposed in the present invention has effectiveness and superiority.

[0145] Table 1: Comparison results of the indicators of BLEAQ, NGA, and MHBLEA

[0146]

[0147] The system of the integrated model of production planning and scheduling based on the two-layer Stackelberg game of the present invention implements the above method. The system includes: a building module and an optimization module;

[0148] The building module is used to perform the following steps:

[0149] S1. Construct the upper-layer objective function TD to minimize the total order tardiness; at the same time, construct the lower-layer objective function TC to minimize the total production cost of the order; form a two-layer integrated model of production planning and scheduling through the upper and lower-layer objective functions, and set the constraint conditions;

[0150] The optimization module is used to perform the following steps:

[0151] S2. Construct a two-layer evolutionary algorithm based on Markov chain learning and heuristic rules to solve the integrated model of production planning and scheduling in S1.

[0152] Although the present invention has been shown and described through embodiments, those of ordinary skill in the art should understand that various modifications, changes, or substitutions can be made to these embodiments without departing from the principles and spirit of the present invention. Therefore, the protection scope of the present invention should be defined by the appended claims and their equivalents.

Claims

1. An optimization method for an integrated model of production planning and scheduling based on a two-layer game model, characterized in that, It includes the following steps: S1. Construct the upper-layer objective function TD to minimize the total order tardiness; at the same time, construct the lower-layer objective function TC to minimize the total production cost of the order; form a two-layer integrated production planning and scheduling model through the upper and lower-layer objective functions, and set the constraint conditions; S2. Construct a two-layer evolutionary algorithm based on Markov chain learning and heuristic rules to solve the integrated production planning and scheduling model in S1, and output the optimal processing sequence, upper-layer objective function value, and lower-layer objective function value; the upper layer of the two-layer evolutionary algorithm is a genetic algorithm, and the lower layer is a simulated annealing algorithm.

2. The optimization method for the integrated model of production planning and scheduling based on a two-layer game model according to claim 1, characterized in that The expression of the upper-layer objective function TD to minimize the total order tardiness is as follows: where w i represents the weight of order i; D i represents the due date of order i; C i represents the completion time of order i; N e Indicates the set of orders completed in advance, N d Indicates the set of orders completed with delay, N s Indicates the set of selected orders, N indicates the set of orders to be processed; The expression of the lower-layer objective function TC to minimize the total production cost of the order is as follows: In the formula, F represents the set of furnaces; ξ ij represents the number of times order i is processed on furnace j; t ij represents the production cost of processing one furnace of order i on furnace j; σ ij represents the silicon package maintenance cost required for processing order i on furnace j; The setting of the constraint conditions includes: N e = {i | i ∈ N s ; D i - C i ≥ 0}; N d = N s \N e ; Wherein, N e represents the set of orders completed in advance; N s represents the set of selected orders; D i represents the due date of order i; C i represents the completion time of order i; N d represents the set of orders completed with delay; It also includes: 1) Each order can only be assigned to one furnace, and each order can only be assigned once; 2) The completion time of the first order processed on furnace j is equal to the processing time; 3) Starting from the second order processed on furnace j, the completion time of the current order is equal to the completion time of the previous order plus the processing time of the current order; 4) The number of times order i is processed on furnace j; 5) The silicon package repair cost for order i processed on furnace j.

3. An optimization method for an integrated model of production planning and scheduling based on a two-layer game model according to claim 1, characterized in that The steps of constructing a two-layer evolutionary algorithm based on Markov chain learning and heuristic rules to solve the integrated production planning and scheduling model in S1 are as follows: S2.

1. Set the solution vector of the two-layer evolutionary algorithm in a hybrid coding manner; S2.

2. Set the key parameters in the two-layer evolutionary algorithm based on the Markov chain learning model and heuristic rules. The key parameters include the upper-layer population size, optimization iteration times, crossover rate, mutation rate, and learning rate; the lower-layer initial temperature, minimum temperature, annealing rate, and inner-loop iteration times; S2.

3. Initialize the upper-layer population and sort the population according to the due date to form an initial processing sequence; S2.

4. Call the heuristic allocation rule to allocate the order to the furnace, and calculate the unoptimized p of the order ij , C i , TD and TC; where p ij represents the processing time of order i on furnace j; S2.

5. In the upper-layer genetic algorithm, select dominant individuals through tournament selection, generate offspring by single-point crossover, and perform mutation by flipping bits, and update the population at the same time; The genetic algorithm includes: selection operation, crossover operation, and mutation operation; S2.

6. In the lower-layer simulated annealing algorithm, obtain the optimal processing sequence by updating the neighborhood solution; S2.

7. Call the heuristic rules for decoding, allocate the orders to the furnaces, and calculate the optimized p, ij C i TD and TC of the orders, and update the population; S2.

8. Judge whether the learning condition is reached. The learning condition is that the selected orders do not reach 80% of the total orders. If it is reached, execute step S2.

9. If it is not reached, jump to execute step S2.5; S2.

9. Predict the new population by the Markov chain learning model; S2.

10. Judge whether the termination condition is reached. If it is reached, end and output the optimal processing sequence and upper and lower-layer objective function values. If the algorithm termination condition is not reached, return to S2.5 to continue iterative optimization.

4. An optimization method for an integrated model of production planning and scheduling based on a two-layer game model according to claim 4, characterized in that The steps of predicting the new population by the Markov chain learning model are as follows: S2.

9. Predict the new population by the Markov chain learning model, and the steps are as follows: S2.9.

1. Extract the selection status of each order from the input order allocation situation and initialize the order selection status; S2.9.

2. Construct the transition probability matrix; S2.9.

3. Normalize the transition matrix, i.e., calculate the sum of each row and divide each element in the row by the sum of the row, to ensure that the elements in each row represent the probabilities of transitioning from one state to other states; S2.9.

4. According to the updated transition matrix, predict a new selection state for each order through a random mechanism; based on the selection state of the previous order and the probabilities in the transition matrix, determine whether to select the current order. The selection method is as follows: if the random value is less than the corresponding transition probability, select the order; otherwise, do not select it; S2.9.

5. Update the status of the original order according to the new selection state.

5. A system of an integrated model for production planning and scheduling based on a two-layer game model, characterized in that, The system includes: a building module and an optimization module; The building module is used to perform the following steps: S1. Construct an upper-layer objective function TD to minimize the total order tardiness; at the same time, construct a lower-layer objective function TC to minimize the total production cost of the order; form a two-layer integrated production planning and scheduling model through the upper and lower-layer objective functions and set constraint conditions; The optimization module is used to perform the following steps: S2. Construct a two-layer evolutionary algorithm based on Markov chain learning and heuristic rules to solve the integrated production planning and scheduling model in S1.