A reentrant mixed flow shop production scheduling method based on reinforcement learning

By employing a reinforcement learning-based approach and utilizing bidirectional long short-term memory networks and proximal policy optimization algorithms, the multi-objective scheduling problem in a reentrant hybrid flow shop was solved. This approach optimized the maximum completion time and total delay time, improving production efficiency and resource utilization, and promoting the intelligent development of the manufacturing industry.

CN119690009BActive Publication Date: 2025-10-24SHENYANG INST OF AUTOMATION - CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411817753.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-11
Publication Date
2025-10-24
Estimated Expiration
2044-12-11

AI Technical Summary

Technical Problem

Traditional scheduling methods struggle to simultaneously optimize and minimize maximum completion time and total delay time in reentrant hybrid flow shops, resulting in insufficient production efficiency and accuracy, high computational complexity, and difficulty in handling multi-objective scheduling problems.

Method used

A reinforcement learning-based approach is adopted, utilizing a bidirectional long short-term memory network to construct a scheduling strategy. The model is trained through a proximal policy optimization algorithm to achieve intelligent decision-making that minimizes the maximum completion time and total delay time. By combining Markov process and reward function design, production scheduling is optimized.

Benefits of technology

Significantly improve production efficiency, optimize resource allocation, increase production flexibility and accuracy, reduce production costs, and promote the intelligent upgrading of manufacturing industry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119690009B_ABST
    Figure CN119690009B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of reentrant mixed flow shop production scheduling method based on reinforcement learning, first the production demand of reentrant mixed flow shop is analyzed, the information of current product production is determined;Then determine optimization goal, minimize maximum completion time and minimize total delay time as optimization goal, determine constraint condition and parameter, and determine the constraint and related parameter variable in optimization problem, establish multi-objective reentrant mixed flow shop scheduling problem model;Optimization solution is carried out using reinforcement learning, based on Markov decision process model is described, constructs state space, action space, to overcome the challenge brought by reward sparse problem in reinforcement learning, adopt the way of layered reward, construct global external reward and internal reward after completing each stage production;BiLSTM is used to build scheduling strategy, extract internal scheduling information, train the model using proximal policy optimization (PPO), realize the intelligent decision of action selection.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of production scheduling, and specifically relates to a reentrant mixed flow shop production scheduling method based on reinforcement learning. BACKGROUND

[0002] In the manufacturing industry, the reentrant mixed flow shop scheduling problem is a complex and critical problem that involves the reasonable arrangement of production tasks among multiple production lines and multiple processes. This problem has a high degree of dynamicity and uncertainty, as production tasks often have different processing sequences, process requirements, and resource constraints. Traditional scheduling methods, such as heuristic algorithms and genetic algorithms, often struggle to obtain satisfactory scheduling results when faced with such complex problems. Many existing scheduling methods mainly focus on a single objective, such as minimizing the production cycle or maximizing equipment utilization. However, in actual production, multiple objectives often need to be considered simultaneously, such as minimizing the maximum completion time and minimizing the total delay time. Single-objective optimization methods often cannot balance these objectives, resulting in unsatisfactory scheduling results. For large-scale, multi-variety, and high-complexity production tasks, the computational complexity of traditional scheduling methods is often very high. This not only increases the computational cost and time, but also may affect the accuracy and feasibility of the scheduling results.

[0003] Multi-objective optimization in the reentrant mixed flow shop scheduling problem refers to considering both the minimization of the maximum completion time and the minimization of the total delay time. Minimizing the maximum completion time means trying to shorten the completion time of the latest task among all production tasks to improve overall production efficiency. Minimizing the total delay time means trying to reduce the deviation between the actual completion time and the planned completion time of production tasks to improve the accuracy and reliability of production planning. There is a mutual constraint and balance between these two objectives. For example, to minimize the maximum completion time, the priority of some production tasks may need to be sacrificed, resulting in an increase in the total delay time. Conversely, to minimize the total delay time, the allocation and priority of production tasks may need to be adjusted, affecting the maximum completion time. Therefore, how to balance and optimize between the two objectives is a major challenge in the reentrant mixed flow shop scheduling problem. In summary, the existing technology has many limitations in solving the multi-objective optimization of the reentrant mixed flow shop scheduling problem. SUMMARY

[0004] Aiming at the equipment selection and process sequencing problem of re-entrant hybrid flow shop production, a re-entrant hybrid flow shop production scheduling method based on reinforcement learning is proposed, the scheduling objectives are to minimize the maximum completion time and minimize the total delay, based on Markov process, a bidirectional long short-term memory network is used to construct the scheduling strategy, extract internal scheduling information, and use proximal policy optimization (PPO) to train the model to realize intelligent decision-making of action selection.

[0005] The technical scheme adopted by the present application to achieve the above-mentioned purpose is:

[0006] A re-entrant hybrid flow shop production scheduling method based on reinforcement learning, comprising the following steps:

[0007] 1) Based on the relevant information of the production workshop, a mathematical model of the re-entrant hybrid flow shop scheduling problem is constructed;

[0008] 2) Based on the mathematical model of the re-entrant hybrid flow shop scheduling problem, a reinforcement learning model is constructed;

[0009] 3) Based on the bidirectional long short-term memory network, potential scheduling information is extracted, a scheduling strategy model is constructed, proximal policy optimization method is used to train the scheduling strategy model, and the trained scheduling strategy model is used to obtain intelligent decision-making of action selection.

[0010] The step 1) comprises the following steps:

[0011] 1.1) Determine the processing equipment information and workpiece information;

[0012] 1.2) Minimize the maximum completion time and minimize the total delay time as the optimization objective of the re-entrant hybrid flow shop production scheduling problem;

[0013] 1.3) According to the hypothesis and the actual processing process, determine the constraint conditions and related parameter variables in the re-entrant hybrid flow shop production scheduling problem.

[0014] The constraint conditions include:

[0015]

[0016]

[0017]

[0018]

[0019]

[0020]

[0021] where i is the workpiece index, I is the workpiece set, I = {1, 2,..., n}, n is the total number of workpieces, j is the production stage index, J is the production stage set, J = {1, 2,..., p}, p is the total number of production stages, l is the machining level index, L is the machining level set, L = {1, 2,..., t}, t is the total number of machining levels, k is the machine index, K j is the set of available machines for production stage j, j j is the total number of available machines for production stage j, j P ijl is the machining time of workpiece i at the first level of production stage j, ijl S ijl is the start machining time of workpiece i at the first level of production stage j, i C max is the completion machining time of workpiece i at the first level of production stage j, ijlk d i is the due date of workpiece i, ij C ijl is the maximum completion time, A is a positive number, x ijl and are decision variables.

[0022] The step 2) comprises the following steps:

[0023] 2.1) determining the state space based on Markov process and mathematical model of the problem;

[0024] 2.2) determining the action space;

[0025] 2.3) designing the reward function in a hierarchical reward manner according to the optimization objective of the model.

[0026] The state space S is:

[0027]

[0028] where W i represents workpiece i, O ij represents the machining operation of workpiece i at production stage j, each operation in the job has four characteristics: characteristic 1. The number of layers of the operation, indicating the machining layer number of the current operation; characteristic 2. The scheduling state of the operation, indicating the start machining time S ijl of workpiece i at the first level of production stage j; characteristic 3. The machining time of the operation, indicating the machining time P ijl of workpiece i at the first level of production stage j; characteristic 4. The arrangement equipment of the operation.

[0029] The action space A is:

[0030] A = A o × A m

[0031] wherein A o represents an operation selection action space, containing all operations that meet the requirements and do not meet the requirements, A m represents a machine selection action space, also containing all machines that meet the requirements and do not meet the requirements, if the workpiece has not been scheduled, the operation that meets the requirements is the first operation; if the workpiece has been partially scheduled, the operation that meets the requirements is the next operation after the operation that has been scheduled, and for the unassigned operation, the machine that meets the requirements is the machine that can perform the operation.

[0032] The reward function R is:

[0033]

[0034] wherein ω1 represents the reward weight of the maximum completion time, and ω2 represents the reward weight of the total delay time.

[0035] To achieve the goal of minimizing the maximum completion time and the total delay time, after the agent completes a stage of processing, the agent will obtain two stage rewards R p1 and R p2 :

[0036] R p1 =-(S ijl -C ij-1l )

[0037]

[0038] wherein C iJl,Max represents the latest completion processing time of the workpiece i in the production stage j of the lth layer, when the actual completion processing time exceeds C iJl,Max , the workpiece has a delay risk, then the total stage reward R p obtained by the agent after completing a stage of processing is:

[0039] R p =ω p1 R p1 +ω p2 R p2

[0040] wherein ω p1 represents the reward weight of the stage reward R p1 , and ω p2 represents the reward weight of the stage reward R p2 .

[0041] The step 3) comprises the following steps:

[0042] 3.1) input the scheduling scene features into the BiLSTM-based scheduling strategy model;

[0043] 3.2) The scheduling strategy model outputs a feature matrix that combines global information and local information;

[0044] 3.3) Perform probability distribution of network computing operation sequencing and machine selection sub-problems, and select the action with the maximum probability among the eligible actions;

[0045] 3.4) Assign the selected job to the selected machine and set it at the earliest workable position.

[0046] The present application has the following advantages and benefits:

[0047] 1. Significantly improve production efficiency: By minimizing the maximum completion time, this method can ensure the close connection between each process on the production line, reduce the bottleneck link in the production cycle, and significantly improve the overall production efficiency. At the same time, minimizing the total delay time helps to reduce the delay in order delivery, improve customer satisfaction, and further promote the improvement of production efficiency.

[0048] 2. Intelligent decision-making and adaptive ability: Using a bidirectional long short-term memory network (Bi-LSTM) to build a scheduling strategy can capture the timing dependency and internal scheduling information in the production process, providing strong support for intelligent decision-making. Through the proximal policy optimization (PPO) algorithm to train the model, the model has the ability to adaptively adjust the scheduling strategy in complex and variable production environments, so as to cope with various uncertain factors.

[0049] 3. Optimize resource allocation: This method can accurately predict the time and resources required for each process, thereby optimizing resource allocation and reducing resource waste. Through intelligent scheduling, it can ensure the balanced use of resources such as equipment and personnel on the production line, improving resource utilization.

[0050] 4. Promote the intelligent upgrading of manufacturing industry: The successful application of this method will promote the development of manufacturing industry towards intelligence and automation, and improve the core competitiveness of enterprises. By introducing advanced reinforcement learning technology and intelligent scheduling algorithm, it can bring more efficient and flexible production mode for enterprises.

[0051] 5. Reduce production cost: By optimizing production process and resource allocation, this method helps to reduce production cost and improve enterprise's profitability. At the same time, reducing production delay and waste also helps to reduce the operating cost of enterprises. BRIEF DESCRIPTION OF DRAWINGS

[0052] Figure 1 Framework diagram of multi-action scheduling strategy based on bidirectional long short-term memory network;

[0053] Figure 2 Flowchart of bidirectional long short-term memory network operation. DETAILED DESCRIPTION

[0054] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0055] like Figure 1 and Figure 2 As shown in Figure 1, consider a reentrant hybrid flow shop scheduling problem. There are N workpieces to be processed, with a total of p processing stages. Each workpiece undergoes t processing layers, passing through these p stages in sequence at each layer, resulting in t-1 reentrancy. Each processing stage has a processor dedicated to a specific process, and at least one stage has multiple parallel machines. All processors at the same stage can complete the processing of each workpiece, and all parallel machines processing the same process have the same processing rate. A workpiece can only enter the next stage after completing the corresponding process at each stage. When faced with parallel processors, a workpiece can only be processed by one of the parallel machines. The processing time for each process on the corresponding processor is known. The choice of parallel processor for each workpiece at each stage and the order in which workpieces are processed on the same machine are determined to minimize both maximum processing time and total delay.

[0056] Problem Assumptions:

[0057] The process flow of N workpieces is exactly the same, and they are processed t times in sequence through p processing stages.

[0058] The machine is available at any time, regardless of machine failure and maintenance.

[0059] All workpieces arrive at time 0 and can be processed.

[0060] There is no priority between workpieces, that is, once any process of all workpieces is started, it cannot be stopped until the process is completed.

[0061] The transportation time of the workpiece between the workstations at each stage is not taken into account.

[0062] Each machine can only process one workpiece at a time, and each workpiece can only be processed by one machine at any time.

[0063] The present invention provides a production scheduling method for a reentrant hybrid flow shop based on reinforcement learning, which comprises the following steps:

[0064] S1: Determine the current product production quantity, process requirements, equipment quantity, equipment processability and other information;

[0065] The input of the optimization method comprises processing equipment information and workpiece information, and the output is a sorted processing procedure arrangement sequence, corresponding maximum completion time and delay time; the processing equipment information comprises processing equipment type, processing equipment quantity, and processing equipment processable procedure; the workpiece information comprises workpiece quantity, workpiece production stage, workpiece production in-out layer number, workpiece procedure sequence, and workpiece production stage processing time; and the optimization target of the model is the minimum maximum completion time and the minimum total delay time of all workpieces.

[0066] S2: determining an optimization target, taking the minimum maximum completion time and the minimum total delay time as the optimization target of the reentrant hybrid flow shop production scheduling problem;

[0067] S3: determining constraint conditions and parameters, determining the constraints and related parameter variables in the reentrant hybrid flow shop production scheduling problem based on the proposed hypothesis and actual processing process;

[0068] The parameter symbol definitions of the mathematical model are as follows:

[0069] i: workpiece index;

[0070] I: workpiece set, I = {1, 2,..., n}, n is the total number of workpieces;

[0071] j: production stage index;

[0072] J: production stage set, J = {1, 2,..., p}, p is the total number of production stages;

[0073] l: processing level index;

[0074] L: processing level set, L = {1, 2,..., t}, t is the total number of processing levels;

[0075] k: machine index;

[0076] K j : available machine set of production stage j, K j = {1, 2,..., m j} m j is the total number of available machines of production stage j;

[0077] P ijl : processing time of workpiece i in production stage j of the lth layer;

[0078] S ijl : starting processing time of workpiece i in production stage j of the lth layer;

[0079] C ijl : completion processing time of workpiece i in production stage j of the lth layer;

[0080] d i : delivery date of workpiece i;

[0081] C max : maximum completion time;

[0082] A: a sufficiently large positive number;

[0083] x ijlk : decision variable, if workpiece i is processed on machine k in production stage j of layer l, it takes the value 1, otherwise 0;

[0084] decision variable, if workpiece i1 is processed before workpiece i2 on machine k in production stage j of layer l, it takes the value 1, otherwise 0.

[0085] The mathematical model of the scheduling problem is represented as follows:

[0086] Optimization objective

[0087] min f1 = min C max (1)

[0088]

[0089] Constraint conditions

[0090]

[0091]

[0092]

[0093]

[0094]

[0095]

[0096]

[0097]

[0098] wherein, optimization objective (1) is to minimize the maximum completion time, optimization objective (2) is to minimize the maximum delay time; constraint condition (3) indicates that the value of the maximum completion time should be greater than or equal to the completion time of the last process of all workpieces at the last level; constraint condition (4) indicates that each workpiece at each level can only be processed by one machine at each stage; constraint condition (5) is that any process of all workpieces cannot stop once the process is started until the process is completed; constraint conditions (6) and (7) are process priority constraints, that is, the workpiece can only start processing of the next stage at the level after completing the immediately preceding process, and can only start processing of the first process at the next level after completing the last stage at the previous level; constraint condition (8) ensures that a subsequent workpiece can only be processed after a workpiece is processed; constraint condition (9) is that the start time of each production stage of each workpiece at each level should be no earlier than 0 o'clock; and constraint condition (10) is the value range of the decision variable.

[0099] S4: determining a state space, determining a state space in the reinforcement learning process based on a Markov process and a mathematical model of the problem, including scheduling state characteristics of each operation, processing time characteristics and processing machine characteristics;

[0100] Further, the Markov decision process is defined as a four-tuple <S, A, R, P>, wherein S represents a state space, A represents an action space, R represents a reward function, and P represents a state transition probability matrix.

[0101] Further, the state space S: at each scheduling decision point, the environmental characteristics include all job characteristics. Each job characteristic includes a set of operation sequences, as shown in formula (11),

[0102]

[0103] wherein W i represents a workpiece i, O ij represents a processing operation of the workpiece i at the production stage j. Each operation in the job has four characteristics: characteristic 1. The number of layers of the operation, indicating the number of processing layers of the current operation, denoted as l; characteristic 2. The scheduling state of the operation, indicating the start processing time S ijl of the workpiece i at the production stage j at the lth layer; characteristic 3. The processing time of the operation, indicating the processing time P ijl of the workpiece i at the production stage j at the lth layer; and characteristic 4. The arrangement equipment of the operation.

[0104] S5: determining an action space, when applying reinforcement learning to solve the scheduling problem, an operation to be scheduled is first selected, and then a machine to execute the operation is selected, so a multi-action space design is adopted;

[0105] The action space is represented as: A = A o × A m , where A o represents the operation selection action space, including all eligible and ineligible operations, A m represents the machine selection action space, also including all eligible and ineligible machines. If the workpiece has not been scheduled, its eligible operation is the first operation; if it has been partially scheduled, its eligible operation refers to the next operation after the already scheduled operation. For unassigned operations, its eligible machine is the machine that can perform the operation.

[0106] S6: Determine the reward function, and design the reward function in a hierarchical reward manner according to the optimization objective of the model;

[0107] Further, the external reward formulated based on the optimization objective of the model is represented as: Where ω1 represents the reward weight of the maximum completion time, and ω2 represents the reward weight of the total delay time.

[0108] To achieve the objectives of minimizing the maximum completion time and minimizing the total delay time, after the agent completes a stage of processing, it will obtain two stage rewards R p1 and R p2 : R p1 = -(S ijl -C ij-1l ), where C iJl,Max represents the latest completion processing time of workpiece i in the production stage j of the lth layer, and when the actual completion processing time exceeds C iJl,Max , the workpiece has a risk of delay. The total stage reward obtained by the agent after completing a stage of processing is R p = ω p1 R p1 + ω p2 R p2 , where ω p1 represents the reward weight of the stage reward R p1 , and ω p2 represents the reward weight of the stage reward R p2 .

[0109] S7: Form an optimization scheme, based on Markov process, use bidirectional Long-short Memory Network (BiLSTM) to construct a scheduling strategy, extract internal scheduling information, use Proximal Policy Optimization (PPO) to train the model, and realize intelligent decision-making of action selection.

[0110] BiLSTM is used as a bidirectional version of LSTM to extract hidden features in the scheduling scenario. First, the agent observes the scheduling environment and inputs the scheduling scenario features into the BiLSTM-based feature extraction network; second, the feature extraction network outputs a feature matrix that combines global and local information; then, the action network calculates the probability distribution of the operation sequencing and machine selection sub-problems and selects the action with the highest probability among the eligible actions; finally, the selected job is assigned to the selected machine and set at the earliest processable position. It mainly consists of two steps: job feature extraction and environment feature fusion. In the job feature extraction step, the state features of all operation sequences of job i are input into the BiLSTM network to obtain the forward feature h f and the reverse feature h r . Then the two features are connected to form the state features of job i (h f , h r ). Since the network input structure from different jobs is consistent and the gating parameters within the network are consistent, the extracted encoding is uniform. In the environment feature fusion step, the average of all job features is calculated to represent the environment state feature H.

[0111] In summary, the reentrant hybrid flow shop production scheduling method based on reinforcement learning has significant benefits. It not only improves production efficiency and optimizes resource allocation, but also improves production flexibility and response speed, reduces production cost, and promotes the intelligent upgrading of manufacturing industry. These advantages will bring significant economic and social benefits to enterprises.

Claims

1. A method for production scheduling of a re-entrant hybrid flow shop based on reinforcement learning, characterized in that, The method comprises the following steps: 1) constructing a mathematical model of a re-entrant hybrid flow shop scheduling problem based on production workshop related information; 2) constructing a reinforcement learning model based on the mathematical model of the re-entrant hybrid flow shop scheduling problem; 3) extracting potential scheduling information based on a bidirectional long short-term memory network, constructing a scheduling strategy model, training the scheduling strategy model by using a proximal policy optimization method, and obtaining an intelligent decision of action selection by using the trained scheduling strategy model; The step 1) comprises the following steps: 1.1) determining processing equipment information and workpiece information; 1.2) taking minimization of maximum completion time and minimization of total delay time as optimization objectives of the re-entrant hybrid flow shop scheduling problem; 1.3) determining constraint conditions and related parameter variables in the re-entrant hybrid flow shop scheduling problem according to assumptions and actual processing processes; The constraint conditions comprise: Where i is the workpiece index, I is the workpiece set, I = {1, 2, ..., n}, n is the total number of workpieces, j is the production stage index, J is the production stage set, J = {1, 2, ..., p}, p is the total number of production stages, l is the processing level index, L is the processing level set, L = {1, 2, ..., t}, t is the total number of processing levels, k is the machine index, K j is the set of available machines in production stage j, K j ={1,2,...m j }, m j is the total number of available machines in production stage j, P ijl is the processing time of workpiece i in production stage j at level l, S ijl is the starting processing time of workpiece i in production stage j at level l, C ijl is the completion time of workpiece i in production stage j at level l, d i is the delivery time of workpiece i, C max is the maximum completion time, A is a positive number, x ijlk and is the decision variable.

2. The method of claim 1, wherein, The step 2) comprises the following steps: 2.1) determining a state space based on a Markov process and a problem mathematical model; 2.2) determining an action space; 2.3) designing a reward function in a hierarchical reward manner according to an optimization objective of the model.

3. The method of claim 2, wherein, The state space S is: wherein W i represents the workpiece i, O ij represents the machining operation of the workpiece i production stage j, each operation in the job has four characteristics: characteristic 1. The number of layers of the operation, the machining layer where the current operation is located is represented by l; characteristic 2. The scheduling state of the operation, the starting machining time S ijl of the workpiece i in the production stage j of the lth layer is represented; characteristic 3. The machining time of the operation, the machining time P ijl of the workpiece i in the production stage j of the lth layer is represented; characteristic 4. The arrangement equipment of the operation.

4. The method of claim 2, wherein, The action space A is: A = A o x A m where A o denotes the operation selection action space, containing all the operations that are in compliance and not in compliance with the requirements, A m denotes the machine selection action space, also containing all the machines that are in compliance and not in compliance with the requirements, if the workpiece has not been scheduled, its operation in compliance is the first operation; if the workpiece has been partially scheduled, its operation in compliance is the next operation after the operation that has been scheduled, and for the unassigned operation, the machine in compliance is the machine that can perform the operation.

5. The method of claim 2, wherein, The reward function R is: Wherein, ω1 represents a reward weight of the maximum completion time, and ω2 represents a reward weight of the total delay time.

6. The method of claim 5, wherein, To achieve the goal of minimizing the maximum completion time and the total delay time, the agent will obtain two stage rewards R p1 with R p2 : R p1 =-(S ijl -C ij-1l ) where C ijl,Max represents the latest completion machining time of workpiece i at the production stage j of the lth layer, and when the actual completion machining time exceeds C ijl,Max , the workpiece has a delay risk, and the total stage reward R p obtained by the intelligent agent after completing a stage of machining is: R p = ω p1 R p1 + ω p2 R p2 where ω p1 represents the reward weight of the stage reward R p1 . ω p2 represents the reward weight of the stage reward R p2 .

7. The method of claim 1, wherein, The step 3) comprises the following steps: 3.1) inputting scheduling scene features into a scheduling strategy model based on BiLSTM; 3.2) outputting a feature matrix combining global information and local information by the scheduling strategy model; 3.3) performing network calculation operation ordering and probability distribution of machine selection two sub-problems, and selecting an action with the maximum probability from the actions meeting the conditions; 3.4) assigning the selected job to the selected machine, and setting the earliest processable position.

Citation Information

Patent Citations

  • Method and device for optimizing scheduling problem of reentrant hybrid flow shop

    CN115222107A

  • Workshop production real-time scheduling method and system based on AI industrial Internet of Things

    CN116560323A