Distributed assembly replacement flow shop scheduling method based on deep reinforcement learning
By employing a deep reinforcement learning method based on the PPO algorithm, the scheduling problem of distributed assembly replacement flow shop was solved, the job scheduling was optimized, and the total processing time was reduced while the intelligence of the scheduling process was improved.
Patent Information
- Application Number
- CN202510921254.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-11-11
AI Technical Summary
Existing deep reinforcement learning algorithms have not yet effectively solved the distributed assembly replacement flow shop scheduling problem. In particular, under the challenges of multi-shop task correlation, performance differences and spatial and temporal complexity, how to achieve stable and efficient production resource optimization remains a difficult problem.
A deep reinforcement learning method based on the PPO algorithm is adopted to optimize job scheduling by setting a reward function and composite scheduling rules. This includes designing the state space, action space and reward mechanism, and combining factory allocation and job sorting rules to optimize the total processing time.
It effectively reduced the total processing time of the distributed assembly replacement flow workshop scheduling, improved the adaptability and intelligence of the scheduling process, and achieved faster status quo resolution.
Smart Images

Figure CN120930981A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of manufacturing and scheduling, specifically relating to a distributed assembly replacement flow shop scheduling method based on deep reinforcement learning. Background Technology
[0002] Emerging manufacturing models such as cloud manufacturing, intelligent manufacturing, and service-oriented manufacturing have brought new opportunities and challenges to the manufacturing industry. To reduce costs and improve efficiency, many manufacturers have adopted distributed manufacturing models that span multiple regions, factories, and workshops. Unlike traditional single-factory manufacturing, distributed manufacturing emphasizes the allocation and coordination of overall manufacturing resources. The scheduling problem of distributed workshops is characterized by the correlation of tasks across multiple workshops, the differences in performance between different workshops, spatial dispersion, and temporal complexity. How to achieve stable and efficient optimization of production and manufacturing resource allocation in a multi-workshop distributed scheduling scenario is the core problem and research focus of distributed manufacturing. To solve this problem, a new scheduling problem has emerged: the Distributed Assembly Permutation Flowshop Scheduling Problem (DAPFSP).
[0003] The Distributed Assembly Replacement Flow Shop Scheduling Problem (DAPFSP) combines the characteristics of both distributed shop scheduling and assembly replacement flow shop scheduling, and has been widely applied in industries such as automotive and textiles. Compared to traditional distributed shop scheduling and assembly replacement flow shop scheduling problems, DAPFSP is more complex and more difficult to solve, making it a typical NP-hard problem. Its complexity lies in the fact that, in addition to the constraints of conventional PFSP, it also includes constraints related to product assembly and multiple factories. The former means that assembling a product can only proceed if all operations required for assembling that product are completed and the assembly machines are idle; the latter means that operations can be performed in any factory.
[0004] To address these challenges, researchers have proposed a variety of intelligent scheduling algorithms aimed at improving the efficiency and flexibility of the scheduling process.
[0005] Deep Reinforcement Learning (DRL) algorithms combine the feature extraction capabilities of deep neural networks with the decision optimization capabilities of reinforcement learning, providing a novel approach to solving complex scheduling problems. In particular, the Proximal Policy Optimization (PPO) algorithm, due to its good stability and efficiency, has been increasingly applied to distributed shop floor scheduling problems, significantly improving the adaptability and intelligence of the scheduling process. However, current research has not yet utilized deep reinforcement learning algorithms to solve the distributed assembly replacement flow shop scheduling problem. Summary of the Invention
[0006] To address the aforementioned issues, this invention provides a distributed assembly displacement flow shop scheduling method based on deep reinforcement learning. It is based on the PPO algorithm and aims to minimize the total processing time. By setting a reward function and composite scheduling rules, it optimizes job scheduling.
[0007] The present invention adopts the following technical solution:
[0008] A distributed assembly permutation flow shop scheduling method based on deep reinforcement learning includes the following steps:
[0009] Step 1: Build the environmental framework of the distributed assembly replacement flow workshop and initialize the environmental parameters, set the constraints of workshop scheduling, and set the scheduling objective to minimize the total flow time;
[0010] Step 2: Define the state space, action space, and reward mechanism, and calculate the impact of different reward mechanisms on the total flow time;
[0011] Step 3: Update the job and machine information;
[0012] Step 4: Run the PPO algorithm for learning and training;
[0013] Step 5: Record the scheduling results and determine whether the training results have reached the termination condition. If yes, proceed to step 6; otherwise, return to step 4.
[0014] Furthermore, in step 1, before setting up the environment, the scheduling problem and scheduling objectives must be clearly defined. This invention considers distributed assembly replacement flow shop scheduling.
[0015] In Step 1, the environment framework for building a distributed assembly replacement flow shop is designed. The flow shop includes two stages: production and assembly. In the production stage, there are n jobs and F identical factories. Job i needs to be assigned to any factory f for processing. Each factory is equipped with m identical machines, which is equivalent to the same flow shop. In this flow shop, each job is processed sequentially on different machines. The processing process on each machine represents a process step. The processing sequence of each process step in each job is fixed and must be processed on the machines in a certain order. In each flow shop, all jobs need to be processed according to the machine sequence and along the same path.
[0016] During the assembly phase, there are s products and one assembly machine. Each product h has a defined set of tasks. Assembly of a product can only begin when all tasks for that product are completed and the assembly machine is idle. Tasks belonging to the same product are processed continuously. The next product's process can only begin after all tasks belonging to that product have been completed.
[0017] Table 1 shows the DAPFSP 8_2_3_3 example, indicating the processing time and product assembly time for each job.
[0018] Table 1DAPFSP 8_2_3_3 calculation example
[0019]
[0020] Using the DAPFSP 8_2_3_3 example described in Table 1, the generated Gantt chart is as follows: Figure 3 As shown.
[0021] Aside from processing, assembly, and waiting times, other time consumption is not considered. The environmental parameters and mathematical symbols involved in this invention are shown below:
[0022] 1. Basic parameters:
[0023] n: Number of assignments
[0024] m: Number of machines
[0025] F: Number of factories
[0026] s: Number of products
[0027] j: Machine index
[0028] i: Job index, i∈1,2,...,n;
[0029] f: Factory index, f∈1,2,...,F;
[0030] h,z: Product index, h,z∈1,2,…,s;
[0031] k: Location index. In the scheduling sequence of factory f, k = 1, 2, ..., n represent the first, second, ..., n jobs respectively.
[0032] M = {M1, M2, ..., M} m}: Represents the set of all machines;
[0033] J = {J1, J2, ..., J} n}: A set of n jobs to be processed;
[0034] P = {P1, P2, ..., P} s}: Represents a collection of products;
[0035] n f The number of jobs assigned to factory f;
[0036] n h : indicates the constituent product P h The number of jobs; Product P h It consists of multiple tasks;
[0037] O i,j Homework J i The j-th operation;
[0038] P i,j Homework J i In machine M j Processing time;
[0039] t h Product P h Assembly time;
[0040] c h Product P h The number of tasks that have been completed;
[0041] r h Product P h The maximum completion time of the completed tasks, i.e., the current ready time;
[0042] C j (f): The completion time of the last job on machine j in factory f;
[0043] l i : The number of operations completed in operation i, i.e., the number of machines used in operation i. i ∈{0,1,2,…,m}
[0044] 2. Auxiliary variables:
[0045] π fπ is the set of processing sequences within factory f. f =(π) t,1 ,π t,2 ,…,π t,n );
[0046] C k,j,f The operation is at the k-th position in factory f, and occurs on machine M. j The time required to complete the processing;
[0047] π h Product P h The set of scheduling sequences for the middle job, π h =(v h,1, π h,2 ,..,v h,n );
[0048] D πh Product P h The final completion moment;
[0049] Product P h Homework J i The time when the processing is completed;
[0050] P h Product P h The time between the completion of all processing operations and the assembly stage represents product P. h Release time;
[0051] δ j,f If machine M j If the element belongs to factory f, the value is 1; otherwise, it is 0.
[0052] TF: Total traffic time for all products;
[0053] p f p represents the set of processing times for each job finally assigned in factory f on each machine. f =(p f,1 ,p f,2 ,…,p f,m );
[0054] v h Product number in sequence;
[0055] 3. Decision variables:
[0056] X i,k,f If homework J i If a component occupies position k in factory f, its value is 1; otherwise, it is 0.
[0057] Y h,z If product Ph Must be ranked in product P z If the preceding product is used as its predecessor, the value is 1; otherwise, it is 0.
[0058] Based on the above notation, the Distributed Assembly Replacement Flow Shop Scheduling Problem (DAPFSP) model used in this patent can be mathematically described as follows:
[0059]
[0060]
[0061] Constraint (1) shows the objective function, minimizing the total flow time. Constraint (2) for each job J i It must be assigned to a specific position k in a factory f. Constraint (3) Each position k in each factory f can only be occupied by one job, and each machine M j Only one job can be processed at a time. Constraint (4) Calculate job J i Machine M in factory f j The completion time of the k-th position. Constraints (5) and (6) indicate that each product must have only one predecessor and one successor. Constraint (7) indicates that product P h The release time is equal to the completion time of its last job. Constraint (8) is for product P. h The completion time. Constraint (9) ensures that the product's operation completion time and assembly time are positive values. Constraint (10) defines the range of values for the decision variables.
[0062] Furthermore, in step 2, the state space, action space, and reward mechanism are crucial in the PPO algorithm. The algorithm learns, explores, and trains by observing the initial state, then selecting an action from the action space at each time step and observing the next state and reward returned by the environment. The state space of this invention includes five state values: job completion state, machine state, job scheduling sequence, product assembly state, and pre-constraint state.
[0063]
[0064] S5=(v1,v2,…,v s (15)
[0065] Equation (11) records the completion time of the three most recently assigned jobs in each factory f on each machine j. The positions of the three most recently assigned jobs represent the positions of the last, second to last, and third to last jobs assigned on the machine, respectively. If a job does not exist, it is filled with 0.
[0066] Equation (12) shows the operating status of each machine, indicating whether it is idle or busy. statusj∈{0,1} indicates whether machine j is busy. The value is 0 when idle and 1 when busy.
[0067] Equation (13) describes the scheduling sequence of operations in each factory.
[0068] Status (14) indicates the completion status of each product.
[0069] State (15) reflects the prerequisite relationships between products, and the state space can be represented as:
[0070] S = {S1,S2,S3,S4,S5}.
[0071] This patent considers the action space of DAPFSP, which includes sub-problems of factory allocation and job sequencing within factories. At each decision point, the decision-maker must determine not only the factory selection rules but also the job allocation rules. To this end, 20 composite scheduling rules are proposed, first selecting the factories to be allocated, and then sequencing the jobs within the factories.
[0072] The factory allocation rules include the following five:
[0073] S1 Factory Integrated Load Balancing Rule FCLB: Selects the appropriate factory to allocate new products based on the current load of each factory. When allocating tasks, it considers three key indicators: 1. Machine busyness: The total remaining available time of all machines in the factory is added up. The busier the machines are, the higher the score; 2. Order backlog: The more orders currently in the queue, the higher the score; 3. Stacking of the same product: The more orders of the same model are in the queue, the higher the score.
[0074] The factory's total load score = equipment busyness score + order backlog score + same product accumulation score;
[0075] By calculating the total load score for each factory, the factory with the lowest total load score is finally determined as a candidate; during runtime, all factories with the same and lowest total load score will be filtered, and one of them will be randomly selected for task assignment.
[0076] S2 Factory Integrated Capacity Balancing Rule (FCCB): Selects a suitable factory for allocating specific products based on the factory's production capacity. When allocating tasks, it considers three key indicators: 1. Immediately available capacity: the more machines currently idle, the higher the score; 2. Recent release capacity: the more machines predicted to be idle in the near future, the higher the score; 3. Overruns of similar products: the more jobs processing the same product currently, the more points are deducted.
[0077] The overall capacity score is calculated as follows: score of immediately available capacity + score of recently released capacity + deduction for suppressed similar products. The factory with the highest overall capacity score is selected as a candidate. During operation, all factories with the same and highest overall capacity score will be screened, and one of them will be randomly selected for task assignment.
[0078] S3 Future Time Window Load Balancing Rule (FTLB): The core of this factory selection strategy is to predict the processing capacity of each factory within a future time period △T and select the factory most likely to process new jobs quickly. The specific calculation involves three steps:
[0079] First, determine the time window range as the current time T + time period △T;
[0080] Then, two key indicators are calculated for each factory: 1. Number of available machines: the more machines that are idle within the time window, the higher the score; 2. Queue load: reflecting the current queue pressure: the more jobs currently being processed and those in the queue, the higher the score.
[0081] Load balancing score = score for available machines + deduction for queue load;
[0082] The factory with the highest load balancing score is ultimately selected; if multiple factories have the same score, a random selection is made.
[0083] S4 Shortest Queue Load Balancing Rule SQ: Calculate the current number of job orders for each factory, find the factory with the shortest queue length, where the queue length is the total number of jobs currently being processed and those in the queue; then randomly select one factory from those with the same queue length as a candidate factory for task allocation.
[0084] S5 Random Factory Rule: Randomly select a factory.
[0085] First, select a factory according to the factory allocation rules, insert the workpieces to be processed, and then rearrange the processing order of all workpieces according to the job sorting rules.
[0086] Assignment sorting rules (four rules):
[0087] 1. Shortest Processing Time Rule (SPT): Sort the job list and, based on the processing time of each job in its current process, place the job with the shortest processing time at the top.
[0088] 2. Longest Processing Time Rule (LPT): Sort the job list and, based on the processing time of each job in its current process, place the job with the longest processing time at the top.
[0089] 3. Bottleneck Shortest Processing Time Rule (BSPT): Pre-calculate the index of the bottleneck process, sort the job list, and rank the job with the shortest processing time at the bottleneck process first.
[0090] 4. Bottleneck Remaining Processing Time Priority Rule (BRPT): First, calculate the remaining waiting time for each job at the bottleneck process. For jobs that have already completed the bottleneck process, their priority is 0; for jobs that have not yet reached the bottleneck process, they are sorted according to their processing time at the bottleneck process and their distance from the bottleneck process, with jobs closer to the bottleneck having higher priority.
[0091] Furthermore, in step 2, the reward mechanism of the present invention comprehensively evaluates the scheduling performance through multiple core and auxiliary items to promote system improvements in reducing total makespan and achieving load balancing.
[0092] First, the rewards have a dynamic weight adjustment mechanism. The Makespan weight (TIME_FACTOR 0.3→0.8) and the load balancing weight (LOAD_FACTOR 0.5→0.1) change as the training progresses, allowing the algorithm to automatically adjust its optimization priorities at different stages. In the early stages, it prioritizes ensuring system stability (load balancing) while exploring more possibilities, and in the later stages, it focuses on optimizing the core metric (Makespan).
[0093] The rewards consist of the following components:
[0094] 1. Dual Makespan Incentive Mechanism: The total flow time Makespan is the total flow time TF of all products. A baseline value for Makespan is set. When the actual Makespan is lower than the baseline, a positive reward is given; otherwise, a penalty is given.
[0095] 2. Continuous improvement mechanism: A reward is given if the current total flow time Makespan is shorter than the previous one;
[0096] 3. Intelligent load assessment mechanism: The load balance is assessed by standard deviation. The smaller the standard deviation, the more balanced the load. The penalty is calculated as np.std(overall load value) * LOAD_FACTOR. The load balance weight LOAD_FACTOR is between 0.1 and 0.5. The overall load value = 0.7 * number of current factory operations + 0.3 * total available time of all machines in the factory. The more balanced the load, the smaller the penalty.
[0097] 4. Assisted penalty optimization: Set delivery time penalty and assembly delay penalty. For the current job, if its completion time exceeds the scheduled delivery time, or if the assembly time is inconsistent with the completion time, penalties will be applied respectively. The greater the total number of delays of all jobs, the greater the penalty for delivery time. When the completion time of the current job is inconsistent with the assembly time of the product belonging to this job, the greater the absolute value of the difference between the two, the greater the penalty will be applied.
[0098] All rewards and penalties are added together to form the total reward, total_reward. According to the above process, the reward is an evaluation of the overall scheduling result of the PPO algorithm selecting the corresponding scheduling rule. The higher the reward, the better the evaluation. When the same single dynamic event occurs, the corresponding scheduling rule will appear more frequently through the output of the PPO algorithm.
[0099] Furthermore, in step 3, the information about the job and the machine is updated. Whenever an action is selected, the start time, actual processing time, and end time of the selected job and machine are first confirmed (where the start time is the greater of the completion time of the same job on the previous machine and the completion time of the previous job on the same machine, i.e., max(C)). k,j-1,f C k-1,j,f The processing end time is max(C). k,j-1,f C k-1,j,f )+P i,j Subsequently, update the machine's processing time C. j(f) Number of completed assignments: l i , Number of unfinished assignments (ml) i Scheduling sequence X i,k,f Whenever a job completes a process, the number of completed processes is incremented by 1, and the number of remaining unfinished jobs is decremented by 1. At this time, it is determined whether the number of processes in the current job exceeds the number of processes n that the job should have. If it does, the set J of unfinished processes of the job is deleted.
[0100] Finally, update the product status. Once all tasks related to this product have been completed, the product can be assembled. The earliest assembly start time is the completion time of the last task belonging to this product. Calculate the total flow time TF after all products are assembled.
[0101] In step 4, the PPO (Proximal Policy Optimization) algorithm used in this invention is based on a standard neural network structure, including an input layer, hidden layers, and an output layer. The dimension of the input layer is consistent with the state space, receiving state information from the scheduling environment. The hidden layer contains two networks: the Actor network consists of two fully connected layers of 64 neurons each, using the ReLU activation function, and outputs the action probability distribution; the Critic network also consists of two fully connected layers of 64 neurons each, outputting a value estimate of the current state. The input and hidden layers of the Actor and Critic networks can share parameters, but the output layers are designed independently to separate the policy from the value assessment. Regarding training configuration, the Adam optimizer is used to set independent learning rates for the Actor and Critic networks. In the loss function, the Actor loss is based on a pruning objective function to ensure stable policy updates, while the Critic loss uses mean squared error (MSE) to reduce the difference between the predicted state value and the actual reward. The algorithm process includes forward propagation, action selection, and policy update. The input state generates an action probability distribution through the Actor network, and the Critic network generates a state value estimate. After selecting an action, it is executed to obtain a new state and reward. The parameters of the Actor and Critic networks are updated using the collected data through an objective function.
[0102] Furthermore, in step 5, the recorded scheduling result is the processing time (makespan) after all jobs have completed one processing cycle, and the training rounds are set to 500 rounds in this invention.
[0103] Finally, in step 5, by comparing the completion times generated in each iteration, the effectiveness and convergence of the algorithm can be clearly seen. The cumulative reward value of each iteration can be seen in the output results, proving that the algorithm is constantly being optimized.
[0104] Furthermore, in step 5, the recorded scheduling results are the time and reward after all jobs have completed one processing cycle, and the training rounds are set to 500 in this invention.
[0105] The beneficial effects of this invention are mainly reflected in:
[0106] 1) Use the PPO algorithm to solve the distributed assembly replacement flow shop scheduling problem.
[0107] 2) A new composite scheduling rule and a new reward mechanism were designed to enable the algorithm to learn better. The design of the composite scheduling rule and reward mechanism helps the algorithm find a better solution more quickly. Attached Figure Description
[0108] Figure 1 This is a flowchart illustrating a distributed assembly replacement flow workshop scheduling method disclosed in an example of the present invention.
[0109] Figure 2 This is a schematic diagram of a distributed assembly and replacement flow workshop processing flow disclosed in an example of the present invention.
[0110] Figure 3 This is a Gantt chart of the DAPFSP 8_2_3_3 example disclosed in this invention.
[0111] Figure 4 This is a schematic diagram of a composite rule disclosed in an example of the present invention.
[0112] Figure 5 This is a schematic diagram of the PPO algorithm framework disclosed in an example of the present invention.
[0113] Figure 6 This is an iterative diagram illustrating the application of the DQN algorithm to solve the distributed assembly replacement flow shop scheduling problem, as disclosed in an example of the present invention.
[0114] Figure 7 This is an iterative diagram illustrating the application of the PPO algorithm to solve the distributed assembly replacement flow shop scheduling problem, as disclosed in an example of the present invention. Detailed Implementation
[0115] The present invention will now be further described with reference to the accompanying drawings.
[0116] refer to Figure 1 A distributed assembly replacement flow shop scheduling method includes the following steps: building the environmental framework of the distributed assembly replacement flow shop and initializing environmental parameters, defining the state space, action space and reward mechanism, designing composite rules, designing reward mechanism, updating information, running the PPO algorithm, recording scheduling results and drawing iterative graphs, etc.
[0117] Step 1: Build the environment framework of the distributed assembly and replacement flow workshop and initialize the environment parameters.
[0118] See Figure 2 Taking a typical fabricated assembly line workshop as an example, this workshop includes two stages: a processing stage and an assembly stage. The processing stage has n jobs and F identical factories. Jobs j∈1,2,…,n need to be assigned to any factory f∈1,2,…,F for processing. Each factory is equipped with m identical machines, equivalent to the same assembly line workshop. The processing sequence of each step in each job is fixed and must be performed on the machines in a specific order. In each assembly line workshop, all jobs need to be processed according to the machine sequence and along the same path.
[0119] In the assembly phase, there are *s* products and one assembly machine. Each product, r∈1,2,…,s, has definite partial operations. Assembly of a product can only begin when all operations belonging to that product are completed and the assembly machine is idle. Operations belonging to the same product are processed consecutively. The next product's operations can only proceed after all operations belonging to that product have been completed. Time consumption other than processing, assembly, and waiting time is not considered.
[0120] Step 3: Design composite rules, see Figure 4 There are five factory allocation rules (FCLB, FCCB, FTLB, SQ, Random) and four job sequencing rules (SPT, LPT, BSPT, BLPT), for a total of 20 composite scheduling rules.
[0121] S1 Factory Integrated Load Balancing Rule (FCLB): This rule selects the appropriate factory to allocate new products based on the current load of each factory. When allocating tasks, it considers three key indicators: 1. Machine Busyness: This sums the remaining available time of all machines in the factory. The less remaining available time, the busier the machine, and the higher the score. The score is calculated as: Remaining available time of all machines in the factory * 1; 2. Order Backlog: The more orders currently in the queue, the higher the score. The score is calculated as: Number of orders currently in the queue * 500 (each pending order is equivalent to adding 500 units of load); 3. Product Backlog: The more orders of the same model in the queue, the higher the score. The score is calculated as: Number of orders of the same model in the queue * 1000 (a backlog of the same model will bring an additional 1000 times the pressure).
[0122] The total load score for a factory is calculated as follows: equipment busyness score + order backlog score + same product accumulation score. By calculating the total load score for each factory, the factory with the lowest total load score is selected as a candidate. During runtime, all factories with the same and lowest total load score are filtered, and one is randomly selected for task allocation.
[0123] S2 Factory Integrated Capacity Balancing Rule (FCCB): Selects a suitable factory for assigning specific products based on the factory's production capacity. When assigning tasks, it considers three key indicators: 1. Immediately available capacity (70% weight): The more currently idle machines, the higher the score; 2. Recent release capacity (30% weight): Machines predicted to be idle within the next 5 minutes; the more machines expected to be idle in the near future, the higher the score; 3. Overstock of similar products (-2 points per order): The more jobs processing the same product, the more points are deducted.
[0124] Overall capacity score = score of immediately available capacity + score of recently released capacity + deduction for suppression of similar products = number of currently available machines * 0.7 + number of available machines in the next 5 minutes * 0.3 - number of similar products * 2.
[0125] The three key metrics used in the FCCB evaluation have different dimensions because the number of jobs processing the same product (i.e., similar products) is small, while the number of currently idle machines is much larger. Therefore, for overall balance, the product of the number of similar products is multiplied by a relatively large coefficient.
[0126] The factory with the highest overall capacity score is selected as a candidate. During runtime, all factories with the same overall capacity score and the highest score will be filtered, and one of them will be randomly selected for task assignment.
[0127] S3 Future Time Window Load Balancing Rule (FTLB): The core of this factory selection strategy is to predict the processing capacity of each factory within a future 50-minute timeframe and select the factory most likely to process new jobs quickly. The specific calculation involves three steps:
[0128] First, define the time window range as the current time T + 50 minutes.
[0129] Then, two key indicators are calculated for each factory: 1. Number of available machines: the more machines that are idle within the time window, the higher the score; 2. Queue load: reflecting the current queue pressure: the more jobs currently being processed and those in the queue, the higher the score.
[0130] Load balancing score = Score for available machines + Deduction for queue load = Number of available machines within the time window - Total number of jobs currently being processed and queued for processing * 0.5 (0.5 is a coefficient for overall balancing).
[0131] The processing capacity of a factory is quantified using a load balancing score equal to the score of available machines plus a deduction for queue load. Available machines directly represent idle resources, while queue load is subtracted because new jobs need to compete for resources with existing queues. The factory with the highest load balancing score is ultimately selected; if multiple factories have the same score, a random selection is made. This ensures that new jobs are assigned to the factory with the strongest processing capacity in the next 50 minutes, considering both immediate idle resources and avoiding excessively long waiting queues.
[0132] S4 Shortest Queue Load Balancing Rule SQ: Calculate the current number of job orders for each factory, find the factory with the shortest queue length, where the queue length is the total number of jobs currently being processed and those in the queue; then randomly select one factory from those with the same queue length as a candidate factory for task allocation.
[0133] S5 Random Factory Rule: Randomly select a factory.
[0134] First, select a factory according to the factory allocation rules, insert the workpieces to be processed, and then rearrange the processing order of all workpieces according to the job sorting rules.
[0135] The assignment sorting rules include the following four:
[0136] 1. Shortest Processing Time Rule (SPT): Sort the job list and, based on the processing time of each job in its current process, place the job with the shortest processing time at the top.
[0137] 2. Longest Processing Time Rule (LPT): Sort the job list and, based on the processing time of each job in its current process, place the job with the longest processing time at the top.
[0138] 3. Bottleneck Process Shortest Processing Time Rule (BSPT): Pre-calculate the index of the bottleneck process, sort the job list, and rank the job with the shortest processing time at the bottleneck process first.
[0139] 4. Bottleneck Remaining Processing Time Priority Rule (BRPT): First, calculate the remaining waiting time of each job at the bottleneck process. For jobs that have completed the bottleneck process, their priority is 0. For jobs that have not yet reached the bottleneck process, they are sorted according to their processing time at the bottleneck process and their distance from the bottleneck process. The closer the job is to the bottleneck, the higher its priority.
[0140] Step 4: Design a reward mechanism. Evaluate scheduling performance comprehensively through multiple core and auxiliary items to improve total makespan and load balancing.
[0141] First, the rewards have a dynamic weight adjustment mechanism. The Makespan weight (dynamic coefficient TIME_FACTOR 0.3→0.8) and the load balancing weight (dynamic coefficient LOAD_FACTOR 0.5→0.1) change as the training progresses, allowing the algorithm to automatically adjust its optimization focus at different stages. In the early stages, it prioritizes ensuring system stability (load balancing) while exploring more possibilities, and in the later stages, it focuses on optimizing the core metric (Makespan).
[0142] The rewards consist of the following components:
[0143] 1. Dual Makespan incentive mechanism (baseline_reward and makespan_penalty): Set the baseline value to 1500 (adjustable according to the scenario). When the actual Makespan is lower than the baseline, the reward is calculated as (baseline value - actual value) * 0.8 * TIME_FACTOR (dynamic coefficient). At the same time, the makespan penalty is calculated as exp(0.001 * makespan (total flow time)) * TIME_FACTOR, using an exponential function to apply the penalty to the total flow time. As the total flow time increases, the penalty will also increase.
[0144] 2. Continuous Improvement Reward Mechanism: If the current total flow time (that is, the total flow time of TF products) is shorter than the previous one, a reward will be given, calculated as (previous round makespan - current round makespan) * 2, to encourage continuous improvement of the system.
[0145] 3. Intelligent load assessment mechanism (load_balance_reward): The load balance is assessed by standard deviation. The penalty is calculated as np.std(overall load value) * LOAD_FACTOR. The overall load value = 0.7 * the current number of operations in the factory + 0.3 * the total available time of all machines in the factory. This can detect the risk of factory overload earlier and avoid misjudgment of "idling machines + long queues".
[0146] 4. Assisted Penalty Optimization (due_penalty and assembly_penalty): Set delivery time penalty and assembly delay penalty. For the current job, if its completion time exceeds the scheduled delivery time, or if the assembly time is inconsistent with the completion time, penalties will be applied respectively. Delivery time penalty = the sum of the delays of all jobs in the entire factory * 0.02. Assembly delay penalty = the absolute value of the difference between the completion time of the current job and the assembly time of the product belonging to this job * 0.2. The two penalties respectively punish the overall delay of the job and the production and assembly incoordination of a specific job.
[0147] All rewards and penalties are added together to form the total reward, total_reward. In this way, the system can comprehensively consider multiple factors, thereby optimizing the scheduling strategy.
[0148] Step 5: Update the relevant information for the job and the machine. Since scheduling requires scheduling based on the processing time of the job and the machine, the status of each job and each machine needs to be updated after processing is completed to ensure the accuracy of the scheduling information.
[0149] The processing start time is the greater of the completion time of the same job on the previous machine and the completion time of the previous job on the same machine. The processing end time is the processing start time plus the actual processing time. Subsequently, the machine's processing time, the number of completed jobs, the number of remaining unfinished jobs, and the scheduling sequence are updated. Whenever a job completes a process, the number of completed processes is incremented by 1, and the number of remaining unfinished jobs is decremented by 1. At this time, it is determined whether the number of processes of the current job exceeds the total number of processes m that the job should have. If it does, the set J of unfinished processing jobs of the job is deleted.
[0150] Finally, update the product status. Once all jobs belonging to the product have been completed, the product can be assembled. The earliest start time for assembly is the completion time of the last job belonging to the product. Calculate the total flow time TF after all products are assembled.
[0151] Step 6: Run the PPO algorithm for learning and training.
[0152] Table 2 Core Experimental Parameters
[0153]
[0154] See Figure 5Before training begins, the environment needs to be initialized (30 products, 200 jobs, 4 factories, 5 machines per factory). This step generates an initial state vector containing factory load, machine status, and job characteristics. The Actor network then receives this state and outputs the probability distribution of choices for each scheduling rule combination. Next, the Agent randomly selects one of 20 composite rules as an action based on this probability distribution and submits it to the environment for execution. The environment advances the simulation time, calculates new factory states and job completion status, and generates reward signals based on multiple objectives such as completion time and load balancing. During this process, it also checks whether the termination condition is met. Afterward, the Agent stores the current state, selected action, obtained reward, new state, and termination flag in the experience pool. Simultaneously, the Critic network evaluates the value of the current state and calculates the advantage value to guide policy optimization. Next, a batch of data is randomly drawn from the experience pool, and the Actor network uses the PPO pruning mechanism (clip_param = 0.3) to control the difference between the old and new policies for parameter updates. Simultaneously, the Critic network also optimizes (GAEλ = 0.95) to improve the accuracy of state value prediction. The Adam optimizer (learning rate 3e-4) combines policy loss, value loss, and entropy reward to calculate the total loss and update the network parameters. This process repeats continuously until 500 rounds are reached. Ultimately, the agent can adaptively select efficient combinations of scheduling rules under various production states, thereby minimizing the maximum completion time.
[0155] Step 7: Record the scheduling results and determine whether the training time has reached the termination condition. After conducting a certain number of experiments, the algorithm is selected to undergo 500 rounds of learning and training.
[0156] Step 8: Draw an iterative graph of total processing time and reward, and analyze the scheduling results. (See attached graph) Figure 6 As shown.
[0157] Following the above process, the PPO algorithm is replaced with the DQN algorithm, and the final iterative result of the scheduling problem is shown below. Figure 7 As shown.
[0158] according to Figure 6 and Figure 7 The PPO algorithm is compared with DQN. An iteration graph is plotted after each method runs for a fixed number of rounds, with the vertical axis representing total processing time and reward, and the horizontal axis representing the number of rounds. Analysis of the iteration graphs shows that the PPO algorithm has better convergence and the ability to find the optimal value than DQN.
[0159] This embodiment discloses a distributed assembly substitution flow shop scheduling method based on deep reinforcement learning. It uses PyCharm software and implements the deep learning model using the gym and PyTorch packages. The deep reinforcement learning model can be run in PyCharm to realize a distributed assembly substitution flow shop scheduling system. The main steps in the process are building the distributed assembly substitution flow shop scheduling environment, designing the reward mechanism, designing composite scheduling rules, and running the PPO algorithm to solve the problem. All of these steps are implemented on the PyCharm platform and will not be elaborated further. The main function of this invention is to solve a distributed assembly substitution flow shop scheduling problem using the PPO algorithm. Figure 6 The iterative graph comparison demonstrates that a distributed assembly permutation flow shop scheduling method based on deep reinforcement learning can effectively reduce the total processing time of the flow shop scheduling and can be applied in practical fields. Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention. In summary, the above descriptions are only preferred embodiments of the present invention, and all equivalent changes and modifications made within the scope of the claims of the present invention should be covered by the patent of the present invention.
Claims
1. A distributed assembly permutation flow shop scheduling method based on deep reinforcement learning, characterized in that... Includes the following steps: Step 1: Build the environmental framework of the distributed assembly replacement flow workshop and initialize the environmental parameters, set the constraints of workshop scheduling, and set the scheduling objective to minimize the total flow time; Step 2: Define the state space, action space, and reward mechanism, and calculate the impact of different reward mechanisms on the total flow time; Step 3: Update the job and machine information; Step 4: Run the PPO algorithm for learning and training; Step 5: Record the scheduling results and determine whether the training results have reached the termination condition. If yes, proceed to step 6; otherwise, return to step 4. Step 6: Draw an iterative graph of total flow time, with the vertical axis representing the total flow time after each scheduling cycle and the horizontal axis representing the number of scheduling cycles completed. Analyze the scheduling results based on the iterative graph to obtain the workshop scheduling result with the minimum total flow time.
2. The distributed assembly permutation flow shop scheduling method based on deep reinforcement learning as described in claim 1, characterized in that... In Step 1, the framework for building a distributed assembly displacement flow shop is designed. The flow shop consists of two phases: production and assembly. The production phase has n jobs and F identical factories. Job i needs to be assigned to any factory f for processing. Each factory is equipped with m identical machines, which is equivalent to the same flow shop. In this flow shop, each job is processed sequentially on different machines. The processing process on each machine represents a process step. The processing order of each process step in each job is fixed and must be processed on the machines in a certain order. In each flow shop, all jobs need to be processed according to the machine order and along the same path. During the assembly phase, there are s products and one assembly machine. Each product has a defined set of tasks. Assembly of a product can only begin when all the tasks it contains are completed and the assembly machine is idle. Tasks belonging to the same product are processed continuously. The next product's process can only begin after all the tasks belonging to that product have been completed. The environmental parameters include the following:
1. Basic parameters: n: Number of assignments m: Number of machines F: Number of factories s: Number of products j: Machine index i: Job index, i∈1,2,...,n; f: Factory index, f∈1,2,...,F; h,z: Product index, h,z∈1,2,…,s; k: Location index. In the scheduling sequence of factory f, k = 1, 2, ..., n represent the first, second, ..., n jobs respectively. M = {M1, M2, ..., M} m }: Represents the set of all machines; J = {J1, J2, ..., J} n }: A set of n jobs to be processed; P = {P1, P2, ..., P} s }: Represents a collection of products; n f The number of jobs assigned to factory f; n h : indicates that the product P is composed of h The number of jobs; Product P h It consists of multiple tasks; O i,j Homework J i The j-th operation; P i,j Homework J i In machine M j Processing time; t h Product P h Assembly time; c h Product P h The number of tasks that have been completed; r h Product P h The maximum completion time of the completed tasks, i.e., the current ready time; C j (f): The completion time of the last job on machine j in factory f; l i : The number of operations completed in operation i, i.e., the number of machines used in operation i. i ∈{0,1,2,…,m} 2. Auxiliary variables: π f π is the set of processing sequences within factory f. f =(π) t,1 ,π t,2 ,…,π t,n ); C k,j,f The task is located at the k-th position in factory f, and is performed on machine M. j The time required to complete the processing; π h Product P h The set of scheduling sequences for the middle job, π h =(π) h,1, π h,2 ,..,π h,n ); D πh Product P h The final completion moment; Product P h Homework J i The processing completion time; R h Product P h The time between the completion of all processing operations and the start of assembly represents the product P. h Release time; δ j,f If machine M j If the element belongs to factory f, the value is 1; otherwise, it is 0. TF: Total traffic time for all products; p f p represents the set of processing times for each job finally assigned in factory f on each machine. f =(p f,1 ,p f,2 ,…,p f,m ); v h Product number in sequence; 3. Decision variables: X i,k,f If homework J i If a component occupies position k in factory f, its value is 1; otherwise, it is 0. Y h,z If product P h Must be ranked in product P z If the preceding product is used as its predecessor, the value is 1; otherwise, it is 0. The goal of workshop scheduling is to minimize the total flow time, as shown in formula (1): The constraint formulas for the constraints of workshop scheduling include equations (2) to (10); Constraint (2) indicates that each job J i It must be assigned to a specific position k in a factory f; Constraint (3) states that each position k in each factory f can only be occupied by one job, and each machine M j Only one job can be processed at a time; Constraint (4) represents the computational task J. i Machine M in factory f j The completion time of the k-th position; Constraints (5) and (6) indicate that each product must have only one predecessor and one successor; Constraint (7) indicates that product P h The release time is equal to the completion time of its last task; Constraint (8) indicates that it is product P h Completion time; Constraint (9) means ensuring that the product's job completion time and assembly time are positive values; Constraint (10) defines the range of values for the decision variables.
3. The distributed assembly permutation flow shop scheduling method based on deep reinforcement learning as described in claim 1, characterized in that... In step 2, the state space includes five state values: job completion state, machine state, job scheduling sequence, product assembly state, and pre-constraint state, which are represented by equations (11)-(15), respectively. S5=(v1,v2,…,v s ) (15) Equation (11) records the completion time of the three most recently assigned jobs in each factory f on each machine j. The positions of the three most recently assigned jobs represent the positions of the last, second to last, and third to last jobs assigned on the machine, respectively. If a job does not exist, it is filled with 0. Equation (12) shows the operating status of each machine, indicating whether it is idle or busy. status j∈{0,1} indicates whether machine j is busy. The value is 0 when idle and 1 when busy. Equation (13) describes the scheduling sequence of operations in each factory; Status (14) indicates the completion status of each product; State (15) reflects the prerequisite relationships between products, and the state space can be represented as: S = {S1,S2,S3,S4,S5}.
4. The distributed assembly permutation flow shop scheduling method based on deep reinforcement learning as described in claim 1, characterized in that... The action space includes two parts: factory allocation rules and job sorting rules within the factory. First, a factory is selected according to the factory allocation rules, and the jobs to be processed are inserted. Then, the processing order of all jobs within the factory is rearranged according to the job sorting rules. The factory allocation rules include the following five: S1 Factory Integrated Load Balancing Rule FCLB: Selects the appropriate factory to allocate new products based on the current load of each factory. When allocating tasks, it considers three key indicators:
1. Machine busyness, which is the sum of the remaining available time of all machines in the factory. The busier the machines are, the higher the score; 2. Order backlog, which is the higher the score if there are more orders currently in the queue.
3. The more identical products there are in the queue, the higher the score. The factory's total load score = equipment busyness score + order backlog score + same product accumulation score; By calculating the total load score for each factory, the factory with the lowest total load score is finally determined as a candidate; during runtime, all factories with the same and lowest total load score will be filtered, and one of them will be randomly selected for task assignment. S2 Factory Integrated Capacity Balancing Rule (FCCB): Selects a suitable factory for allocating specific products based on the factory's production capacity. When allocating tasks, it considers three key indicators:
1. Immediately available capacity; the more currently idle machines, the higher the score.
2. The more machines that have recently released production capacity and are predicted to be idle in the near future, the higher the score; 3. The more similar products are being processed, the more points will be deducted. The overall capacity score is calculated as follows: score of immediately available capacity + score of recently released capacity + deduction for suppressed similar products. The factory with the highest overall capacity score is selected as a candidate. During operation, all factories with the same and highest overall capacity score will be screened, and one of them will be randomly selected for task assignment. S3 Future Time Window Load Balancing Rule (FTLB): The core of this factory selection strategy is to predict the processing capacity of each factory within a future time period △T and select the factory most likely to process new jobs quickly. The specific calculation involves three steps: First, determine the time window range as the current time T + time period △T; Then, two key indicators are calculated for each factory:
1. Number of available machines: the more machines that are idle within the time window, the higher the score; 2. Queue load: reflecting the current queue pressure: the more jobs currently being processed and those in the queue, the higher the score. Load balancing score = score for available machines + deduction for queue load; The factory with the highest load balancing score is ultimately selected; if multiple factories have the same score, a random selection is made. S4 Shortest Queue Load Balancing Rule SQ: Calculate the current number of job orders for each factory, find the factory with the shortest queue length, where the queue length is the total number of jobs currently being processed and those in the queue; then randomly select one factory from those with the same queue length as a candidate factory for task allocation. S5 Random Factory Rule: Randomly select a factory.
5. A distributed assembly permutation flow shop scheduling method based on deep reinforcement learning as described in claim 4, characterized in that... The job sorting rules include the following four:
1. Shortest Processing Time Rule (SPT): Sort the job list and, based on the processing time of each job in its current process, place the job with the shortest processing time at the top.
2. Longest Processing Time Rule (LPT): Sort the job list and, based on the processing time of each job in its current process, place the job with the longest processing time at the top.
3. Bottleneck Process Shortest Processing Time Rule (BSPT): The bottleneck process is the process with the longest production cycle time in the production process of a certain product. The index of the bottleneck process is calculated in advance, the job list is sorted, and the job with the shortest processing time is placed first according to the processing time of each job in the bottleneck process.
4. Bottleneck Remaining Processing Time Priority Rule (BRPT): First, calculate the remaining waiting time of each job at the bottleneck process. For jobs that have completed the bottleneck process, their priority is 0. For jobs that have not yet reached the bottleneck process, they are sorted according to their processing time at the bottleneck process and their distance from the bottleneck process. The closer the job is to the bottleneck, the higher its priority.
6. The distributed assembly permutation flow shop scheduling method based on deep reinforcement learning as described in claim 1, characterized in that... The reward mechanism includes the following aspects:
1. Dual Makespan Incentive Mechanism: The total flow time Makespan is the total flow time TF of all products. A baseline value for Makespan is set. When the actual Makespan is lower than the baseline, a positive reward is given; otherwise, a penalty is given.
2. Continuous improvement mechanism: A reward is given if the current total flow time Makespan is shorter than the previous one; 3. Intelligent load assessment mechanism: Comprehensive load value = 0.7 * current number of operations in the factory + 0.3 * total available time of all machines in the factory. Calculate the standard deviation of the comprehensive load value. The smaller the standard deviation, the more balanced the load, and the smaller the penalty.
4. Auxiliary Penalty Optimization: Auxiliary penalty optimization: Set delivery time penalty and assembly delay penalty. For the current job, if its completion time exceeds the scheduled delivery time, or if the assembly time is inconsistent with the completion time, penalties will be applied respectively. The greater the total number of delays for all jobs, the greater the penalty applied to the delivery time. When the completion time of the current job is inconsistent with the assembly time of the product belonging to this job, the greater the absolute value of the difference between the two, the greater the penalty applied. All rewards and penalties are added together to form the total reward, total_reward. According to the above process, the reward is an evaluation of the overall scheduling result of the PPO algorithm selecting the corresponding scheduling rule. The higher the reward, the better the evaluation. When the same single dynamic event occurs, the corresponding scheduling rule will appear more frequently through the output of the PPO algorithm.
7. The distributed assembly permutation flow shop scheduling method based on deep reinforcement learning as described in claim 1, characterized in that... In step 3, the information of the job and machine is updated. Whenever an action is selected, the start time, actual processing time, and end time of the selected job and machine are first confirmed. The start time is the greater of the completion time of the same job on the previous machine and the completion time of the previous job on the same machine, which is max(C). k,j-1,f C k-1,j,f The processing end time is max(C). k,j-1,f C k-1,j,f )+P i,j Subsequently, update the machine's processing time C. j(f) Number of completed assignments: l i , Number of unfinished assignments (ml) i Scheduling sequence X i,k,f Whenever a job completes a process, the number of completed processes is incremented by 1, and the number of remaining unfinished jobs is decremented by 1. At this time, it is determined whether the number of processes in the current job exceeds the total number of processes m that the job should have. If it does, the set J of unfinished processes of the job is deleted. Finally, update the product status. Once all tasks related to this product have been completed, the product can be assembled. The earliest assembly start time is the completion time of the last task belonging to this product. Calculate the total flow time TF after all products are assembled.
Citation Information
Cited By
Batch delivery distributed assembly intelligent scheduling method based on reinforcement learning
CN122288252A