Hierarchical dynamic scheduling method based on strategy optimization DDQN algorithm

By adopting a hierarchical dynamic scheduling method based on strategy optimization in a reentrant hybrid flow workshop with batch machine, the problem of insufficient scheduling efficiency and optimization under multi-batch stages and dynamic disturbances is solved, and efficient workpiece scheduling and machine resource utilization are achieved.

CN119987302AActive Publication Date: 2025-05-13SOUTHWEST JIAOTONG UNIV

Patent Information

Application Number
CN202510060040.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-13
Estimated Expiration
2045-01-15

AI Technical Summary

Technical Problem

The prior art is difficult to effectively solve the dynamic scheduling problem of reentrant and hybrid flow workshops with batch machines, especially in the case of multi-batch processing stages and dynamic disturbances, and the scheduling efficiency and optimization are insufficient.

Method used

A DDQN algorithm based on strategy optimization is adopted to construct a hierarchical dynamic scheduling method, and the artifact batch and scheduling are optimized through collaborative decision-making between batch agents and scheduling agents, combined with Markov decision-making process model and heuristic variable threshold batch strategy.

Benefits of technology

It realizes efficient scheduling in multi-batch stages and dynamic disturbances, significantly improves workpiece completion time, reduces weighted total drag period, and maximizes the average machine utilization rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119987302A_ABST
    Figure CN119987302A_ABST
Patent Text Reader

Abstract

The invention discloses a hierarchical dynamic scheduling method based on a strategy optimization DDQN algorithm, which is used for solving a dynamic scheduling problem of a reentrant hybrid flow shop with a batch processor, and comprises the following steps: firstly, determining an objective function and a constraint condition of the scheduling problem; secondly, by introducing a self-attention mechanism, providing a layered structure based on the self-attention mechanism DDQN, and constructing two agents, namely a batching agent and a scheduling agent, to respectively solve batching sub-problems and scheduling sub-problems in the problem; besides, in order to cope with the multi-stage batch processing and reentry scheduling characteristics in the problem, a Markov decision process considering the characteristics of the two agents is designed, and the Markov decision process comprises states, actions and rewards. Furthermore, an action selection strategy and a soft start target network updating strategy based on a mask combined with an epsilon-greedy strategy are provided, so that the efficiency and generalization ability are improved. The method has a remarkable effect on solving the dynamic scheduling problem of the reentrant hybrid flow shop with the batch processor.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of hybrid flow shop scheduling optimization, and in particular relates to a hierarchical dynamic scheduling method based on a strategy-optimized DDQN algorithm. Background Art

[0002] The reentrant hybrid flow shop scheduling problem with batch processing machines is a classic extension of the hybrid flow shop problem, which is closer to production and manufacturing scenarios such as semiconductor manufacturing, blade coating, and steel pipe cold drawing. In recent years, it has also attracted attention due to the simultaneous consideration of batch processing and reentrant constraints. The existing research only considers the problem of a single batch processing stage, but there are multiple batch processing stages in the production scenario. For example, the oxidation and diffusion processes in semiconductor manufacturing are both batch processing stages, and they are not necessarily processed continuously, so it is necessary to study the reentrant hybrid flow shop problem with multiple batch processing stages. In addition, there are many dynamic disturbances in the actual reentrant hybrid flow shop scheduling problem with batch processing machines, such as the arrival of new workpieces, so the problem is expanded to a dynamic reentrant hybrid flow shop scheduling problem with batch processing machines.

[0003] The commonly used dynamic scheduling methods in workshops include priority rules, metaheuristic methods, and deep reinforcement learning. Priority rules and metaheuristic methods each have some shortcomings in dynamic scheduling problems. Priority rules have high efficiency, but do not have re-optimization and global optimization; metaheuristic optimization takes a long time; deep reinforcement learning can make decisions in a short time while ensuring current and long-term scheduling performance by learning the interaction between the control decision model and the scheduling environment. In particular, the DDQN series of deep reinforcement learning methods have a dual Q network structure and experience replay mechanism, which can effectively solve the problems of high dimensional state space, complex action selection, and high real-time requirements in workshop dynamic scheduling problems. Summary of the invention

[0004] In order to overcome the shortcomings of the prior art in solving the dynamic scheduling problem of a reentrant hybrid flow shop with batch processing machines, the present invention provides a hierarchical dynamic scheduling method based on a strategy-optimized DDQN algorithm.

[0005] The hierarchical dynamic scheduling method based on the strategy-optimized DDQN algorithm of the present invention is used to solve the dynamic scheduling problem of a reentrant hybrid flow shop with a batch processor, specifically:

[0006] A. Description and assumptions of the scheduling problem.

[0007] There are n initial and random arrival workpieces of Q types. Each workpiece goes through 1 stage according to its process route, which includes |BS| batch processing stages, 1≤|BS|≤l; there are m s unrelated parallel machines, ms ≥1, the batch processing machines included in the batch processing stage can simultaneously process a limited number of workpieces not exceeding the machine capacity, and the single processing machines included in the single processing stage are single-piece processing; each process workpiece can be processed by a machine in this stage, and there is no priority constraint on the processing sequence between workpieces, but there is a sequence constraint on the processing sequence between each workpiece process. The workpiece may need to be repeatedly processed multiple times at a workstation, and the workpiece is allowed to skip certain stages of processing when it is visited at different times.

[0008] B. Objective function and constraints.

[0009] Objective function:

[0010] Minimizing the completion time, minimizing the weighted total tardiness, and maximizing the average machine utilization are expressed as:

[0011]

[0012] Wherein, formula (1) represents the completion time of all workpieces, j represents the workpiece index, and o j represents the last process of workpiece j, C j,oj represents the completion time of job j. Formula (2) represents the weighted total delay of all jobs, w j represents the urgency of each workpiece j, D j represents the expiration time of each workpiece j. Formula (3) represents the average machine utilization rate, o represents the process index of the workpiece, s is the index of the stage, k is the index of the machine, and m s is the number of machines in stage s, M s is the set of machines in stage s, P j,o,s,k represents the processing time of process o of workpiece j on machine k at stage s, Y j,o,s,k Indicates whether the process o of workpiece j is processed on machine k at stage s. If it is Y j,o,s,k =1, otherwise Y j,o,s,k =0.

[0013] Constraints:

[0014]

[0015] Σ j,j'∈b Z j,j' =1 (11)

[0016]

[0017]

[0018] Among them, formula (4) indicates that all workpieces j must be processed after they arrive, S j,1represents the start time of process 1 of workpiece j, A j represents the arrival time of workpiece j. Formula (5) indicates that each process of each workpiece can only be processed on one machine in one stage. Formula (6) indicates that in the single processing stage, each machine can only process one workpiece at a time, S j,o represents the start processing time of process o of workpiece j, X jo,j’o’,sk Indicates whether the process o of workpiece j is processed before the process o' of workpiece j'. If so, X jo,j’o’,sk =1, otherwise X jo,j’o’,sk =0, N is a sufficiently large positive number, O s represents the set of processes that need to be processed at stage s. Formula (7) represents the ordering constraint between processes. For two operations of the same job with sequence constraints, the latter must be processed after the former is processed. Formula (8) represents the ordering constraint between job j cycles. The job accessed in the latter cycle must be processed after the previous cycle is processed. S j,r represents the starting processing time of the rth cycle of workpiece j, C j,r represents the end processing time of the rth cycle of workpiece j. Formula (9) indicates that batch processing is performed before processing to ensure that each job is assigned to only one batch during batch processing. b represents the batch index, and B s,k represents the set of batches processed on machine k in stage s, α j,b Indicates whether workpiece j is assigned to batch b. If so, α j,b =1, otherwise α j,b = 0. Formula (10) indicates that a batch can only be processed on one batch processing machine, BS represents the batch processing stage set, β b,s,k Indicates whether batch b is processed on machine k at stage s. If yes, β b,s,k =1, otherwise, β b,s,k = 0. Formula (11) indicates that only workpieces j belonging to the same process type can form a batch b, Z j,j' Indicates whether workpiece j and workpiece j' belong to the same process type. If yes, Z j,j' =1, otherwise Z j,j' = 0. Equation (12) represents the capacity limit of the batch processing machine, ensuring that the total number of jobs in the batch processed by the machine does not exceed its maximum capacity, V s,k represents the processing capacity of machine k in stage s. Equation (13-14) represents the order priority constraint, which ensures that different batches processed by the same batch processing machine follow the order, S b,k represents the start time of processing of batch b on machine k, C b,k represents the end processing time of batch b on machine k, γ b,b' Indicates whether batch b is processed before batch b'. If so, γb,b' =1; otherwise γ b,b' =0.

[0019] C. A hierarchical dynamic scheduling method based on strategy-optimized DDQN (HSDDQN) is used to solve the dynamic scheduling problem of a reentrant hybrid flow shop.

[0020] Firstly, a hierarchical structure based on the policy optimization DDQN algorithm is proposed, and a batching agent (BA) and a scheduling agent (SA) are constructed to solve the batching and scheduling sub-problems of workpieces in the dynamic scheduling problem of a reentrant hybrid flow shop with batch processing machines. The batching agent BA decides whether to execute the job batch, while the scheduling agent SA determines the priority order of jobs or batches and the machine selection.

[0021] Secondly, based on the above two agents, a Markov decision process model suitable for the dynamic scheduling problem of a reentrant hybrid flow shop with batch processing machines is designed, which combines multi-stage batch processing and reentrancy characteristics, defines states, rewards and actions, and combines the heuristic variable threshold batch policy (VTBP) with dynamic workpiece batching.

[0022] Markov Decision Process Model:

[0023] (1) Status:

[0024] Considering the stage processing characteristics of the hybrid flow shop, the characteristics of each agent are designed according to the processing stage at the decision time point, and 6 state features are extracted as the input of BA, which are:

[0025]

[0026] 12 state features are extracted as the input of SA, specifically:

[0027]

[0028]

[0029] (2) Action:

[0030] The heuristic variable threshold batch policy VTBP is used to construct the action space of the batching agent. The VTBP strategy is as follows: first, determine the process type of the job, then determine the capacity threshold of the batch, then sort the jobs according to certain rules, and finally form a batch by selecting the threshold number of jobs; when scheduling, all jobs in the processing stage buffer are sorted in ascending order according to the time they arrived at the stage and classified by job type; then, calculate the average waiting time of each type of job; select the type q with the longest average waiting time for batch processing; next, it is necessary to determine the threshold V of the current batch b ; Let Os,q,t represents the set of job operations of type q that need to be batch processed at stage s but have not yet arrived at the current time t, BN s,q Indicates the number of type q jobs in the current stage s; if O s,q,t ≠0, indicating that there are still unfinished batch processing operations in the current stage; in this case, if BN s,q >|O s,q,t |, then V b =min(BN s,q ,V s,k ); if BN s,q ≤|O s,q,t |, then V b =0; otherwise, if O s,q,t =0, then V b =min(BN s,q ,V s,k ).

[0031] The batch schemes built based on different batch rules directly affect the optimization results. When batch processing jobs in the buffer, the consistency of processing time is given priority, that is, jobs with similar processing time are grouped into the same batch, thereby reducing the waste of processing time. In addition, in order to minimize the total delay, completion time and reduce the waiting time of jobs in the buffer, three batch rules are proposed to balance the consistency of processing time with other time-based performance indicators. A weighted batch quality index is introduced. This index takes into account the consistency of delivery time and processing time of jobs in the buffer, as shown in formula (15):

[0032]

[0033] In the formula, buffer s represents the set of jobs in the buffer of stage s during the batch scheduling decision process, q j =q j' It means that only jobs belonging to the same job family can be grouped for batch processing, θ is the weight coefficient, P j,o,s,k represents the processing time required for the current batch operation of job j on machine k, D j represents the delivery time of job j, c j represents the remaining shortest processing time of job j, A s,j represents the arrival time of job j at stage s; batch quality index The smaller the size, the higher the priority of the job in batch processing; jobs are processed in order of Sort in ascending order to determine the batch order.

[0034] The action space of BA is as follows:

[0035]

[0036]

[0037] The scheduling point is set as the machine idle time; the action space of SA includes job selection and machine selection; in the batch processing stage, SA schedules the batches formed by BA; in the individual processing stage, the job selection rules and machine selection rules are as follows:

[0038]

[0039] (3) Rewards:

[0040] The total weighted delay of the workpiece is subdivided into each scheduling decision of the decision model BA. After the decision model BA executes each step, an immediate reward will be generated to reflect the impact of the action. The immediate reward of BA at time t is shown in formula (16). In order to make the two goals of minimizing the maximum completion time and maximizing the machine utilization rate be reflected in the actions of the same decision model, the difference between the average machine utilization rate and the completion time difference between each step of scheduling is taken as the immediate reward of the decision model SA. The immediate reward of SA at time t is shown in formula (17). In order to ensure that the average machine utilization rate and the completion time difference between each step of scheduling have the same order of magnitude in guiding the behavior of the decision model, α1 and α2 are introduced.

[0041]

[0042] r SA =α1[U ave (t')-U ave (t)]-α2[C step (t')-C step (t)] (17)

[0043] in, represents the number of jobs in the job pool of stage p at time t, represents the number of workpieces in batch b, t′ represents the next decision time. If the expiration date of the workpiece is less than the next decision time, Z is 1, otherwise Z is 0. If the workpiece completes the processing of this process at the next decision time, t m is equal to the completion time of the process of the workpiece, otherwise t m Equal to the next decision time t′.

[0044] Finally, in order to further enhance the effectiveness and generalization ability of the hierarchical dynamic scheduling method, two improvement strategies are proposed:

[0045] a) The mask-based action selection strategy is combined with the ε-greedy strategy to restrict the action space of each agent at different decision points and improve the efficiency of collaborative decision-making.

[0046] b) A soft-start network update strategy is used for offline training to accelerate learning and convergence.

[0047] The beneficial technical effects of the present invention are:

[0048] 1. This paper aims at the dynamic scheduling problem of reentrant hybrid flow shop with batch processing machines, and constructs a dynamic scheduling model, aiming to improve the scheduling effect by minimizing the completion time, total delay and maximizing the average machine utilization. The model fully considers the arrival of new orders, the restrictions of incompatible workpiece families and the capacity characteristics of batch processing machines, and constructs a comprehensive scheduling framework. In addition, considering the complexity and dynamic characteristics of the problem, a corresponding Markov decision model is designed. The state, action space and reward mechanism of the intelligent agent are defined in the model, and an efficient scheduling decision process is achieved by constructing a highly adaptable model network.

[0049] 2. The dynamic scheduling model and Markov decision model proposed in the present invention can more accurately reflect the dynamic changes of the workshop environment, especially under complex conditions such as the arrival of new orders and the incompatibility of workpiece families, and can achieve more accurate scheduling optimization. Secondly, the dual-depth Q network based on the self-attention mechanism can effectively handle complex scheduling problems, and significantly improve the efficiency and accuracy of scheduling decisions through collaboration and interaction between intelligent agents. Finally, through the heuristic variable threshold batching strategy and intelligent agent improvement strategy, not only the dynamics of workpiece batching is improved, but also the collaborative ability, generalization ability and decision-making accuracy of the intelligent agent are enhanced, thereby further optimizing the system performance and solving the multi-objective optimization problems that are difficult to deal with by traditional scheduling methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 This is the principle framework of the hierarchical dynamic scheduling method based on strategy optimization DDQN of the present invention.

[0051] Figure 2 It is the network structure of intelligent agent. DETAILED DESCRIPTION

[0052] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0053] The hierarchical dynamic scheduling method based on the strategy-optimized DDQN algorithm of the present invention is used to solve the dynamic scheduling problem of a reentrant hybrid flow shop with a batch processor, specifically:

[0054] A. Description and assumptions of the scheduling problem.

[0055] There are n initial and random arrival workpieces of Q types. Each workpiece goes through 1 stage according to its process route, which includes |BS| batch processing stages, 1≤|BS|≤l; there are m s unrelated parallel machines, m s ≥1, the batch processing machines included in the batch processing stage can simultaneously process a limited number of workpieces not exceeding the machine capacity, and the single processing machines included in the single processing stage are single-piece processing; each process workpiece can be processed by a machine in this stage, and there is no priority constraint on the processing sequence between workpieces, but there is a sequence constraint on the processing sequence between each workpiece process. The workpiece may need to be repeatedly processed multiple times at a workstation, and the workpiece is allowed to skip certain stages of processing when it is visited at different times.

[0056] B. Objective function and constraints.

[0057] Objective function:

[0058] Minimizing the completion time, minimizing the weighted total tardiness, and maximizing the average machine utilization are expressed as:

[0059]

[0060] Wherein, formula (1) represents the completion time of all workpieces, j represents the workpiece index, and o j represents the last process of workpiece j, C j,oj represents the completion time of job j. Formula (2) represents the weighted total delay of all jobs, w j represents the urgency of each workpiece j, D j represents the expiration time of each workpiece j. Formula (3) represents the average machine utilization rate, o represents the process index of the workpiece, s is the index of the stage, k is the index of the machine, and m s is the number of machines in stage s, M s is the set of machines in stage s, P j,o,s,k represents the processing time of process o of workpiece j on machine k at stage s, Y j,o,s,k Indicates whether the process o of workpiece j is processed on machine k at stage s. If it is Y j,o,s,k =1, otherwise Y j,o,s,k =0.

[0061] Constraints:

[0062]

[0063] ∑ j,j'∈b Z j,j' =1(11)

[0064]

[0065]

[0066] Among them, formula (4) indicates that all workpieces j must be processed after they arrive, S j,1 represents the start time of process 1 of workpiece j, A j represents the arrival time of workpiece j. Formula (5) indicates that each process of each workpiece can only be processed on one machine in one stage. Formula (6) indicates that in the single processing stage, each machine can only process one workpiece at a time, S j,o represents the start processing time of process o of workpiece j, X jo,j’o’,sk Indicates whether the process o of workpiece j is processed before the process o' of workpiece j'. If so, X jo,j’o’,sk =1, otherwise X jo,j’o’,sk =0, N is a sufficiently large positive number, O s represents the set of processes that need to be processed at stage s. Formula (7) represents the ordering constraint between processes. For two operations of the same job with sequence constraints, the latter must be processed after the former is processed. Formula (8) represents the ordering constraint between job j cycles. The job accessed in the latter cycle must be processed after the previous cycle is processed. S j,r represents the starting processing time of the rth cycle of workpiece j, C j,r represents the end processing time of the rth cycle of workpiece j. Formula (9) indicates that batch processing is performed before processing to ensure that each job is assigned to only one batch during batch processing. b represents the batch index, and B s,k represents the set of batches processed on machine k in stage s, α j,b Indicates whether workpiece j is assigned to batch b. If so, α j,b =1, otherwise α j,b = 0. Formula (10) indicates that a batch can only be processed on one batch processing machine, BS represents the batch processing stage set, β b,s,k Indicates whether batch b is processed on machine k at stage s. If yes, β b,s,k =1, otherwise, β b,s,k = 0. Formula (11) indicates that only workpieces j belonging to the same process type can form a batch b, Z j,j' Indicates whether workpiece j and workpiece j' belong to the same process type. If yes, Z j,j' =1, otherwise Z j,j' = 0. Equation (12) represents the capacity limit of the batch processing machine, ensuring that the total number of jobs in the batch processed by the machine does not exceed its maximum capacity, V s,k represents the processing capacity of machine k in stage s. Equation (13-14) represents the order priority constraint, which ensures that different batches processed by the same batch processing machine follow the order, S b,krepresents the start time of processing of batch b on machine k, C b,k represents the end processing time of batch b on machine k, γ b,b' Indicates whether batch b is processed before batch b'. If so, γ b,b' =1; otherwise γ b,b' =0.

[0067] C. Hierarchical dynamic scheduling method based on strategy optimization DDQN (HSDDQN) solves the dynamic scheduling problem of reentrant hybrid flow shop. The principle of hierarchical dynamic scheduling method based on strategy optimization DDQN is as follows Figure 1 shown.

[0068] Firstly, a hierarchical structure based on the policy optimization DDQN algorithm is proposed, and a batching agent (BA) and a scheduling agent (SA) are constructed to solve the batching and scheduling sub-problems of workpieces in the dynamic scheduling problem of a reentrant hybrid flow shop with batch processing machines. The batching agent BA decides whether to execute the job batch, while the scheduling agent SA determines the priority order of jobs or batches and the machine selection.

[0069] Secondly, based on the above two agents, a Markov decision process model suitable for the dynamic scheduling problem of a reentrant hybrid flow shop with batch processing machines is designed, which combines multi-stage batch processing and reentrancy characteristics, defines states, rewards and actions, and combines the heuristic variable threshold batch policy (VTBP) with dynamic workpiece batching.

[0070] Markov Decision Process Model:

[0071] (1) Status:

[0072] Considering the stage processing characteristics of the hybrid flow shop, the characteristics of each agent are designed according to the processing stage at the decision-making time point, and 6 state features are extracted as the input of BA. The detailed description of each state feature is shown in Table 1.

[0073] Table 1 BA state design

[0074]

[0075] 12 state features are extracted as the input of SA, and the detailed description of each state feature is shown in Table 2.

[0076] Table 2 SA state design

[0077]

[0078]

[0079] (2) Action:

[0080] The heuristic variable threshold batch policy VTBP is used to construct the action space of the batching agent. The VTBP strategy is as follows: first, determine the process type of the job, then determine the capacity threshold of the batch, then sort the jobs according to certain rules, and finally form a batch by selecting the threshold number of jobs; when scheduling, all jobs in the processing stage buffer are sorted in ascending order according to the time they arrived at the stage and classified by job type; then, calculate the average waiting time of each type of job; select the type q with the longest average waiting time for batch processing; next, it is necessary to determine the threshold V of the current batch b ; Let O s,q,t represents the set of job operations of type q that need to be batch processed at stage s but have not yet arrived at the current time t, BN s,q Indicates the number of type q jobs in the current stage s; if O s,q,t ≠0, indicating that there are still unfinished batch processing operations in the current stage; in this case, if BN s,q >|O s,q,t |, then V b =min(BN s,q ,V s,k ); if BN s,q ≤|O s,q,t |, then V b =0; otherwise, if O s,q,t =0, then V b =min(BN s,q ,V s,k ).

[0081] The batch schemes built based on different batch rules directly affect the optimization results. When batch processing jobs in the buffer, the consistency of processing time is given priority, that is, jobs with similar processing time are grouped into the same batch, thereby reducing the waste of processing time. In addition, in order to minimize the total delay, completion time and reduce the waiting time of jobs in the buffer, three batch rules are proposed to balance the consistency of processing time with other time-based performance indicators. A weighted batch quality index is introduced. This index takes into account the delivery time consistency and processing time consistency of jobs within the buffer, as shown in formula (15).

[0082] Decision makers can adjust the weight coefficients to align with their preferred optimization goals.

[0083]

[0084] In the formula, buffer s represents the set of jobs in the buffer of stage s during the batch scheduling decision process, q j =q j'It means that only jobs belonging to the same job family can be grouped for batch processing, θ is the weight coefficient, P j,o,s,k represents the processing time required for the current batch operation of job j on machine k, D j represents the delivery time of job j, c j represents the remaining shortest processing time of job j, A s,j represents the arrival time of job j at stage s; batch quality index The smaller the size, the higher the priority of the job in batch processing; jobs are processed in order of The action space of BA is shown in Table 3.

[0085] Table 3 Action space of BA

[0086]

[0087] When V b =0, BA selects the "None" action and the job will continue to wait;

[0088] The scheduling point is set to the machine idle time; the action space of SA includes job selection and machine selection; in the batch processing stage, SA schedules the batches formed by BA. In the individual processing stage, the job selection rules and machine selection rules are shown in Table 4.

[0089] Table 4 Action space of SA

[0090]

[0091] (3) Rewards:

[0092] The total weighted delay of the workpiece is subdivided into each scheduling decision of the decision model BA. After the decision model BA executes each step, an immediate reward will be generated to reflect the impact of the action. The immediate reward of BA at time t is shown in formula (16). In order to make the two goals of minimizing the maximum completion time and maximizing the machine utilization rate be reflected in the actions of the same decision model, the difference between the average machine utilization rate and the completion time difference between each step of scheduling is taken as the immediate reward of the decision model SA. The immediate reward of SA at time t is shown in formula (17). In order to ensure that the average machine utilization rate and the completion time difference between each step of scheduling have the same order of magnitude in guiding the behavior of the decision model, α1 and α2 are introduced.

[0093]

[0094] r SA =α1[U ave (t')-U ave (t)]-α2[C step (t')-C step(t)] (17)

[0095] in, represents the number of jobs in the job pool of stage p at time t, represents the number of workpieces in batch b, t′ represents the next decision time. If the expiration date of the workpiece is less than the next decision time, Z is 1, otherwise Z is 0. If the workpiece completes the processing of this process at the next decision time, t m is equal to the completion time of the process of the workpiece, otherwise t m Equal to the next decision time t′.

[0096] Finally, in order to further enhance the effectiveness and generalization ability of the hierarchical dynamic scheduling method, two improvement strategies are proposed:

[0097] a) The mask-based action selection strategy is combined with the ε-greedy strategy to restrict the action space of each agent at different decision points and improve the efficiency of collaborative decision-making.

[0098] In this problem, BA needs to determine the processing mode of the current stage and select actions from different available action sets according to the mode. When the processing stage is in batch mode, BA can select appropriate job batch rules. However, in a single processing stage, batch rules cannot be included in the set of available actions. In order to reduce the dimension of the action space in complex environments and help the decision model learn task knowledge more effectively and efficiently, a mask mechanism is adopted. This mechanism divides the overall action space into multiple regions to form specific action spaces for different states.

[0099] The ultimate learning goal of the decision agent is the distribution of action scores. Simply preventing the decision agent from selecting inappropriate actions does not accelerate the convergence process of the agent. Therefore, before selecting an action, a mask strategy is usually added to the value function vector, and then the masked value is passed into the neural network for training, combined with the ε-greedy strategy in the DDQN algorithm. The action constraint selection method combining the mask mechanism and the ε-greedy strategy is shown in Algorithm 1.

[0100]

[0101]

[0102] b) A soft-start network update strategy is used for offline training to accelerate learning and convergence.

[0103] This paper proposes a soft start target network update strategy, which adopts progressive update. In the early stage of training, the target network update frequency is high, and as the training progresses, the update frequency gradually decreases. The soft start strategy adjusts the target network update frequency according to the training progress, and more flexibly adapts to the training needs of different stages, thereby accelerating the convergence of the agent and reducing the instability of the agent. The pseudo code of the soft start strategy is shown in Algorithm 2, where C = 20, C f =200, δ=random(0.75,1).

[0104]

[0105] Network structure:

[0106] Fully connected networks are often widely used in scheduling problems because they can flexibly model complex environments and tasks. However, traditional network structures often have bloated architectures and are prone to falling into local optimality, especially when the input feature dimension is high. Therefore, this description introduces a self-attention mechanism that enables the model to dynamically adjust its attention distribution, enhancing the flexibility and efficiency of processing information in complex environments. This description optimizes the network structure by combining a self-attention mechanism layer and a fully connected layer that gradually reduces the number of neurons, thereby improving the efficiency and accuracy of the HSDDQN algorithm. The network structure of each agent uses Relu as the activation function, such as Figure 2 shown.

[0107] Example:

[0108] 1. Experimental design:

[0109] The experiment was implemented in Python and ran on a computer equipped with an AMD Ryzen 7 5800H@3.20GHz CPU, 64GB memory, and an NVIDIA GeForce RTX 3070 graphics card. Since there are no existing benchmarks for the dynamic scheduling problem of a reentrant hybrid flow shop with batch machines, the present invention constructs a series of datasets to support the verification of HSDDQN and its improvements. The initial shop has multiple jobs, after which new jobs arrive according to a Poisson distribution, and the interval between the arrival of two consecutive new workpieces follows an exponential distribution "exp(1 / λ)". The training and test benchmarks used in the present invention are randomly generated based on the parameters listed in Table 5, where "randi" and "randf" represent the uniform distribution of integers and real numbers, respectively, and the total number of machines M is obtained by summing the number of machines in each stage. The parameter settings for the training process are listed in Table 6.

[0110] Table 5. Training and test case generation rules

[0111]

[0112] Table 6. Hyperparameters of training process

[0113]

[0114]

[0115] The embodiment uses two performance indicators, generation distance (GD) and inverse generation distance (IGD), to evaluate the quality of the solution obtained under the Pareto optimal solution. GD and IGD are defined as follows:

[0116]

[0117] Among them, d i,A,P represents the minimum Euclidean distance of the i-th solution in set P to the solution in set A. In the numerical experiments of the present invention, all Pareto optimal solutions obtained by multiple runs of HSDDQN and other comparative methods are merged, and then the set of non-dominated solutions is taken out as an approximation of the true Pareto front P. This method ensures that P can be regarded as an approximation of the true Pareto front, which is sufficient for performance comparison. In order to eliminate the influence of different dimensions, the objective function value is normalized.

[0118] 2. Comparison of algorithm results

[0119] (1) Comparison results with the governing rules

[0120] In order to verify the effectiveness and superiority of the proposed HSDDQN algorithm, the performance of the method is compared with known scheduling rules at a specific scale. The study selected the four most effective combinations of six job scheduling rules and two machine selection rules as benchmark rules. At the same time, in order to confirm whether the proposed HSDDQN has learned a feasible rule selection strategy, the random strategy is also compared, that is, the job scheduling rule and machine selection rule are randomly selected at each rescheduling point. Each method is repeated 20 times on each test instance, and the performance indicators of the final Pareto optimal frontier obtained by the comparison method are shown in Tables 7 and 8, where the best results are marked in bold.

[0121] Table 7 GD values ​​of HSDDQN compared with the rules

[0122]

[0123]

[0124] Table 8 IGD values ​​of HSDDQN compared with the rules

[0125]

[0126]

[0127] Compared with the random action selection strategy, the proposed HSDDQN achieves better results on all indicators in almost all instances. This proves that HSDDQN has learned an effective strategy for selecting feasible goals and scheduling rules at each rescheduling point. Next, compared with other known scheduling rules, HSDDQN can achieve the best results on GD and IGD indicators on 75% of the instances, and also achieves good results on 68.75% of the instances. This shows that the solution obtained by HSDDQN is closest to the true Pareto optimal front in terms of convergence.

[0128] (2) Comparison results with deep reinforcement learning

[0129] To evaluate the convergence performance and learning ability of the proposed HSDDQN, the present invention compares it with four dynamic deep reinforcement scheduling methods: HMAPPO, THDQN, MA-IDDQN and HRLDDQN. The comparison is carried out under the condition of consistent design of state, action and reward function. The purpose is to show the effectiveness of the proposed algorithm and environment settings. All deep reinforcement learning algorithms are trained using the same test cases. To evaluate the generalization ability of the trained HSDDQN on large-scale problems compared with other scheduling methods, we introduced new jobs of different sizes and increased the frequency of workpieces entering the production system. This resulted in higher work-in-process levels, increased preemption of production equipment, and more complex production scheduling scenarios. The parameter settings are consistent with those listed in Table 5. The experiment was repeated 20 times for each test instance. Tables 9-10 show the performance indicators of each method and highlight the best results.

[0130] Table 9 GD values ​​of HSDDQN compared with other deep reinforcement learning

[0131]

[0132]

[0133] Table 10 IGD values ​​of HSDDQN compared with other deep reinforcement learning

[0134]

[0135]

[0136] The results show that HSDDQN outperforms other methods in GD value in 85.4% of instances and in IGD metric in 69% of instances, indicating that HSDDQN achieves the best compromise among the three research objectives.

[0137] (3) Comparison results of workpiece batching strategies

[0138] In order to verify the effectiveness of the proposed variable threshold batching policy (VTBP), it is compared with the common fixed capacity threshold policy. The common fixed capacity threshold can be divided into fixed capacity threshold policy (FCTBP) and fixed time threshold policy (FTTBP). Two equipment capacity thresholds are set - 70% and 100%, denoted as FCTBP1 and FCTBP2; two time thresholds are set - 1*PT and 2*PT, denoted as FTTBP1 and FTTBP2 (PT is the average working time of all workpieces in the same batch process).

[0139] Tables 11 and 12 respectively show the GD index and IGD index of the four strategies on nine representative examples.

[0140] Table 11 Comparison of average GD indicators of strategies

[0141]

[0142]

[0143] Table 12 Comparison of average IGD indicators of strategies

[0144]

[0145] It can be seen from the experimental results that the VTBP strategy proposed in the present invention is 40.69% and 50.76% better than the FCTBP strategy and the FTTBP strategy in terms of GD index; and is 23.72% and 44.39% better than the FCTBP strategy and the FTTBP strategy in terms of IGD index, respectively, indicating that the solution obtained by the VTBP strategy is more advantageous than other strategies in terms of diversity and convergence.

Claims

1. A hierarchical dynamic scheduling method based on a policy-optimized DDQN algorithm, characterized in that: It is used to solve the dynamic scheduling problem of a reentrant hybrid flow shop with batch processing machines, specifically: A. Description and assumptions of the scheduling problem; There are n initial and random arrival workpieces of Q types. Each workpiece goes through 1 stage according to its process route, which includes |BS| batch processing stages, 1≤|BS|≤l; there are m s unrelated parallel machines, m s ≥1, the batch processing machines included in the batch processing stage can simultaneously process a limited number of workpieces not exceeding the machine capacity, and the single processing machines included in the single processing stage are single-piece processing; each process workpiece can be processed by a machine in this stage, and there is no priority constraint between the processing order of workpieces, but there is a priority constraint between the processing order of each workpiece process. The workpiece may need to be repeatedly processed multiple times at a station, and the workpiece is allowed to skip some stages of processing when the number of visits is different; B. Objective function and constraints; Objective function: Minimizing the completion time, minimizing the weighted total tardiness, and maximizing the average machine utilization are expressed as: Wherein, formula (1) represents the completion time of all workpieces, j represents the workpiece index, and o j represents the last process of workpiece j, C j,oj represents the completion time of job j; Formula (2) represents the weighted total delay of all jobs, w j represents the urgency of each workpiece j, D j represents the expiration time of each workpiece j; Formula (3) represents the average machine utilization rate, o represents the process index of the workpiece, s is the index of the stage, k is the index of the machine, and m s is the number of machines in stage s, M s is the set of machines in stage s, P j,o,s,k represents the processing time of process o of workpiece j on machine k at stage s, Y j,o,s,k Indicates whether the process o of workpiece j is processed on machine k at stage s. If it is Y j,o,s,k =1, otherwise Y j,o,s,k =0; Constraints: Among them, formula (4) indicates that all workpieces j must be processed after they arrive, S j,1 represents the start time of process 1 of workpiece j, A j represents the arrival time of workpiece j; Formula (5) indicates that each process of each workpiece can only be processed on one machine in one stage; Formula (6) indicates that in the single processing stage, each machine can only process one workpiece at a time, S j,o represents the start processing time of process o of workpiece j, X jo,j’o’,sk Indicates whether the process o of workpiece j is processed before the process o' of workpiece j'. If so, X jo,j’o’,sk =1, otherwise X jo,j’o’,sk =0, N is a sufficiently large positive number, O s represents the set of processes that need to be processed in stage s; Formula (7) represents the ordering constraint between processes. For two operations of the same job with sequence constraints, the latter must be processed after the former is processed; Formula (8) represents the ordering constraint between job j cycles. The job accessing the latter cycle must be processed after the previous cycle is processed. S j,r represents the starting processing time of the rth cycle of workpiece j, C j,r represents the end processing time of the rth cycle of workpiece j; Formula (9) indicates that batch processing is performed before processing to ensure that each job is assigned to only one batch during batch processing, b represents the batch index, B s,k represents the set of batches processed on machine k in stage s, α j,b Indicates whether workpiece j is assigned to batch b. If so, α j,b =1, otherwise α j,b =0; Formula (10) indicates that a batch can only be processed on one batch processing machine, BS represents the batch processing stage set, β b,s,k Indicates whether batch b is processed on machine k at stage s. If yes, β b,s,k =1, otherwise, β b,s,k = 0; Formula (11) indicates that only workpieces j belonging to the same process type can form a batch b, Z j,j' Indicates whether workpiece j and workpiece j' belong to the same process type. If yes, Z j,j' =1, otherwise Z j,j' = 0; Formula (12) represents the capacity limit of the batch processing machine, ensuring that the total number of jobs in the batch processed by the machine does not exceed its maximum capacity, V s,k represents the processing capacity of machine k in stage s; Formula (13-14) represents the order priority constraint, which ensures that different batches processed by the same batch processing machine follow the order, S b,k represents the start time of processing of batch b on machine k, C b,k represents the end processing time of batch b on machine k, γ b,b' Indicates whether batch b is processed before batch b'. If so, γ b,b' =1; otherwise γ b,b' =0; C. A hierarchical dynamic scheduling method based on strategy-optimized DDQN is used to solve the dynamic scheduling problem of a reentrant hybrid flow shop; Firstly, a hierarchical structure based on the strategy optimization DDQN algorithm is proposed, and a batching agent BA and a scheduling agent SA are constructed to solve the batching and scheduling sub-problems of the reentrant hybrid flow shop dynamic scheduling problem with batch processing machines. The batching agent BA decides whether to execute the job batch, while the scheduling agent SA determines the priority order of the job or batch and the machine selection. Secondly, based on the above two agents, a Markov decision process model suitable for the dynamic scheduling problem of a reentrant hybrid flow shop with batch processing machines is designed, which combines multi-stage batch processing and reentrancy characteristics, defines states, rewards and actions, and combines the heuristic variable threshold batch policy VTBP with dynamic workpiece batching; Finally, in order to further enhance the effectiveness and generalization ability of the hierarchical dynamic scheduling method, two improvement strategies are proposed: a) The mask-based action selection strategy is combined with the ε-greedy strategy to limit the action space of each agent at different decision points and improve the efficiency of collaborative decision-making; b) A soft-start network update strategy is used for offline training to accelerate learning and convergence.

2. According to claim 1, a hierarchical dynamic scheduling method based on a policy-optimized DDQN algorithm is characterized in that: The Markov decision process model is specifically: (1) Status: Considering the stage processing characteristics of the hybrid flow shop, the characteristics of each agent are designed according to the processing stage at the decision time point, and 6 state features are extracted as the input of BA, which are: 12 state features are extracted as the input of SA, specifically: (2) Action: The heuristic variable threshold batch policy VTBP is used to construct the action space of the batching agent. The VTBP strategy is as follows: first, determine the process type of the job, then determine the capacity threshold of the batch, then sort the jobs according to certain rules, and finally form a batch by selecting the threshold number of jobs; when scheduling, all jobs in the processing stage buffer are sorted in ascending order according to the time they arrived at the stage and classified by job type; then, calculate the average waiting time of each type of job; select the type q with the longest average waiting time for batch processing; next, it is necessary to determine the threshold V of the current batch b ; Let O s,q,t represents the set of job operations of type q that need to be batch processed at stage s but have not yet arrived at the current time t, BN s,q Indicates the number of type q jobs in the current stage s; if O s,q,t ≠0, indicating that there are still unfinished batch processing operations in the current stage; in this case, if BN s,q >|O s,q,t |, then V b =min(BN s,q ,V s,k ); if BN s,q ≤|O s,q,t |, then V b =0; Otherwise, if O s,q,t =0, then V b =min(BN s,q ,V s,k ); The batch schemes built based on different batch rules directly affect the optimization results. When batch processing jobs in the buffer, the consistency of processing time is given priority, that is, jobs with similar processing time are grouped into the same batch, thereby reducing the waste of processing time. In addition, in order to minimize the total delay, completion time and reduce the waiting time of jobs in the buffer, three batch rules are proposed to balance the consistency of processing time with other time-based performance indicators. A weighted batch quality index is introduced. This index takes into account the consistency of delivery time and processing time of jobs in the buffer, as shown in formula (15): In the formula, buffer s represents the set of jobs in the buffer of stage s during the batch scheduling decision process, q j =q j' It means that only jobs belonging to the same job family can be grouped for batch processing, θ is the weight coefficient, P j,o,s,k represents the processing time required for the current batch operation of job j on machine k, D j represents the delivery time of job j, c j represents the remaining minimum processing time of job j, A s,j represents the arrival time of job j at stage s; batch quality index The smaller the size, the higher the priority of the job in batch processing; jobs are processed in order of Sort in ascending order to determine the batch order; The action space of BA is as follows: When V b =0, BA selects the "None" action and the job will continue to wait; The scheduling point is set as the machine idle time; the action space of SA includes job selection and machine selection; in the batch processing stage, SA schedules the batches formed by BA; in the individual processing stage, the job selection rules and machine selection rules are as follows: (3) Rewards: The total weighted delay of the workpiece is subdivided into each scheduling decision of the decision model BA. After the decision model BA executes each step, an immediate reward will be generated to reflect the impact of the action. The immediate reward of BA at time t is shown in formula (16). In order to make the two goals of minimizing the maximum completion time and maximizing the machine utilization rate be reflected in the actions of the same decision model, the difference between the average machine utilization rate and the completion time difference between each step of scheduling is taken as the immediate reward of the decision model SA. The immediate reward of SA at time t is shown in formula (17). In order to ensure that the average machine utilization rate and the completion time difference between each step of scheduling have the same magnitude in guiding the behavior of the decision model, α1 and α2 are introduced. r SA =α1[U ave (t')-U ave (t)]-α2[C step (t')-C step (t)] (17) in, represents the number of jobs in the job pool of stage p at time t, represents the number of workpieces in batch b, t′ represents the next decision time. If the expiration date of the workpiece is less than the next decision time, Z is 1, otherwise Z is 0. If the workpiece completes the processing of this process at the next decision time, t m is equal to the completion time of the process of the workpiece, otherwise t m Equal to the next decision time t′.

Citation Information

Patent Citations

  • Scheduling method and system for reentrant hybrid flow shop with batch processor

    CN115129002A

  • Flow shop new order insertion optimization method based on deep reinforcement learning

    CN115793583A

  • Two-stage assembly flow shop dynamic scheduling method based on deep reinforcement learning

    CN117369393A

  • Reentrant hybrid flow shop scheduling method based on IMOEA / D algorithm

    CN118642444A

  • Deep reinforcement learning based assistive adjustment system for dry weight of hemodialysis patient

    WO2023202500A1

Cited By

  • Knowledge and data hybrid-driven intelligent scheduling method under random arrival of batch

    CN120218566A

  • Computing-storage stream joint scheduling optimization method and system based on DDQN and heuristic strategy

    CN120371483A

  • Multi-agent batch processing machine and variable sub-batch hybrid flow shop scheduling method

    CN121258124A

  • Method for minimizing maximum completion time of curing and post-processing operation in aviation composite material manufacturing

    CN122175294A