A job shop batching scheduling method based on D3QN and genetic algorithm

By combining the hierarchical iterative optimization strategy of D3QN and genetic algorithm in batch scheduling in the work workshop, the problems of batch division and process sorting are solved, and efficient scheduling strategies and production efficiency are improved.

CN115826530BActive Publication Date: 2025-06-27HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211609644.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-14
Publication Date
2025-06-27
Estimated Expiration
2042-12-14

AI Technical Summary

Technical Problem

The prior art is difficult to effectively solve the problems of batch division and process sorting in batch scheduling of the work workshop, resulting in low scheduling efficiency and high cost.

Method used

A hierarchical iterative optimization strategy based on D3QN and genetic algorithm is adopted to determine the batch scheme through the genetic algorithm, and a trained D3QN model is used to provide an adaptive scheduling strategy for sub-batch process sorting to minimize the maximum completion time.

Benefits of technology

A high-quality scheduling strategy is obtained in a short period of time, which significantly improves production efficiency, reduces costs, and improves the solution speed and generalization ability of batch scheduling problems in the operation workshop.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115826530B_ABST
    Figure CN115826530B_ABST
Patent Text Reader

Abstract

The present invention discloses a job shop batch scheduling method based on D3QN and genetic algorithm. Steps of the present invention: 1. Construct a mathematical model for the job shop batch scheduling problem, and the scheduling objective is to minimize the makespan; 2. The batch division scheme of the job shop batch scheduling problem is determined by the genetic algorithm; 3. Perform crossover and mutation operations on the chromosomes in the population; 4. Decode the chromosomes to obtain the batch division scheme of the workpieces; represent the sub-batch operation sequencing problem with a disjunctive graph model; 5. Use a graph neural network to perform representation learning on the feature information of the disjunctive graph nodes and extract the feature states of the operation sequencing problem; 6. Design a D3QN model structure with prioritized experience replay and train the model; 7. Judge whether the termination condition is satisfied. The present invention adopts a hierarchical iterative optimization strategy, and uses the trained D3QN model to provide an adaptive scheduling strategy for the inner-layer sub-batch operation sequencing problem, so as to minimize the makespan and improve production efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent production scheduling, and proposes a job shop batching scheduling method based on D3QN and genetic algorithm. Background Technique

[0002] The manufacturing industry is the pillar industry of the national economy and also reflects the country's comprehensive strength. China's manufacturing industry is gradually transforming towards intelligence and digitization. Shop scheduling is an important link in the production and manufacturing process. How to achieve an efficient and intelligent scheduling system is one of the keys to enhancing the competitiveness of enterprises. The job shop scheduling problem is a key research direction in the current field of shop production scheduling. Most of the research on the job shop scheduling problem focuses on individual workpieces, while in actual production, in order to improve scheduling efficiency and reduce costs, a batching scheduling mode is often adopted. Therefore, the research on the job shop batching scheduling problem has important practical significance for improving enterprise production efficiency, reducing costs, and obtaining more economic benefits.

[0003] The job shop batching scheduling problem needs to solve two sub-problems: batch division and operation sequencing, which greatly increases the complexity of the problem and the search space of the solution. In recent years, domestic and foreign scholars have mostly used genetic algorithms to solve the job shop batching scheduling problem. The existing technologies have the following deficiencies: 1. The batching scheduling problem is complex and the solution space is huge, and it is difficult for the genetic algorithm to find an ideal solution within a limited time; 2. For different batching schemes, the genetic algorithm needs to re-iterate and solve when solving the sub-batch operation sequencing problem, and the generalization is poor.

[0004] In recent years, deep reinforcement learning has achieved good results in combinatorial optimization fields such as the traveling salesman problem and path optimization problems. Therefore, applying deep reinforcement learning to the scheduling problem is a novel research direction. Currently, no research on using deep reinforcement learning to solve the job shop batching scheduling problem has been found. The D3QN (Dueling Double DQN) algorithm is a new type of deep reinforcement learning algorithm, which combines the advantages of the Dueling DQN and Double DQN algorithms and further improves the traditional DQN algorithm. The present invention uses the D3QN algorithm to solve the sub-batch operation sequencing problem to quickly obtain a better adaptive scheduling strategy, thereby optimizing the makespan and greatly improving the solution speed of the job shop batching scheduling problem. Summary of the Invention

[0005] Aiming at the defects in the prior art, the present invention provides a job shop batching scheduling method based on D3QN and genetic algorithm. This method adopts a hierarchical iterative optimization strategy. The batching scheme of the job shop batching scheduling task is determined by the outer genetic algorithm. Based on the results of the outer batch division, the trained D3QN model is used to provide an adaptive scheduling strategy for the inner sub-batch process sequencing problem, so as to minimize the makespan and improve production efficiency.

[0006] To achieve the above object, the technical solution adopted by the present invention includes the following steps:

[0007] S1: Construct a mathematical model of the job shop batching scheduling problem according to the number of machines in the production workshop, the types of workpieces to be processed, and the number of each type of workpiece to be processed. The goal of scheduling is to minimize the makespan.

[0008] S2: The batch division scheme of the job shop batching scheduling problem is determined by the genetic algorithm: Each batch of workpieces is randomly split into several sub-batches of different sizes for combination based on the coding form of the real number sequence, and the initial population is generated accordingly.

[0009] S3: Perform crossover and mutation operations on the chromosomes in the population to increase the diversity of the population.

[0010] S4: Decode the chromosomes to obtain the batch division scheme of the workpieces; represent the process sequencing problem of the sub-batches after the workpiece batch division with a disjunctive graph model. Based on this disjunctive graph model, establish a Markov decision process, and design the states, actions, and rewards of the process.

[0011] S5: Use a graph neural network to perform representation learning on the obtained disjunctive graph node feature information, capture the implicit relationship between processes, and effectively extract the feature states of the process sequencing problem.

[0012] S6: Design a D3QN model with prioritized experience replay and train the model to provide an adaptive scheduling strategy and makespan for the process sequencing problem of the sub-batches after the workpiece batch division, and use the reciprocal of the makespan as the fitness function value of the genetic algorithm.

[0013] S7: Judge whether the iteration of the genetic algorithm meets the termination condition. If it meets, output the optimal batching scheme and scheduling strategy of the job shop batching scheduling problem. Otherwise, use the roulette wheel method to select the optimal individual in the population to enter the next generation, and execute step S3.

[0014] The establishment process of the mathematical model of the job shop batching scheduling in step S1 is as follows:

[0015] 1-1. The job shop batch scheduling problem is described as follows: Assume that there are n types of workpieces in the workshop waiting to be processed on m machines. Each type of workpiece can be batch-divided into several sub-batches, and the number of workpieces in each sub-batch is randomly allocated. Each sub-batch of each type of workpiece contains k processes. Given the processing machine and processing time of each process, the goal of scheduling is to reasonably divide the workpieces into batches and arrange the processing order of each sub-batch process so as to minimize the makespan.

[0016] 1-2. Symbol definitions for establishing the mathematical model of the job shop batch scheduling problem based on the above problem description: n: The types of workpieces to be processed;

[0017] m: The serial number of the machine, and the machine set M = {M1, M2,..., M m};

[0018] j: The process number of workpiece i;

[0019] k: The sub-batch number of workpiece i;

[0020] B i : The total number of workpiece i;

[0021] A max : The maximum batch quantity of the workpiece;

[0022] P i : The number of sub-batches of workpiece i;

[0023] S ik : The quantity of the k-th batch of workpiece i;

[0024] O ij : The j-th process of workpiece i, j = 1, 2,..., m;

[0025] PT ijm : The processing time of the j-th process of workpiece i on machine m;

[0026] ST ikjm : The start time of the j-th process of the k-th batch of workpiece i on machine m;

[0027] ET ikjm : The end time of the j-th process of the k-th batch of workpiece i on machine m;

[0028] E ikm : The end time of the last process of the k-th batch of workpiece i on machine m;

[0029] X ijm : Decision variable, which is 1 if the j-th process of workpiece i is processed on machine m, otherwise 0; C i : The completion time of workpiece i;

[0030] C max : The maximum completion time when all workpieces are completed;

[0031] 1 - 3. According to the above definition, the following mathematical model can be established for the job - shop batch scheduling problem:

[0032] Objective function:

[0033] min C max = minmax{C i | i = 1, 2,..., n} (1)

[0034] Constraints:

[0035]

[0036] 1 ≤ P i ≤ A max (3)

[0037] ET ikjm = ST ikjm + PT ijm × S ik (4)

[0038] ST ik(j+1)m >= ET ikjm (5)

[0039]

[0040] C max = max{C i | i = 1, 2,..., n} (7)

[0041] Equation (1) indicates that the objective of model optimization is to minimize the completion time; Equation (2) indicates that the sum of the quantities of all sub - batches of workpiece i must be equal to the total quantity of workpiece i; Equation (3) indicates that the number of sub - batch partitions cannot exceed the maximum number of predefined partitions; Equation (4) indicates that the completion time of the j - th operation of the k - th batch of workpiece i on machine m is equal to the sum of the processing start time and the processing time of this batch of operations, where the processing time of this batch of operations is the product of the operation processing time and the batch quantity; Equation (5) indicates that the same workpiece can only be processed in the next operation after the previous operation is completed; Equation (6) indicates that the j - th operation of workpiece i can only be processed on one machine; Equation (7) indicates that the maximum completion time is equal to the maximum value of the completion times of all workpieces.

[0042] The crossover in step S3 refers to randomly generating an integer r (1 ≤ r ≤ n), and then exchanging the chromosome genes corresponding to the workpiece numbered r in two parent chromosomes to obtain two new chromosomes.

[0043] The mutation in step S3 means randomly generating an integer r (1 ≤ r ≤ n), and then randomly selecting two positions of the chromosome gene corresponding to the workpiece numbered r to perform the operations of adding 1 and subtracting 1 to obtain a new chromosome.

[0044] In step S4, to construct a disjunctive graph model for the operation sequencing problem, the specific process is as follows:

[0045] 4-1. The disjunctive graph G = (V, C ∪ D) is a mixed graph, where V represents the set of all processing operation nodes; C represents the set of connection arcs, that is, the sequential constraint relationship between different operations of the same workpiece; D represents the set of disjunctive arcs, and the two operation nodes connected by the disjunctive arc can be processed on the same machine; the scheduling can be regarded as determining the directions of all disjunctive arcs in the graph while minimizing the makespan;

[0046] 4-2. According to the real-time state of the scheduling process, add the following characteristic information to each operation node in the disjunctive graph:

[0047] (1) Operation status: represented by a one-hot vector, where [1, 0, 0] means not completed yet, [0, 1, 0] means being processed, and [0, 0, 1] means completed;

[0048] (2) Operation processing time;

[0049] (3) Operation expected completion time;

[0050] (4) Operation waiting time;

[0051] (5) Operation remaining processing time;

[0052] (6) Workpiece operation completion rate;

[0053] Normalize the time-related data among them, and map their values to the range of [0, 1].

[0054] The design of the actions in step S4 means using 8 heuristic rules (FIFO, LIFO, MOR, LOR, LPT, SPT, LTPT, STPT) as the action space.

[0055] The design of the rewards in step S4, its specific calculation process is as follows:

[0056]

[0057] In the formula, U tRepresents the utilization rate of the machine at time t, which is calculated as: Machine utilization rate = Total working time of the machine / (Current time - Start processing time of the first process); C is a constant related to the scheduling scale, makespan is the actual completion time, and T ini Is the initial estimated completion time, and its calculation method is: L is the time step when the scheduling is completed.

[0058] In step S5, the graph neural network is used to perform representation learning on the disjunctive graph node feature information, and the calculation method of its node representation is as follows:

[0059]

[0060] In the formula Represents the k-th generation node feature of the target node v, h o Represents the connecting arc node feature, h d Represents the disjunctive arc node feature; f θ Represents the update function of the target node v, f o Represents the connecting arc node update function, f d Represents the disjunctive arc node update function; || represents the vector concatenation operator; after K iterations, each node in the disjunctive graph contains the feature states of K-hop neighbor nodes. The sum of the feature states of each node in the graph is taken and then averaged to obtain the feature state of the disjunctive graph, that is: h G =Σ u∈V h v K / |V|.

[0061] In step S6, the D3QN model is trained to provide an adaptive scheduling strategy and completion time for the sub-batch process sequencing problem after workpiece batch division, and the specific process is as follows:

[0062] 6-1. Initialize parameters: Initialize the current Q-network parameter θ, initialize the target Network parameter θ - , and assign the Q-network parameter to the target Network, θ→θ - , the total number of iteration rounds T, the discount factor γ, the exploration rate ∈, the target Network parameter update frequency P, experience replay capacity N, priority experience replay parameters α and β;

[0063] 6-2. Input the current state s into the D3QN network, calculate the Q values corresponding to each action, select the action a using the ∈-greedy algorithm, and the system gives the reward r;

[0064] 6-3. Store the current state s, action a, reward r, and next state s' obtained during the training process in the form of a quadruple (s, a, r, s') in the prioritized experience replay pool to provide experience data for subsequent scheduling decisions;

[0065] 6-4. Sample min-batch samples from the experience pool M using the prioritized experience replay method for training. The probability of sample j being sampled is calculated as follows:

[0066]

[0067] where α represents the priority weight. When α = 0, it represents uniform sampling, and p j represents the priority metric, whose value is related to |TD-error|, and p j The specific calculation formula is as follows:

[0068]

[0069] p j = |δ j | + ∈(12)

[0070] where δ j is the TD-error, and ∈ is a very small positive number to prevent the sampling probability from approaching 0, which is used to ensure that all samples have a probability of being sampled;

[0071] When performing priority sampling, different samples are assigned different probabilities, and the original expected distribution is thus changed. Importance sampling is used to eliminate this bias. The importance sampling weight is calculated as follows:

[0072] ω j = (N·P j ) -β / max i ω i (13)

[0073] where β is a hyperparameter used to adjust the degree of bias;

[0074] Input the sampled data into the D3QN model, calculate the time difference error TD-error, and then update the priority in the prioritized replay experience mechanism;

[0075] 6-5. Calculate the loss function and continuously update the weight parameters of the D3QN network through the stochastic gradient descent method. The calculation method of the loss function is as follows:

[0076]

[0077] where γ represents the discount factor.

[0078] 6-6. Cycle training until the cumulative reward converges to obtain a scheduling model for the final operation sequencing problem. The trained model supports the rapid solution of the operation sequencing problem and can adapt to operation sequencing problems of different scales, with good generalization performance.

[0079] Compared with the prior art, the present invention has the following advantages and effects:

[0080] 1. The present invention decomposes the job shop batch scheduling problem into two sub-problems: batch division and operation sequencing of sub-batches, and adopts a hierarchical iterative optimization strategy to solve the two sub-problems, reducing the complexity of the problem. The genetic algorithm is used to determine the batch division scheme of workpieces, and the trained D3QN model is used to solve the operation sequencing problem of sub-batches, and a high-quality scheduling strategy can be obtained in a short time.

[0081] 2. The present invention performs representation learning on the disjunctive graph through a graph neural network, and designs an effective graph node representation calculation method to adapt to different scheduling environments. Therefore, the trained D3QN model has good generalization ability, that is, the model trained under small-scale problems can be directly used for large-scale problems without repeated training, and has good generalization performance.

[0082] 3. Compared with the traditional genetic algorithm, the method designed by the present invention has better solution effect under the same number of iterations; in terms of solution speed, this method is one order of magnitude faster than the traditional genetic algorithm, because when solving the operation sequencing problem of sub-batches, the genetic algorithm needs to re-iterate and solve each time, which is very time-consuming, while the D3QN model trained by the present invention can provide a better scheduling strategy within milliseconds. Description of the Drawings

[0083] Figure 1 is a framework diagram of a job shop batch scheduling method based on D3QN and genetic algorithm.

[0084] Figure 2 is a schematic diagram of genetic algorithm chromosome encoding, crossover, and mutation.

[0085] Figure 3 is a disjunctive graph model of a 3×3 operation sequencing problem.

[0086] Figure 4 is a framework diagram of a graph neural network combined with a D3QN model. Detailed Embodiments

[0087] In order to make the purpose, technical solution and advantages of the present invention clearer, the following further explains the present invention with reference to the drawings and embodiments.

[0088] Figure 1It is a framework diagram for solving the job shop batching scheduling problem using D3QN and genetic algorithms. The implementation of the present invention provides a job shop batching scheduling method based on D3QN and genetic algorithms, and the specific steps are as follows:

[0089] S1. Construct a job shop batching scheduling problem model based on the number of machines in the production workshop, the types of workpieces to be processed, and the processing quantity of each type of workpiece. The goal of scheduling is to minimize the makespan, which specifically includes the following steps:

[0090] 1) Make the following description for the job shop batching scheduling problem: Assume that there are n types of workpieces in the workshop waiting to be processed on m machines. Each type of workpiece can be batch-divided into several sub-batches, and the number of workpieces in each sub-batch is randomly allocated. Each sub-batch of each type of workpiece contains k processes. Given the processing machine and processing time of each process, the goal of scheduling is to reasonably divide the workpieces into sub-batches and arrange the processing order of each sub-batch process to minimize the makespan;

[0091] 2) Give the symbolic definition of the model according to the above problem description:

[0092] n: The types of workpieces to be processed;

[0093] m: The serial number of the machine, and the machine set M = {M1, M2,..., M m};

[0094] j: The process number of workpiece i;

[0095] k: The sub-batch batch number of workpiece i;

[0096] B i : The total quantity of workpiece i;

[0097] A max : The maximum batch quantity of the workpiece;

[0098] P i : The number of sub-batches of workpiece i;

[0099] S ik : The quantity of the k-th batch of workpiece i;

[0100] O ij : The j-th process of workpiece i, j = 1, 2,..., m;

[0101] PT ijm : The processing time of the j-th process of workpiece i on machine m;

[0102] ST ikjm : The start time of the j-th process of the k-th batch of workpiece i on machine m;

[0103] ET ikjm: The completion time of the j-th operation of the k-th batch of workpiece i on machine m;

[0104] E ikm : The completion time of the last operation of the k-th batch of workpiece i on machine m;

[0105] X ijm : Decision variable, which is 1 if the j-th operation of workpiece i is processed on machine m, otherwise 0;

[0106] C i : The completion time of workpiece i;

[0107] C max : The makespan when all workpieces are completed;

[0108] (3) According to the above symbol definitions, the following model can be established for the job shop batching scheduling problem:

[0109] Objective function:

[0110] minC max = min max{C i | i = 1, 2,..., n} (1)

[0111] Constraints:

[0112]

[0113] 1 ≤ P i ≤ A max (3)

[0114] ET ikjm = ST ikjm + PT ijm × S ik (4)

[0115] ST ik(j+1)m >= ET ikjm (5)

[0116]

[0117] C max = max{C i | i = 1, 2,..., n} (7)

[0118] Equation (1) indicates that the goal of model optimization is to minimize the makespan; Equation (2) indicates that the sum of the quantities of all sub - batches of workpiece i must be equal to the total quantity of workpiece i; Equation (3) indicates that the number of sub - batch divisions cannot exceed the maximum predefined number of division batches; Equation (4) indicates that the completion time of the j - th operation of the k - th batch of workpiece i processed on machine m is equal to the sum of the processing start time and the operation processing time of this batch, where the operation processing time of this batch is the product of the operation processing time and the batch quantity; Equation (5) indicates that the same workpiece must be processed in the next operation only after the previous operation is completed; Equation (6) indicates that the j - th operation of workpiece i can be processed on only one machine; Equation (7) indicates that the makespan is equal to the maximum value of the completion times of all workpieces.

[0119] S2. The batch division scheme of the job - shop batching scheduling problem is determined by a genetic algorithm. Based on the encoding form of real - number sequences, each batch of workpieces is randomly split into several sub - batches of different sizes for combination, and an initial population is generated accordingly; for example, Figure 2 the encoding (3, 4, 3, 5, 5, 0, 2, 4, 4) in [reference] means that workpiece A, B, and C are at most divided into 3 sub - batches. The specific scheme is: A is divided into 3 batches, and the number of workpiece A in each sub - batch is 3, 4, 3; B is divided into 2 batches, and the number of workpiece B in each sub - batch is 2, 2; C is divided into 3 batches, and the number of workpiece C in each sub - batch is 2, 4, 4.

[0120] S3. Perform crossover and mutation operations on the chromosomes in the population to increase the diversity of the population; decode the chromosomes to obtain the batch division scheme of the workpieces;

[0121] The so - called crossover means randomly generating an integer r (1 ≤ r ≤ n), and then exchanging the chromosome genes corresponding to the workpiece numbered r in two parent chromosomes to obtain two new chromosomes; for example, Figure 2 the crossover in [reference] means that the randomly generated r = 2, and then the chromosome genes corresponding to the workpiece numbered 2 in two parent chromosomes q1 and q2 are exchanged to obtain two chromosomes q1′ and q2′.

[0122] The so - called mutation means randomly generating an integer r (1 ≤ r ≤ n), and then randomly selecting two positions in the chromosome genes corresponding to the workpiece numbered r to perform the operations of adding 1 and subtracting 1 to obtain a new chromosome; for example, Figure 2 the mutation in [reference] means that the randomly generated r = 3, and then the 1st position and the 2nd position in the chromosome genes corresponding to the workpiece numbered 3 in chromosome q are selected to perform the operations of adding 1 and subtracting 1 to obtain a new chromosome q′.

[0123] S4. Represent the problem of sub-batch process sequencing after batch partitioning of workpieces using a disjunctive graph model. Based on the constructed disjunctive graph model of the process sequencing problem, establish a Markov decision process, and design the states, actions, and rewards of the process. The specific process is as follows:

[0124] 1) The disjunctive graph G = (V, C ∪ D) is a mixed graph, where V represents the set of all processing operation nodes; C represents the set of connecting arcs, that is, the sequential constraint relationship between different operations of the same workpiece; D represents the set of disjunctive arcs, and the two operation nodes connected by the disjunctive arc can be processed on the same machine. The scheduling can be regarded as determining the directions of all disjunctive arcs in the graph while minimizing the makespan; Figure 3 is a disjunctive graph model for a process sequencing problem;

[0125] 2) According to the real-time state of the scheduling process, add the following characteristic information to each operation node in the disjunctive graph: (1) Operation state: represented by a one-hot vector, such as [1, 0, 0] indicating not completed, [0, 1, 0] indicating being processed, [0, 0, 1] indicating completed; (2) Operation processing time; (3) Expected completion time of the operation; (4) Waiting time of the operation; (5) Remaining processing time of the operation; (6) Completion rate of the workpiece operation; Normalize the time-related data and map its value to the range [0, 1] to reduce the variance and improve the robustness of the model;

[0126] 3) The action in the above step refers to using 8 heuristic rules (FIFO, LIFO, MOR, LOR, LPT, SPT, LTPT, STPT) as the action space.

[0127] 4) The reward in the above step, its specific calculation process is as follows:

[0128]

[0129] In the formula, U t represents the utilization rate of the machine at time t, and its calculation method is: Machine utilization rate = Total working time of the machine / (Current time - Start processing time of the first operation); C is a constant related to the scheduling scale, makespan is the actual completion time, T ini is the initial expected completion time, and its calculation method is: L is the time step when the scheduling is completed.

[0130] S5. Use a graph neural network to perform representation learning on the characteristic information of the disjunctive graph nodes, capture the implicit relationship between operations, and effectively extract the characteristic states of the process sequencing problem;

[0131] In the above steps, graph neural network is used to perform representation learning on the disjunctive graph node feature information. The specific calculation method of node representation is as follows:

[0132]

[0133] In the formula represents the k-th generation node feature of the target node v, h o represents the connection arc node feature, h d represents the disjunctive arc node feature; f θ represents the update function of the target node v, f o represents the connection arc node update function, f d represents the disjunctive arc node update function; || represents the vector concatenation operator; after K iterations, each node in the disjunctive graph contains the feature states of K-hop neighbor nodes. The feature states of each node in the graph are summed and then averaged to obtain the feature state of the disjunctive graph, that is: h G = Σ u∈V h v K / |V|.

[0134] S6. Design a D3QN model structure with prioritized experience replay and train the model to provide an adaptive scheduling strategy and makespan for the sub-batch operation sequencing problem after workpiece batch division, and use the reciprocal of the makespan as the fitness function value of the genetic algorithm Figure 4 is the framework diagram of the D3QN model; the specific process of D3QN model training is as follows:

[0135] 1) Initialize parameters: Initialize the current Q-network parameters θ, initialize the target network parameters θ - , and assign the Q-network parameters to the target network, θ → θ - , the total number of iteration rounds T, the discount factor γ, the exploration rate ∈, the target network parameter update frequency P, the experience replay capacity N, the prioritized experience replay parameters α and β;

[0136] 2): Input the current state s into the D3QN network, calculate the Q values corresponding to each action, select the action a using the ∈-greedy algorithm, and the system gives the reward r;

[0137] 3) Store the current state s, action a, reward r, and the next state s′ obtained during the training process in the form of a quadruple (s, a, r, s′) in the prioritized experience replay pool to provide experience data for subsequent scheduling decisions;

[0138] 4) Samples of min - batch are sampled from the experience pool M using the prioritized experience replay method for training. The probability that sample j is sampled is calculated as follows:

[0139]

[0140] where α represents the priority weight. When α = 0, it represents uniform sampling, and p j represents the priority metric, whose value is related to |TD - error|. The specific calculation formula of p j is as follows:

[0141]

[0142] p j = |δ j | + ∈

[0143] where δ j is the TD - error, and ∈ is a very small positive number to prevent the sampling probability from approaching 0, which is used to ensure that all samples have a probability of being sampled;

[0144] When performing priority sampling, different samples are assigned different probabilities, and the original expected distribution is thus changed. Importance sampling is used to eliminate this bias. The importance sampling weight is calculated as follows:

[0145] ω j = (N·P j ) -β / max i ω i

[0146] where β is a hyperparameter used to adjust the degree of bias;

[0147] The sampled data is input into the D3QN model to calculate the time - difference error TD - error, and then the priorities in the prioritized replay experience mechanism are updated;

[0148] 5) Calculate the loss function, and continuously update the weight parameters of the D3QN network by the stochastic gradient descent method. The calculation method of the loss function is as follows:

[0149]

[0150] where γ represents the discount factor.

[0151] 6) Train in a loop until the cumulative reward converges to obtain the scheduling model for the final process - sequencing problem. The trained model supports the fast solution of the process - sequencing problem and can adapt to process - sequencing problems of different scales, with good generalization;

[0152] S7. Determine whether the iteration count has been reached. If so, output the optimal batch scheduling plan and scheduling strategy for the job shop batch scheduling problem; otherwise, use the roulette wheel method to select the optimal individual in the population to enter the next generation, and execute step S3.

[0153] Use a specific job shop batch scheduling example to prove the effectiveness and practicality of the present invention. Table 1 shows an example of a job shop scheduling problem of size 4×4. This scheduling problem has 4 types of workpieces, which are processed on 4 machines, with 10 of each type of workpiece. Each workpiece contains 4 processes, and the processing machine and processing time corresponding to each process are known.

[0154] Table 1 Example of job shop scheduling problem

[0155]

[0156] The solution steps using this method are as follows:

[0157] 1) Initialize the genetic algorithm parameters: population size N = 30, number of generations T = 50, crossover probability p c = 0.8, mutation probability p m = 0.1, and let the current iteration count t = 1;

[0158] 2) Use real number sequence encoding and generate the initial population;

[0159] 3) Perform crossover and mutation operations on the chromosomes;

[0160] 3) Decode to obtain the batch division plan of the workpieces. Based on the result of the batch division, use the trained D3QN model to provide an adaptive scheduling strategy and the makespan for sequencing the sub-batch processes;

[0161] 4) Take the reciprocal of the makespan as the fitness function value, and use the roulette wheel method to select N optimal individuals to enter the next generation population;

[0162] 5) Judge whether t >= T. If it holds, end the algorithm and output the better batch scheduling plan and scheduling strategy; otherwise, jump to step 3.

[0163] The makespan obtained by solving with this method is 322, and the batch quantity of each type of workpiece is: (1, 2, 3, 4), (1, 1, 1, 7), (1, 2, 3, 4), (2, 2, 3, 3); the makespan obtained by solving with the genetic algorithm is 348; while the makespan obtained by non-batch solving is 360. It can be seen that workpiece batching is beneficial to improving production efficiency. In addition, in terms of the solving time, using the method provided by the present invention takes 936 seconds, while using the genetic algorithm to obtain an ideal solution takes 2 hours and 16 minutes, which is about 8 times the time taken by the method provided by the present invention.

[0164] From the results of the above embodiments, it can be concluded that this method can effectively solve the job-shop batching scheduling problem and has a higher solution accuracy than the genetic algorithm. Additionally, in terms of the solution time, it takes a large amount of time to obtain a relatively good solution using the genetic algorithm, while using this method can obtain a relatively good solution in a shorter time, which is one order of magnitude faster than the traditional genetic algorithm in terms of the solution speed.

Claims

1. A job shop batching scheduling method based on D3QN and genetic algorithm, characterized in that It includes the following steps: S1: Construct a mathematical model of the job shop batch scheduling problem based on the number of machines in the production workshop, the types of workpieces to be processed, and the processing quantity of each type of workpiece. The goal of scheduling is to minimize the makespan; S2: The batch division scheme of the job shop batch scheduling problem is determined by a genetic algorithm: based on the encoding form of a real number sequence, each batch of workpieces is randomly split into several sub-batches of different sizes for combination, and an initial population is generated accordingly; S3: Perform crossover and mutation operations on the chromosomes in the population to increase the diversity of the population; S4: Decode the chromosomes to obtain the batch division scheme of the workpieces; represent the operation sequencing problem of the sub-batches after the workpiece batch division using a disjunctive graph model. Based on this disjunctive graph model, establish a Markov decision process, and design the states, actions, and rewards of the process; S5: Use a graph neural network to perform representation learning on the obtained disjunctive graph node feature information, capture the implicit relationship between operations, and effectively extract the feature states of the operation sequencing problem; S6: Design a D3QN model with prioritized experience replay and train the model to provide an adaptive scheduling strategy and makespan for the operation sequencing problem of the sub-batches after the workpiece batch division, and use the reciprocal of the makespan as the fitness function value of the genetic algorithm; S7: Determine whether the iteration of the genetic algorithm meets the termination condition. If it meets, output the optimal batch division scheme and scheduling strategy of the job shop batch scheduling problem. Otherwise, use the roulette wheel method to select the optimal individual in the population to enter the next generation, and execute step S3; 2. The job shop batching scheduling method based on D3QN and genetic algorithm according to claim 1, characterized in that The process of establishing the mathematical model of job shop batch scheduling in step S1 is as follows: 1-1. Make the following description for the job shop batch scheduling problem: Assume that there are n types of workpieces in the workshop waiting to be processed on m machines. Each type of workpiece can be batch-divided into several sub-batches, and the number of workpieces in each sub-batch is randomly allocated. Each sub-batch of each type of workpiece contains k operations. Given the processing machine and processing time of each operation, the goal of scheduling is to reasonably divide the workpieces into batches and arrange the processing order of each sub-batch operation to minimize the makespan; 1-2. Establish the symbolic definition of the mathematical model of the job shop batch scheduling problem according to the above problem description: n: The types of workpieces to be processed; m: The serial number of the machine, and the machine set M = {M1, M2,..., M m}; j: The operation number of workpiece i; k: The sub-batch number of workpiece i; B i : The total number of workpiece i; A max : The maximum number of batches of workpieces; P i : The sub - lot quantity of workpiece i; S ik : The quantity of the k-th batch of workpiece i; O ij : The j-th process of workpiece i, where j = 1, 2,..., m; PT ijm : The processing time of the j-th operation of workpiece i on machine m; ST ikjm : The start time of the j-th process of the k-th batch of workpiece i processed on machine m; ET ikjm : The completion time of the j-th operation of the k-th batch of workpiece i processed on machine m; E ikm : The completion time of the last operation of the k-th batch of workpiece i processed on machine m; X ijm : Decision variable, which is 1 if the j-th operation of workpiece i is processed on machine m, and 0 otherwise; C i : Completion time of workpiece i; C max : The makespan when all jobs are completed; 1-3. According to the above definitions, the following mathematical model can be established for the job shop batch scheduling problem: Objective function: minC max = minmax{C i | i = 1, 2, ..., n}(1) Constraint: 1 ≤ P i ≤ A max (3) ET ikjm = ST ikjm + PT ijm × S ik (4) ST ik(j+1)m ≥ ET ikjm (5) C max = max{C i | i = 1, 2, ..., n}(7) Equation (1) indicates that the goal of model optimization is to minimize the completion time; Equation (2) indicates that the sum of the quantities of all sub-batches of workpiece i must be equal to the total quantity of workpiece i; Equation (3) indicates that the number of sub-batch partitions cannot exceed the maximum number of pre-specified partitions; Equation (4) indicates that the completion time of the j-th process of the k-th batch of workpiece i processed on machine m is equal to the sum of the processing start time and the processing time of the processes in this batch, where the processing time of the processes in this batch is the product of the process processing time and the batch quantity; Equation (5) indicates that the same workpiece must be processed in the next process after the previous process is completed; Equation (6) indicates that the j-th process of workpiece i can only be processed on one machine; Equation (7) indicates that the makespan is equal to the maximum value of the completion times of all workpieces.

3. A job shop batching scheduling method based on D3QN and genetic algorithm according to claim 2, characterized in that The crossover in step S3 refers to randomly generating a crossover position r, 1 ≤ r ≤ n, and then exchanging the chromosome genes corresponding to the workpiece numbered r in two parent chromosomes to obtain two new chromosomes.

4. A job shop batching scheduling method based on D3QN and genetic algorithm according to claim 3, characterized in that The mutation in step S3 refers to randomly generating a mutation position r, 1 ≤ r ≤ n, and then randomly selecting two positions in the chromosome genes corresponding to the workpiece numbered r to perform the operations of adding 1 and subtracting 1 to obtain a new chromosome.

5. A job shop batching scheduling method based on D3QN and genetic algorithm according to claim 3 or 4, characterized in that In step S4, a disjunctive graph model for constructing the operation sequencing problem is as follows: 4-1. The disjunctive graph G = (V, C ∪ D) is a mixed graph, where V represents the set of all processing operation nodes; C represents the set of connection arcs, that is, the sequential constraint relationship between different operations of the same workpiece; D represents the set of disjunctive arcs, and the two operation nodes connected by the disjunctive arc are processed on the same machine; the scheduling can be regarded as determining the directions of all disjunctive arcs in the graph while minimizing the makespan; 4-2. According to the real-time state of the scheduling process, the following characteristic information is added to each operation node in the disjunctive graph: (1) Operation status: represented by a one-hot vector, where [1, 0, 0] represents not completed, [0, 1, 0] represents being processed, and [0, 0, 1] represents completed; (2) Operation processing time; (3) Expected completion time of the operation; (4) Operation waiting time; (5) Remaining processing time of the operation; (6) Completion rate of the workpiece operation; Normalize the time-related data among them and map its value to the range of [0, 1].

6. The job shop batching scheduling method based on D3QN and genetic algorithm according to claim 5, characterized in that The design of the actions in step S4 refers to using 8 heuristic rules as the action space.

7. A job shop batching scheduling method based on D3QN and genetic algorithm according to claim 5, characterized in that The design of the rewards in step S4 has the following specific calculation process: In the formula, U t represents the utilization rate of the machine at time t, and its calculation method is: machine utilization rate = total working time of the machine / (current time - start processing time of the first process); C is a constant related to the scheduling scale, makespan is the actual completion time, and T ini is the initial estimated completion time, and its calculation method is: L is the time step when the scheduling is completed.

8. A job shop batching scheduling method based on D3QN and genetic algorithm according to claim 7, characterized in that In step S5, a graph neural network is used to perform representation learning on the disjunctive graph node characteristic information, and its node representation calculation method is as follows: where represents the k-th generation node feature of the target node v, h o represents the connection arc node feature, h d represents the disjunctive arc node feature; f θ represents the update function of the target node v, f o represents the update function of the connection arc node, f d represents the disjunctive arc node update function; || represents the vector concatenation operator; after K iterations, each node in the disjunctive graph contains the feature states of its K-hop neighbor nodes. Summing up the feature states of all nodes in the graph and then taking the average gives the feature state of the disjunctive graph, that is: h G = ∑ u∈V h v K / |V|.

9. A job shop batching scheduling method based on D3QN and genetic algorithm according to claim 7 or 8, characterized in that In step S6, the D3QN model is trained to provide an adaptive scheduling strategy and completion time for the sub-batch operation sequencing problem after workpiece batch partitioning, which specifically includes the following process: 6-1. Initialize parameters: Initialize the current Q-network parameters θ, initialize the target network parameters θ - , and assign the Q-network parameters to the target network, θ → θ - , the total number of iterations T, the discount factor γ, the exploration rate ∈, the target network parameter update frequency P, the experience replay capacity N, the priority experience replay parameters α and β; 6-2. Input the current state s into the D3QN network, calculate the Q values corresponding to each action, select the action a using the ∈-greedy algorithm, and the system gives the reward r; 6-3. Store the current state s, action a, reward r, and next state s′ obtained during the training process in the form of a quadruple (s, a, r, s′) in the prioritized experience replay pool to provide experience data for subsequent scheduling decisions; 6-4. Sample min-batch samples from the experience pool M using the prioritized experience replay method for training. The probability of sample j being sampled is calculated as follows: where α represents the priority weight, and when α = 0, it represents uniform sampling, p j represents the priority index, and its value is related to |TD-error|, p j The specific calculation formula is as follows: p j = |δ j | + ∈ (12) where δ j i.e., the TD-error, and ∈ is a very small positive number to prevent the sampling probability from approaching 0 and to ensure that all samples have a probability of being sampled; When performing prioritized sampling, different samples are assigned different probabilities, and the original expected distribution is thus changed. Importance sampling is used to eliminate this bias. The importance sampling weight is calculated as follows: ω j =(N·P j ) -β / max i ω i (13) In the formula, β is a hyperparameter used to adjust the degree of bias; Input the sampled data into the D3QN model, calculate the temporal difference error TD-error, and then update the priorities in the prioritized replay experience mechanism; 6-5. Calculate the loss function and continuously update the weight parameters of the D3QN network by the stochastic gradient descent method. The calculation method of the loss function is as follows: In the formula, γ represents the discount factor; 6-6. Repeat the training until the cumulative reward converges to obtain the scheduling model for the final process sequencing problem.

Citation Information

Patent Citations

  • DDQN-based intelligent workshop dynamic adaptive scheduling method and system

    CN114037341A

  • Job-shop adaptive scheduling method based on deep reinforcement learning

    CN114707881A