Novel flexible job shop dynamic batch scheduling method
Through the hierarchical iteration strategy, combined with genetic algorithms and A2C models, the problem of batch scheduling in the flexible work workshop is solved, efficient workpiece group batch scheme and process sorting is achieved, the maximum completion time is reduced, and good generalization ability is provided.
Patent Information
- Application Number
- CN202510039388.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-09
AI Technical Summary
The existing workpiece group workshop scheduling methods are difficult to effectively solve the batch scheduling problem, especially when facing complex workpiece groups and multiple process routes, it is difficult to find an ideal solution and lack generalization.
Using a hierarchical iteration strategy, the outer layer uses a genetic algorithm to determine the batch scheme of the flexible work workshop, and the inner layer uses the trained A2C model to provide a scheduling scheme for the process sorting of sub-batches, combining the dual attention network extraction process and machine features.
It realizes efficiently finding the global optimal artifact group batch solution within a limited time, reduces the maximum completion time, and has good generalization capabilities, and can obtain high-quality solutions in examples of different sizes.
Smart Images

Figure CN119962891A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of combinatorial optimization and workshop scheduling and is a flexible job-shop scheduling problem (Flexible Job-shop Scheduling Problem, FJSP), and in particular relates to a new dynamic batch scheduling method for a flexible job-shop. Background Art
[0002] In modern discrete manufacturing systems, the workshop scheduling of all production order workpiece groups in a factory or workshop within a period of time is one of the central links in the management of the manufacturing process. However, most of the existing workpiece group workshop scheduling methods do not sort each workpiece in the workpiece group in batches when solving the workpiece group scheduling plan, that is, each workpiece is implemented as a batch of unified flexible process planning, or only sorts a single workpiece without considering the batch.
[0003] The population optimization algorithm has always been the mainstream algorithm for solving combinatorial optimization problems due to its advantages such as simple implementation and fast convergence speed. However, it has the following shortcomings: 1. When facing batch scheduling problems such as workshop scheduling of workpiece groups, the population optimization algorithm is difficult to find an ideal solution within a limited time due to the complexity of the problem and the huge solution space; 2. For different batching schemes, the population optimization algorithm needs to iterate again when solving the sub-batch process sorting problem, and it does not have generalization.
[0004] In recent years, research on deep reinforcement learning to solve workshop scheduling problems has rapidly emerged. Due to its good generalization and high solution quality, it has been increasingly applied. In the face of complex workshop scheduling problems, deep reinforcement learning is expected to provide an effective solution to the problem. However, no research results have been found so far on the joint solution of genetic algorithms and deep reinforcement learning to solve the problem of scheduling workpieces with flexible process routes and dynamic batching. Based on this background demand, this patent proposes a new method for dynamic batch scheduling of flexible job shops. Summary of the invention
[0005] The present invention proposes a new method for dynamic batch scheduling of flexible job shops, which adopts a hierarchical iterative strategy: the outer layer solves the batch optimization problem of different types of workpieces, and the inner layer solves the process sorting optimization problem of each sub-batch of workpieces after batching. Among them, the batching scheme of the dynamic batch scheduling of the flexible job shop is determined by the genetic algorithm of the outer layer, and the biological genetic mechanism is used instead of the random generation mechanism to optimize the batching scheme solution set of the workpiece group. Compared with the random generation mechanism to optimize the batching scheme of the workpiece group, the use of the biological genetic mechanism can perform a global search in a limited solution space, which helps to efficiently find the global optimal batching scheme of the workpiece group; based on the batching scheme of the outer layer, the A2C model trained in the inner layer is used to provide a scheduling scheme for the process sorting of the sub-batch. Compared with the use of traditional meta-heuristic algorithms to solve the process sorting problem, the present invention proposes an A2C algorithm based on a dual attention network with good generalization ability, which can achieve a model trained on a small-scale instance. When facing the solution of instances of any size, it can obtain high-quality solutions, thereby achieving the goal of minimizing the maximum completion time.
[0006] A new flexible job shop dynamic batch scheduling method mainly includes the following steps:
[0007] S1: Based on a group of workpieces with different types and different batches of each type having flexible process routes, a flexible job shop dynamic batch scheduling model is established with the goal of minimizing the maximum completion time;
[0008] S2: Set the initial parameters of the genetic algorithm, including the population size M and the crossover probability p c And the mutation probability p m , maximum number of iterations iter max ;
[0009] S3: Randomly generate a batching scheme for each workpiece in the workpiece group, and encode the batching scheme corresponding to each workpiece using integer encoding; input the number A of each workpiece j j and the maximum batch quantity B set by the dispatching system max Then randomly determine the batching plan b1, b2, ..., b for each workpiece n ;
[0010] Where n is the number of workpiece types, batch plan N ju represents the number of workpieces in the u-th sub-batch of the j-th workpiece, b j represents the batching plan for the jth workpiece, each b j The length of the vector is B max ;
[0011] S4: Encode the batching schemes of all the workpieces in the workpiece group into a chromosome, and combine the batching schemes of all types of workpieces into a chromosome, which is represented by X i =[b1, b2, …, b n ], X i The length is n×B max ;
[0012] Among them, chromosome X i An initial integration batching scheme representing the artifact population;
[0013] S5: Determine whether the currently generated chromosome data meets the number required for population initialization. If not, return to step S3. If it meets the condition, combine M chromosomes into a population as the initialization population of the genetic algorithm; initialize the batch scheme chromosome population X of all types of workpieces = {X1, X2, ..., X M}, and use it as the initialization input of the batching scheme of the workpiece group, M represents the population size;
[0014] S6: According to the chromosome crossover probability p c , mutation probability p m Perform crossover and mutation operations on chromosomes in the population;
[0015] S7: Decode the chromosome to obtain the batching scheme of the workpiece, convert the sub-batch information after the batch division into a standard data format, and provide scheduling data for the subsequent A2C model;
[0016] S8: The problem of sorting the workpiece processes of all sub-batches obtained after decoding the workpiece group represented by each chromosome is expressed as a Markov decision process;
[0017] S9: Use the dual attention mechanism network to extract process features and machine features to train the A2C model to solve the process sorting problem after batch division of workpieces;
[0018] S10: using the A2C model trained in step S9 to solve the process sequencing problem after the workpiece batch division, obtain the maximum completion time under different batching schemes, and use the reciprocal of the maximum completion time as the fitness function value of the genetic algorithm;
[0019] S11: If the termination condition is not met, the roulette selection method is executed and then the process returns to S6; if the termination condition is met, the current global optimal solution is updated; the termination condition is that in the past few generations, there is no obvious improvement in the individual and the optimal fitness value of the population has almost no fluctuation, that is, the algorithm converges;
[0020] S12: If iter<iter max, then return to S3; otherwise, output the optimal batching plan and process scheduling plan; where iter is the number of current iterations.
[0021] The mathematical model establishment process of the dynamic batch scheduling of the flexible job shop in step S1 is as follows:
[0022] 1-1. The dynamic batch scheduling problem for a flexible job shop of n×m size can be described as follows: n types of workpieces J = {J1, J2, ..., J n}On m machines M={M1,M2,......,M m}, the number of each workpiece is A={A j |j=1,2,......,n}. Each workpiece contains several processes. There are process order constraints between the same workpieces. There are multiple processing machines to choose from for each process. The processing time corresponding to different processing machines is also different. The workpieces are processed in batches, and the number of sub-batches is B={B j |j=1,2,......,n}, and each sub-batch is considered as a whole during production scheduling. The scheduling goal is to minimize the maximum completion time through reasonable workpiece batch division and sub-batch process arrangement.
[0023] 1-2. Based on the above problem description, the symbolic definition of the mathematical model of the flexible job shop batch scheduling problem is established:
[0024] n: number of workpiece types;
[0025] m: total number of machines;
[0026] i: machine number, i=1,2,3,...,m;
[0027] j: workpiece type number, j = 1, 2, 3, ..., n;
[0028] A j : The number of the jth type of workpieces;
[0029] B j : The number of sub-batches of the j-th workpiece;
[0030] B max : The maximum batch quantity set by the scheduling;
[0031] f,l: process number, f=1,2,3,..., l=1,2,3,...;
[0032] u: sub-batch number, u=1,2,3,...,B j ;
[0033] N ju : The batch size of the u-th sub-batch of the j-th workpiece;
[0034] P ijf :Machine M i The processing time of the fth process for processing the jth workpiece;
[0035] S ijuf :Machine M i The start time of the f-th operation for processing the u-th sub-batch of the j-th workpiece;
[0036] E ijuf :Machine M i The end time of the f-th operation of processing the u-th sub-batch of the j-th workpiece;
[0037] X ijuf :The fth process of the jth workpiece and the uth sub-batch on machine M i 1 when processing, otherwise 0;
[0038] Y ijufku'l :The fth process of the jth workpiece and the uth sub-batch on machine M i When the start time of the upper processing is earlier than the lth process of the u'th sub-batch of the kth workpiece, the value is 1, otherwise it is 0;
[0039] L: a sufficiently large positive number;
[0040] C max : Maximum completion time;
[0041] 1-3. Based on the above definition, the following mathematical model can be established for the flexible job shop batch scheduling problem:
[0042] f=min(C max ) (1)
[0043]
[0044] S ijuf ≥E i'ju(f-1) (3)
[0045]
[0046] E ijuf =S ijuf +P ijf ×N ju (5)
[0047] E ijuf ≤S iku'l +L(1-Y ijufku'l ) (6)
[0048] 1≤B j ≤B max (7)
[0049] E ijuf ≥0,S ijuf ≥0,P ijf ≥0 (8)
[0050] Formula (1) indicates that the scheduling objective is to minimize the maximum completion time; Formula (2) indicates that the total number of workpieces remains unchanged after being divided into batches; Formula (3) indicates that the next process can only be started after the current process of each batch is completed; Formula (4) indicates that each process of each part sub-batch can only be processed by one machine; Formula (5) indicates that the fth process of the uth sub-batch of the jth part is processed on machine M. i The end time of the previous processing is equal to the sum of the start time and the processing time; Formula (6) represents the processing time constraint between sub-batches processed by the same machine; Formula (7) represents that the maximum number of batches of workpieces cannot exceed the set maximum number of batches; Formula (8) represents a non-negative constraint.
[0051] As a further improvement, the crossover in step S6 adopts two crossover modes:
[0052] 6-1. Perform a crossover operation on the chromosome and randomly generate a probability p (0≤p≤1);
[0053] 6-1-1. If 0≤p<0.5, generate a random number r (1≤r≤n), perform a pairwise crossover on the batch scheme of the rth workpiece of the selected chromosome, and perform a pairwise crossover on the parent chromosome. Crossover to generate offspring individuals
[0054] 6-1-2. If 0.5<p≤1, two different random numbers r1 and r2 are generated (1≤r1, r2≤n, r1≠r2), and the batching schemes of the r1th and r2th workpieces of the selected chromosomes are crossed in pairs, that is, the difference ΔA between the number of workpieces of the r1th workpiece of the first chromosome and the r2th workpiece of the second chromosome is taken as the partial code value, and the batching schemes of r1 and r2 of the two chromosomes are crossed respectively, and the partial code value ΔA is used to uniformly adjust the batching schemes after crossing, that is, the batch number of workpieces r1 and r2 of the original chromosome remains unchanged after crossing, and each sub-batch is adjusted by ± Operation, that is, after crossing, perform + for each sub-batch of the workpiece type r1 and r2 with more workpieces Operation, make up for the missing number of workpieces, and execute for each sub-batch of the workpiece type r1 and r2 with the smaller number of workpieces - Operation, cutting off the extra part of the workpiece;
[0055] in, for The number of batches in the batching scheme;
[0056] As a further improvement, the mutation in step S6 adopts four mutation modes:
[0057] 6-2. Perform mutation operation on chromosomes and randomly generate probability p;
[0058] 6-2-1. If 0≤p<0.2, generate a random number r (1≤r≤n), mutate the batching scheme of the rth workpiece of the selected chromosome, and recode its batching scheme to adjust it to a non-batching state;
[0059] 6-2-2. If 0.2≤p<0.4, generate a random number r (1≤r≤n), mutate the batching scheme of the rth workpiece of the selected chromosome, and recode its batching scheme, that is, generate two different random numbers r1, r2 (1≤r1, r2≤length(b r ), r1≠r2), where r1 and r2 represent b r Different sub-batches in the batching scheme are summed up to obtain new sub-batches;
[0060] 6-2-3. If 0.4≤p<0.7, generate a random number r (1≤r≤n), mutate the batching scheme of the rth workpiece of the selected chromosome, and mutate the batching scheme b corresponding to workpiece r. r Re-coding is done by dividing the code into equal parts;
[0061] 6-2-4. If 0.7≤p<1, generate a random number r (1≤r≤n), mutate the batching scheme of the rth workpiece of the selected chromosome, and randomly change the batching scheme of workpiece r by using a two-point mutation method, specifically: randomly select two positions, compare the number of workpieces at the two positions to obtain the number difference Δb, randomly generate a number r1 (0<r1<Δb), increase r1 workpieces for the sub-batch with a small number of workpieces, and reduce r1 workpieces for the sub-batch with a large number of workpieces;
[0062] As a further improvement, the reward design in step S8 is calculated as follows:
[0063]
[0064] Among them, C est (O t ,s t ) is action O t The estimated maximum completion time at time t, C est (O t ,s t+1 ) is action O t The estimated maximum completion time at time t+1, C max is the actual maximum completion time for all actions, C iniis the estimated maximum completion time at time t = 0, which is calculated as: P jf is the processing time of the f-th process for processing the j-th workpiece, and L is the time step for scheduling completion.
[0065] As a further improvement, the generalized advantage estimate GAE is used in the A2C decision network in step S9 to balance the variance and bias of the gradient estimate, which is calculated as follows:
[0066]
[0067] Among them, V(s t ) is the state value function. According to the Bellman equation, the state value function is specifically γ and λ are discount factors. Let γ = 0.99 and λ = 0.98 to better calculate the importance of the current action based on the future state value.
[0068] As a further improvement, in step S9, the dual attention network is used to extract process features and machine features to train the A2C model. The training process is as follows:
[0069] 9-1. Initialization parameters: Initialize actor network parameters θ old , and the critic network parameter δ old , total number of iterations T, discount factors γ, λ, learning rate lr, step size α, β.
[0070] 9-2. Distribution according to current strategy Sampling action a t , get reward r t and the new state s t .
[0071] 9-3.Critic network estimates the current and future state values V(s t+l ), l∈[0,∞), and calculate GAE, and use GAE to update the critic network parameters δ→δ old , reduce the error, and use GAE to calculate the advantage function
[0072]
[0073] Among them, r t is the reward at time t, γ and λ are discount factors, φ(s t ,a t ) is the feature vector of the current state sequence.
[0074] 9-4. Calculate the gradient of the actor network to improve policy performance. Use the updated gradient to update the actor network parameters θ→θold The GAE policy gradient formula is as follows:
[0075]
[0076] 9-5. Update status t →s t+1 , and return to step 9-2. After multiple rounds of iterations, the final scheduling model for the process scheduling problem is obtained. The trained model supports process scheduling problems of different scales and supports fast solution, with good generalization performance.
[0077] Compared with the existing technology, the present invention has the following advantages and effects:
[0078] 1. The present invention divides the dynamic batch scheduling problem of the flexible job shop into two sub-problems: the batch segmentation sub-problem and the sub-batch process sequencing sub-problem, and adopts a hierarchical iterative method to solve the two sub-problems, thereby reducing the complexity of the problem. The batching scheme of the dynamic batch scheduling of the flexible job shop is determined by the genetic algorithm of the outer layer. Based on the batching scheme of the outer layer, the A2C model trained in the inner layer is used to provide a scheduling scheme for the process sequencing of the sub-batch, so as to achieve the goal of minimizing the maximum completion time.
[0079] 2. The present invention designs a flexible job shop scheduling method based on the attention mechanism and the A2C model. The method has good generalization ability and can achieve high-quality solutions when the model trained on small-scale instances is applied to solving large-scale instances.
[0080] 3. The present invention designs six crossover and mutation methods that can optimize batch segmentation during the algorithm iteration process. Compared with the traditional batch scheduling method, the batch segmentation of workpieces can be continuously optimized during the algorithm iteration process, thereby expanding the solution space and achieving the goal of minimizing the maximum completion time with higher efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0081] The present invention will be further described below with reference to the accompanying drawings and implementation examples, in which:
[0082] Figure 1 It is a framework diagram of a new flexible job shop dynamic batch scheduling method;
[0083] Figure 2 It is a schematic diagram of chromosome encoding of genetic algorithm;
[0084] Figure 3 This is a schematic diagram of the genetic algorithm chromosome individual crossover;
[0085] Figure 4 This is a schematic diagram of individual chromosome variation in genetic algorithm;
[0086] Figure 5It is a diagram combining dual attention network and A2C model architecture;
[0087] Figure 6 It is a table showing the correspondence between the decoded workpiece code and the actual workpiece sub-batch. DETAILED DESCRIPTION
[0088] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is further described below in conjunction with the accompanying drawings and embodiments.
[0089] Figure 1 The present invention implements a dynamic batch scheduling method for a flexible job shop based on a genetic algorithm and A2C, and the specific steps are as follows:
[0090] S1: Model the number of machines in the production workshop, different types of workpieces, and the number of workpieces to be processed, and establish a flexible job shop dynamic batch scheduling problem with the goal of minimizing the maximum completion time. The specific steps are as follows:
[0091] 1) The dynamic batch scheduling problem for a flexible job shop of n×m size can be described as follows: n types of workpieces J = {J1, J2, ..., J n}On m machines M={M1,M2,......,M m}, the number of each workpiece is A={A j |j=1,2,......,n}. Each workpiece contains several processes. There are process order constraints between the same workpieces. There are multiple processing machines to choose from for each process. The processing time corresponding to different processing machines is also different. The workpieces are processed in batches, and the number of sub-batches is B={B j |j=1,2,......,n}, and each sub-batch is considered as a whole during production scheduling. The scheduling goal is to minimize the maximum completion time through reasonable workpiece batch division and sub-batch process arrangement.
[0092] 2) Based on the above problem description, the symbolic definition of the mathematical model of the flexible job shop dynamic batch scheduling problem is established:
[0093] n: number of workpiece types;
[0094] m: total number of machines;
[0095] i: machine number, i=1,2,3,...,m;
[0096] j: workpiece type number, j = 1, 2, 3, ..., n;
[0097] A j : The number of the jth type of workpieces;
[0098] B j : The number of sub-batches of the j-th workpiece;
[0099] B max : The maximum batch quantity set by the scheduling;
[0100] f,l: process number, f=1,2,3,…,l=1,2,3,…;
[0101] u: sub-batch number, u=1,2,3,…,B j ;
[0102] N ju : The number of workpieces in the u-th sub-batch of the j-th workpiece;
[0103] P ijf :Machine M i The processing time of the fth process for processing the jth workpiece;
[0104] S ijuf :Machine M i The start time of the f-th operation for processing the u-th sub-batch of the j-th workpiece;
[0105] E ijuf :Machine M i The end time of the f-th operation of processing the u-th sub-batch of the j-th workpiece;
[0106] X ijuf :The fth process of the jth workpiece and the uth sub-batch on machine M i 1 when processing, otherwise 0;
[0107] Y ijufku'l :The fth process of the jth workpiece and the uth sub-batch on machine M i When the start time of the upper processing is earlier than the lth process of the u'th sub-batch of the kth workpiece, the value is 1, otherwise it is 0;
[0108] L: a sufficiently large positive number;
[0109] C max : Maximum completion time;
[0110] 3) According to the above definition, the following mathematical model can be established for the dynamic batch scheduling problem of flexible job shops:
[0111] f=min(C max )
[0112]
[0113] S ijuf ≥E i'ju(f-1)
[0114]
[0115] E ijuf =S ijuf +P ijf ×N ju
[0116] E ijuf ≤S iku'l +L(1-Y ijufku'l )
[0117] 1≤B j ≤B max
[0118] E ijuf ≥0,S ijuf ≥0,P ijf ≥0
[0119] Formula (1) indicates that the scheduling objective is to minimize the maximum completion time; Formula (2) indicates that the total number of workpieces remains unchanged after being divided into batches; Formula (3) indicates that the next process can only be started after the current process of each batch is completed; Formula (4) indicates that each process of each part sub-batch can only be processed by one machine; Formula (5) indicates that the fth process of the uth sub-batch of the jth part is processed on machine M. i The end time of the previous processing is equal to the sum of the start time and the processing time; Formula (6) represents the processing time constraint between sub-batches processed by the same machine; Formula (7) represents that the maximum number of batches of workpieces cannot exceed the set maximum number of batches; Formula (8) represents a non-negative constraint.
[0120] S2: Set the initial parameters of the genetic algorithm, including the population size M and the crossover probability p c And the mutation probability p m , maximum number of iterations iter max ;
[0121] S3: Randomly generate a batching scheme for each workpiece in the workpiece group, and encode the batching scheme corresponding to each workpiece using integer encoding; input the number of each workpiece A j and the maximum batch quantity B set by the scheduling system max Then randomly determine the batching plan b1, b2, ..., n for each workpiece n ,After the division, the number of workpieces in each sub-batch is different, which increases the diversity of the initial solution;
[0122] Where n is the number of workpiece types, batch plan N ju represents the number of workpieces in the u-th sub-batch of the j-th workpiece, b j represents the batching plan for the jth workpiece, each b jThe length of the vector is B max ; The specific batch method is as follows:
[0123] 3-1. Assume that the number of workpieces of the jth type is A j , and set the maximum batch size of the scheduling system to B max ;
[0124] 3-2. Generate a random integer r as the batch number of this type of workpiece, where r∈[1,B max ];
[0125] 3-3. Initialize i=0, b=1, where b represents the current b-th sub-batch of the same type of workpiece (1≤b≤r), and i represents the number of workpieces that have been assigned to the sub-batch.
[0126] 3-4. Generate a random integer s (0<s≤A j -i) as the number of workpieces in the bth sub-batch of type j workpieces, after each sub-batch is allocated, b←b+1;
[0127] 3-5. Reassign i (i = i + s) and determine whether the value of i satisfies the condition i + rb ≤ A j , if satisfied, proceed to step 3-6; otherwise, assign a value to i (i=is) and return to step 3-4;
[0128] 3-6. Determine whether b satisfies the condition b=r-1. If so, the remaining undivided workpieces are divided into the last sub-batch, and the batch division ends; if not, b is assigned a value of b←b+1, and the process returns to step 3-4;
[0129] S4: Encode the batching schemes of all the workpieces in the workpiece group into a chromosome, and combine the batching schemes of all types of workpieces into a chromosome, which is represented by X i =[b1, b2, ..., b n ], X i The length is n×B max ;
[0130] Among them: Chromosome X i An initial integration batching scheme representing the artifact population;
[0131] S5: Determine whether the currently generated chromosome data meets the number required for population initialization. If not, return to step S3. If it meets the conditions, combine multiple chromosomes into one population as the initialization population of the genetic algorithm. Initialize the batch scheme chromosome population X of all types of workpieces = {X1, X2, ..., X M}, and use it as the initialization input of the batching scheme of the workpiece group, M represents the population size;
[0132] S6: According to the chromosome crossover probability p c , mutation probability p m Perform crossover and mutation operations on chromosomes in the population;
[0133] Furthermore, when performing the crossover operation on chromosomes, two crossover methods are designed: first, the probability p (0≤p≤1) is randomly generated;
[0134] If 0≤p<0.5, the crossover method is as follows Figure 3 As shown in (a), a random number r (1≤r≤n) is generated, and the batch scheme of the rth workpiece of the selected chromosome is crossovered two by two, and the parent chromosome Crossover to generate offspring individuals
[0135] If 0.5<p≤1, the crossover method is as follows Figure 3 As shown in (b), two different random numbers r1 and r2 (1≤r1, r2≤n, r1≠r2) are generated, and the batching schemes of the r1-th and r2-th workpieces of the selected chromosomes are crossed in pairs, that is, the difference ΔA between the number of workpieces of the r1-th workpiece of the first chromosome and the r2-th workpiece of the second chromosome is taken as the partial code value, and the batching schemes of r1 and r2 of the two chromosomes are crossed respectively, and the partial code value ΔA is used to uniformly adjust the batching schemes after crossing, that is, the original batch numbers of r1 and r2 of the two chromosomes after crossing remain unchanged, and each sub-batch is adjusted by ± Operation, that is, after crossing, perform + for each sub-batch of the workpiece type r1 and r2 with more workpieces Operation, make up for the missing number of workpieces, and execute for each sub-batch of the workpiece type r1 and r2 with the smaller number of workpieces - Operation, cutting off the extra part of the workpiece;
[0136] Furthermore, four mutation methods are used in the chromosome mutation operation: first, the probability p (0≤p≤1) is randomly generated;
[0137] If 0≤p<0.2, the variation is as follows Figure 4 As shown in (a), a random number r (1≤r≤n) is generated, the batching scheme of the rth workpiece of the selected chromosome is mutated, and its batching scheme is recoded to adjust it to a non-batched state;
[0138] If 0.2≤p<0.4, the variation is as follows Figure 4As shown in (b), a random number r (1≤r≤n) is generated, and the batching scheme of the rth workpiece of the selected chromosome is mutated and its batching scheme is recoded, that is, two different random numbers r1 and r2 (1≤r1, r2≤length(b r ), r1≠r2), where r1 and r2 represent b r Different sub-batches in the batching scheme are summed up to obtain new sub-batches;
[0139] If 0.4≤p<0.7, the variation is as follows Figure 4 As shown in (c), a random number r (1≤r≤n) is generated, and the batching scheme of the rth workpiece of the selected chromosome is mutated, and the batching scheme b corresponding to workpiece r is changed. r Re-coding is done by dividing the code into equal parts;
[0140] If 0.7≤p<1, then the variation is as follows Figure 4 As shown in (d), a random number r (1≤r≤n) is generated, and the batching scheme of the rth workpiece of the selected chromosome is mutated. The batching scheme of workpiece r is randomly changed by a two-point mutation method. Specifically, two positions are randomly selected, the number of workpieces at the two positions is compared to obtain the number difference Δb, a number r1 (0≤r1≤Δb) is randomly generated, r1 workpieces are added to the sub-batch with a small number of workpieces, and r1 workpieces are reduced to the sub-batch with a large number of workpieces;
[0141] S7: Decode the chromosomes to obtain the batching scheme of the workpieces, convert the sub-batch information after the batch division into a standard data format, and provide scheduling data for the subsequent A2C model; regard each decoded sub-batch as an independent processing workpiece, and convert it into a standard BenchMark example format, in which the decoded workpiece encoding part is represented by the accumulation method, such as Figure 6 As shown in the figure, three types of workpieces A, B, and C are divided into 3 sub-batches, 2 sub-batches, and 2 sub-batches, respectively. The first sub-batch of workpiece A is numbered "1", the second and third sub-batches are numbered "2, 3", respectively, and the first sub-batch of workpiece B is numbered "4", and so on. The corresponding sub-batch processing time is the product of the number of workpieces in the sub-batch and the processing time of the independent workpiece. The sub-batch process sequencing problem of workpieces after chromosome decoding is essentially a flexible job shop scheduling FJSP problem.
[0142] S8: Formulate the decoded sub-batch process ordering problem as a Markov decision process;
[0143] Furthermore, the MDP of the present invention is defined as follows:
[0144] 1) Status: Set and is the set of related processes and machines at time t, is the set of process-machine pairs that can be processed, and the state s t is a set of feature vectors, including And each process machine pair Among them, ij is the j-th process of the i-th workpiece.
[0145] 2) Action: A collection of all available process machine pairs It is composed of, which represents the action space at time t.
[0146] 3) Reward: The goal is to minimize the maximum completion time, so the reward is also related to the makespan. The reward r t Defined as That is, the time required to complete all the workpieces is the shortest, so the agent tends to choose the one that can reduce C max The action of the value. Among them, C est (O t ,s t ) is action O t The estimated maximum completion time at time t, C ini is the estimated maximum completion time at time t=0.
[0147] 4) Strategy: Use random strategy π(a t |s t ), which is generated by the subsequent DRL algorithm training.
[0148] S9: Design an A2C model based on dual attention mechanism to extract features and train the model to solve the process sorting problem after batch division of workpieces;
[0149] Going further, Figure 5 To combine the dual attention block and the A2C model architecture diagram, the steps shown in the figure use the dual attention network to extract process features and machine features from the environment. For the process attention block, the specific calculation method is as follows:
[0150]
[0151] Among them, O ij is the input process feature of the jth process of the ith workpiece, and formula (9) is ij The previous process O i,j-1 and post-process O i,j+1 The relationship between the two vectors is modeled and the attention coefficient between adjacent nodes is calculated. The [*||*] operation connects the two vectors to form a larger vector. T For a dimension The weight vector of For a dimension The weight matrix of Represents the set of all processes; Formula (10) is normalized by the softmax function to obtain the normalized attention coefficient α i,j,p ,in is the neighbor set of workpiece node i; Formula (11) transforms the input feature after linear transformation into and Perform a weighted linear transformation to obtain the output feature vector
[0152] For the machine attention block, the process is similar to that of the process attention block, except that it mainly calculates the attention coefficients between different machines. The specific calculation method is as follows:
[0153]
[0154] Among them, c kq Indicates machine M k and M q The set of competing processes. T For a dimension The weight vector of For a dimension The weight matrix of For a dimension The weight matrix of Represents the set of all machines. The average pooling of process features and machine features is the global feature of the FJSP problem instance.
[0155] A2C is a reinforcement learning algorithm based on the policy gradient method. It combines the ideas of Actor-Critic structure and advantage function. It shows good performance in tasks that require continuous action space and complex state space processing, and has high learning efficiency and stability. In step S9, the A2C model is trained based on the features extracted by the dual attention network. The training process is as follows:
[0156] 9-1. Initialization parameters: Initialize actor network parameters θ old , and the critic network parameter δ old , total number of iterations T, discount factors γ, λ, learning rate lr, step size α, β.
[0157] 9-2. Distribution according to current strategy Sampling action a t , get reward r t and the new state s t .
[0158] 9-3.Critic network estimates the current and future state values V(s t+l ), l∈[0,∞), and calculate GAE, and use GAE to update the critic network parameters δ→δ old , reduce the error, and use GAE to calculate the advantage function
[0159]
[0160] Among them, r t is the reward at time t, γ and λ are discount factors, φ(s t ,a t ) is the feature vector of the current state sequence.
[0161] 9-4. Calculate the gradient of the actor network to improve policy performance. Use the updated gradient to update the actor network parameters θ→θ old The GAE policy gradient formula is as follows:
[0162]
[0163] 9-5. Update status t →s t+1 , and return to step 9-2. After multiple rounds of iterations, the final scheduling model for the process scheduling problem is obtained. The trained model supports process scheduling problems of different scales and supports fast solution, with good generalization performance.
[0164] S10: Use the trained A2C model to solve the process sequencing problem after the workpiece batch division, obtain the maximum completion time under different batching schemes, and use the inverse of the maximum completion time as the fitness function value of the genetic algorithm;
[0165]
[0166] S11: Determine whether the termination condition is met. If the termination condition is not met, execute the roulette selection method and return to S6; if the termination condition is met, update the current global optimal solution; the termination condition is that in the past few generations, there is no obvious improvement of the individual, and the optimal fitness value of the population has almost no fluctuation, that is, the algorithm converges;
[0167] S12: Determine whether the maximum number of iterations has been reached. If not, return to step S3 and enter the next cycle. Otherwise, output the optimal batching plan and process scheduling plan.
[0168] The above is only a preferred embodiment of the present invention, and does not limit the present invention in any form. Although the present invention has been disclosed as a preferred embodiment, it is not used to limit the present invention. Any person skilled in the art should be able to use the technical content disclosed above to make equivalent embodiments that are equivalent changes by slight changes or modifications without departing from the scope of the technical solution of the present invention. However, any brief modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. A new method for dynamic batch scheduling of flexible job shops, characterized in that: The following steps are involved: S1: Based on a group of workpieces with different types and different batches of each type having flexible process routes, a flexible job shop dynamic batch scheduling model is established with the goal of minimizing the maximum completion time; S2: Set the initial parameters of the genetic algorithm, including the population size M and the crossover probability p c And the mutation probability p m , maximum number of iterations iter max ; S3: Randomly generate a batching scheme for each workpiece in the workpiece group, and encode the batching scheme corresponding to each workpiece using integer encoding; input the number A of each workpiece j j and the maximum batch quantity B set by the dispatching system max Then randomly determine the batching plan b1, b2, ..., b for each workpiece n ; Where n is the number of workpiece types, batch plan N ju represents the number of workpieces in the u-th sub-batch of the j-th workpiece, b j represents the batching plan for the jth workpiece, each b j The length of the vector is B max ; S4: Encode the batching schemes of all the workpieces in the workpiece group into a chromosome, and combine the batching schemes of all types of workpieces into a chromosome, which is represented by X i =[b1, b2, ..., b n ], X i The length is n×B max ; Among them, chromosome X i An initial integration batching scheme representing the artifact population; S5: Determine whether the currently generated chromosome data meets the number required for population initialization. If not, return to step S3. If it meets the conditions, combine M chromosomes into a population as the initialization population of the genetic algorithm; initialize the batch scheme chromosome population X of all types of workpieces = {X1, X2, ..., X M }, and use it as the initialization input of the batching scheme of the workpiece group, M represents the population size; S6: According to the chromosome crossover probability p c , mutation probability p m Perform crossover and mutation operations on chromosomes in the population; S7: Decode the chromosome to obtain the batching plan of the workpiece, convert the sub-batch information after the batch division into a standard data format, and provide scheduling data for the subsequent A2C model; S8: The problem of sorting the workpiece processes of all sub-batches obtained after decoding the workpiece group represented by each chromosome is expressed as a Markov decision process; S9: Use the dual attention mechanism network to extract process features and machine features to train the A2C model to solve the process sorting problem after batch division of workpieces; S10: using the A2C model trained in step S9 to solve the process sequencing problem after the workpiece batch division, obtain the optimal completion time under different batching schemes, and use the reciprocal of the maximum completion time as the fitness function value of the genetic algorithm; S11: If the termination condition is not met, the roulette selection method is executed and then the process returns to S6; if the termination condition is met, the current global optimal solution is updated; the termination condition is that in the past few generations, there is no obvious improvement in the individual and the optimal fitness value of the population has almost no fluctuation, that is, the algorithm converges; S12: If iter <iter max , then return to S3; otherwise, output the optimal batching plan and process scheduling plan; Among them, iter is the number of current iterations.
2. A new flexible job shop dynamic batch scheduling method according to claim 1, characterized in that: The mathematical model establishment process of the dynamic batch scheduling of the flexible job shop in step S1 is as follows: 2-1. The dynamic batch scheduling problem for a flexible job shop of n×m size can be described as follows: n types of workpieces J = {J1, J2, ..., J n }On m machines M={M1,M2,......,M m }, the number of each workpiece is A={A j |j=1,2,......,n}; Each workpiece contains several processes. There are process order constraints between the same workpieces. There are multiple processing machines to choose from for each process. The processing time corresponding to different processing machines is also different. The workpieces are processed in batches, and the number of sub-batches is B={B j |j=1,2,......,n}, and each sub-batch is considered as a whole during production scheduling; the scheduling goal is to minimize the maximum completion time through reasonable workpiece batch division and sub-batch process arrangement; 2-2. Based on the above problem description, the symbolic definition of the mathematical model of the flexible job shop batch scheduling problem is established: n: number of workpiece types; m: total number of machines; i: machine number, i=1,2,3,...,m; j: workpiece type number, j = 1, 2, 3, ..., n; A j : The number of the jth type of workpieces; B j : The number of sub-batches of the j-th workpiece; B max : The maximum batch quantity set by the scheduling; f,l: process number, f=1,2,3,..., l=1,2,3,...; u: sub-batch number, u=1,2,3,...,B j ; N ju : The number of workpieces in the u-th sub-batch of the j-th workpiece; P ijf :Machine M i The processing time of the fth process for processing the jth workpiece; S ijuf :Machine M i The start time of the f-th operation for processing the u-th sub-batch of the j-th workpiece; E ijuf :Machine M i The end time of the f-th operation of processing the u-th sub-batch of the j-th workpiece; X ijuf :The fth process of the jth workpiece and the uth sub-batch on machine M i 1 when processing, otherwise 0; Y ijufku'l :The fth process of the jth workpiece and the uth sub-batch on machine M i When the start time of the upper processing is earlier than the lth process of the u'th sub-batch of the kth workpiece, the value is 1, otherwise it is 0; L: a sufficiently large positive number; C max : Maximum completion time; 2-3. Based on the above definition, the following mathematical model can be established for the dynamic batch scheduling problem of flexible job shops: f=min(C max ) (1) S ijuf ≥E i'ju(f-1) (3) E ijuf =S ijuf +P ijf ×N ju (5) E ijuf ≤S iku'l +L(1-Y ijufku'l ) (6) 1≤B j ≤B max (7) E ijuf ≥0,S ijuf ≥0,P ijf ≥0 (8) Formula (1) indicates that the scheduling objective is to minimize the maximum completion time; Formula (2) indicates that the total number of workpieces remains unchanged after batching; Formula (3) indicates that the next process can only be started after the current process of each batch is completed; Formula (4) indicates that each process of each part sub-batch can only be processed by one machine; Formula (5) indicates that the fth process of the uth sub-batch of the jth part is processed on machine M. i The end time of the previous processing is equal to the sum of the start time and the processing time; Formula (6) represents the processing time constraint between sub-batches processed by the same machine; Formula (7) represents that the maximum number of batches of workpieces cannot exceed the set maximum number of batches; Formula (8) represents a non-negative constraint.
3. A new flexible job shop dynamic batch scheduling method according to claim 1, characterized in that: The crossover in step S6 adopts two crossover modes: Perform a crossover operation on the chromosome and randomly generate a probability p (0≤p≤1); 3-1. If 0≤p<0.5, generate a random number r (1≤r≤n), perform a pairwise crossover on the batch scheme of the rth workpiece of the selected chromosome, and perform a pairwise crossover on the parent chromosome. Crossover to generate offspring individuals 3-2. If 0.5<p≤1, two different random numbers r1 and r2 are generated (1≤r1, r2≤n, r1≠r2), and the batching schemes of the r1th and r2th workpieces of the selected chromosomes are crossed in pairs, that is, the difference ΔA between the number of workpieces of the r1th workpiece of the first chromosome and the r2th workpiece of the second chromosome is taken as the bias code value, and the batching schemes of r1 and r2 of the two chromosomes are crossed respectively, and the bias code value ΔA is used to uniformly adjust the batching schemes after crossing, that is, the batch numbers of the original chromosome workpiece types r1 and r2 remain unchanged after crossing, and each sub-batch is adjusted by ± Operation, that is, after crossing, perform the operation on each sub-batch of the workpiece type r1 and r2 with more workpieces Operation, to make up for the missing number of workpieces, is performed on each sub-batch of the workpiece type r1 and r2 with the smaller number of workpieces. Operation, cutting off the extra part of the workpiece; in, for The number of batches for the batching scheme.
4. The new flexible job shop dynamic batch scheduling method according to claim 1 is characterized in that: The mutation in step S6 adopts four mutation modes: Perform mutation operation on chromosomes and randomly generate probability p; 4-1. If 0≤p<0.2, generate a random number r (1≤r≤n), mutate the batching scheme of the rth workpiece of the selected chromosome, and recode its batching scheme to adjust it to a non-batching state; 4-2. If 0.2≤p<0.4, generate a random number r(1≤r≤n), mutate the batching scheme of the rth workpiece of the selected chromosome, and recode its batching scheme, that is, generate two different random numbers r1, r2(1≤r1, r2≤length(b r ), r1≠r2), where r1 and r2 represent b r Different sub-batches in the batching scheme are summed up to obtain new sub-batches; 4-3. If 0.4≤p<0.7, generate a random number r (1≤r≤n), mutate the batching scheme of the rth workpiece of the selected chromosome, and mutate the batching scheme b corresponding to workpiece r. r Re-coding is done by dividing the code into equal parts; 4-4. If 0.7≤p<1, generate a random number r (1≤r≤n), mutate the batching scheme of the rth workpiece of the selected chromosome, and randomly change the batching scheme of workpiece r by two-point mutation. Specifically, randomly select two positions, compare the number of workpieces at the two positions to obtain the quantity difference Δb, randomly generate a number r1 (0<r1<Δb), increase r1 workpieces for the sub-batch with a small number of workpieces, and reduce r1 workpieces for the sub-batch with a large number of workpieces.
5. The new flexible job shop dynamic batch scheduling method according to claim 1 is characterized in that: The specific calculation process of the reward design in step S8 is as follows: Among them, C est (O t ,s t ) is action O t The estimated maximum completion time at time t, C est (O t ,s t+1 ) is action O t The estimated maximum completion time at time t+1, C max is the actual maximum completion time for all actions, C ini is the estimated maximum completion time at time t = 0, which is calculated as follows: P jf is the processing time of the f-th process for processing the j-th workpiece, and L is the time step for scheduling completion.
6. The new flexible job shop dynamic batch scheduling method according to claim 1 is characterized in that: The generalized advantage estimate GAE is used in the A2C decision network in step S9 to balance the variance and bias of the gradient estimate, which is calculated as follows: Among them, V(s t ) is the state value function. According to the Bellman equation, the state value function is specifically γ and λ are discount factors, and γ = 0.99 and λ = 0.98 are set to better calculate the importance of the current action based on the future state value.
7. The new flexible job shop dynamic batch scheduling method according to claim 1 is characterized in that: In step S9, the dual attention network is used to extract process features and machine features to train the A2C model. The training process is as follows: 7-1. Initialization parameters: Initialize actor network parameters θ old , and the critic network parameter δ old , total number of iterations T, discount factors γ, λ, learning rate lr, step size α, β; 7-2. Distribution based on current strategy Sampling action a t , get reward r t and the new state s t ; 7-3.Critic network estimates the current and future state values V(s t+l ), l∈[0,∞), and calculate GAE, and use GAE to update the critic network parameters δ→δ old , reduce the error, and use GAE to calculate the advantage function Among them, r t is the reward at time t, γ and λ are discount factors, φ(s t ,a t ) is the feature vector of the current state sequence; 7-4. Calculate the gradient of the actor network to improve policy performance; use the updated gradient to update the actor network parameters θ→θ old ; The GAE policy gradient formula is as follows: 7-5. Update status t →s t+1 , and return to step 7-2. After multiple rounds of iterations and convergence, the final scheduling model for the process scheduling problem is obtained; the trained model supports process scheduling problems of different scales and supports fast solution, with good generalization performance.
Citation Information
Cited By
Discrete manufacturing workshop variable batch flexible scheduling method and system
CN120430599A