Multi-target generalized job shop scheduling method considering same-machine operation
By considering the same-machine operation in generalized operation workshop scheduling, combining reinforcement learning and genetic algorithms, the multi-objective optimization problem of forced and flexible parallel batch processing constraints is solved, and the workshop production efficiency and cost reduction are achieved.
Patent Information
- Application Number
- CN202510084652.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to consider both mandatory and flexible parallel batch processing constraints at the same time, resulting in the inability to effectively optimize multi-objective generalized operation workshop scheduling.
A multi-objective generalized operation workshop scheduling method considering the same-machine operation is proposed. By establishing a generalized operation workshop scheduling model, designing the encoding and decoding of processes and machines, using a hybrid heuristic strategy to initialize the population, and introducing cross-mutation operations that adapt to forced and flexible parallel batch processing, adjusting the cross-mutation probability based on reinforcement learning.
The optimization of multiple goals such as maximum completion time, total delay and total carbon emissions has been achieved, which has improved workshop production efficiency, reduced costs, and achieved energy conservation and emission reduction.
Smart Images

Figure CN120069395A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of job shop scheduling optimization, and more particularly, to a generalized job shop scheduling optimization method considering parallel batch processing. Background Art
[0002] In traditional job shop scheduling (JSP), a single machine can only process one workpiece at a time, and the workpiece monopolizes the selected machine during operation. However, in actual production, situations such as the combined processing of mold accessories and the joint inspection of electronic products bring about mandatory parallel batch processing constraints. In the electronic product inspection workshop, different products are batch-tested for high-temperature resistance, and workpieces with similar material property requirements are heated in a single heat treatment furnace, etc., which present flexible parallel batch processing. These situations break the traditional constraints of job shop scheduling. Existing research on the flexible job shop scheduling problem with parallel batch processing often has difficulty in simultaneously handling the constraints involving both mandatory and flexible parallel batch processing tasks, and most of the research focuses on single-objective optimization.
[0003] Reinforcement learning autonomously determines the action strategy through the interaction between the agent and the environment. Each time the agent makes an action strategy, it can obtain the reward feedback from the environment. If it is a positive feedback, the probability of making the same strategy again will increase, otherwise it will decrease. The ultimate training goal is to obtain the maximum environmental reward. The third-generation non-dominated sorting genetic algorithm (NSGA-III) is a genetic algorithm based on the reference point mechanism proposed by K. Deb et al. in 2014. By introducing a widely distributed reference point to maintain the diversity of the population, it is a multi-objective optimization algorithm widely used in the field of workshop scheduling. Summary of the Invention
[0004] The purpose of the present invention is to solve the problem of the multi-objective generalized job shop scheduling optimization problem that cannot solve the constraints considering both mandatory and flexible parallel batch processing, and to provide a multi-objective generalized job shop scheduling method considering same-machine operation, so as to realize the optimization of multiple objectives such as the makespan, total tardiness, and total carbon emissions, improve the workshop production efficiency, reduce costs, and achieve energy conservation and emission reduction.
[0005] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0006] A multi-objective generalized job shop scheduling method considering same-machine operation, comprising the following steps:
[0007] S1: Establish a generalized job shop scheduling model to describe the generalized job shop scheduling problem considering forced and flexible parallel batch processing, and determine the objective function and constraints of the generalized job shop scheduling model;
[0008] S2: Based on the constraints of the generalized job shop scheduling considering forced and flexible parallel batch processing, design the encoding and decoding of processes and machines, and initialize the population using a hybrid heuristic strategy;
[0009] S3: The algorithm introduces crossover and mutation operations adapted to forced and flexible parallel batch processing, adjusts the crossover and mutation probabilities based on reinforcement learning, and obtains a new generation of population through crossover and mutation for iteration;
[0010] S4: Until the algorithm reaches the maximum number of iterations, the generalized job shop scheduling model outputs the optimization results to obtain the Gantt chart of the generalized job shop scheduling with forced and flexible parallel batch processing.
[0011] Furthermore, establish a generalized job shop scheduling model, specifically:
[0012] The generalized job shop scheduling problem considering forced and flexible parallel batch processing is described as follows: N jobs are processed on M machines. Each job contains multiple processes, and there are multiple machines available for each process. The processing time of each job process depends on the selected machine. It is allowed to process multiple jobs simultaneously on the same machine, with forced or flexible same-machine operation. Under the constraints of machine availability, process sequence, and flexible same-machine operation, determine the sequential processing order, the machines used, and the processing time of all processes of all jobs, and obtain the optimization of the makespan, total tardiness, and carbon emission scheduling indicators, and reduce the conflicts between different objectives;
[0013] According to the actual workshop production situation, the following assumptions also need to be met:
[0014] At the initial moment, all machines are in a usable state, all jobs can also start processing immediately, and there is no priority difference between individual jobs. For each process of the same job, its processing sequence is determined according to the established process route. At the same time point, only one process of a job can be in the processing process. Any machine can only process one process at the same moment. The time consumed by auxiliary operations is included in the labor quota of the process. When a process starts processing on the selected machine, only the availability of the current machine and whether the previous process of this process has been completed need to be considered. During the processing of the process, the entire process flow cannot be interrupted.
[0015] Furthermore, the generalized job shop scheduling model includes the following mathematical symbols:
[0016] Index variables: i and h represent the workpiece numbers, where i, h ∈ {1, 2,..., N}; j and g represent the operation numbers, where j, g ∈ {1, 2,..., N i}; k represents the machine number, where k ∈ {1, 2,..., M}; x represents the serial number of the parallel batch processing operation set; N represents the total number of workpieces; M represents the total number of machines; N i represents the total number of operations of workpiece i; I and H represent the workpiece set numbers corresponding to the combined operation; J and G represent the operation set numbers corresponding to the combined operation; U represents an infinite real number; N MPBPO represents the number of operations in the forced parallel batch processing operation set; N FPBPO represents the number of operations in the flexible parallel batch processing operation set;
[0017] Parameter variables: O ij represents the j-th operation of workpiece i; O IJ represents the combined operation, O IJ ={O ij , O hg , O pg ,..., O yz}; T i represents the delivery time of workpiece i; represents the operation time of operation O IJ on machine k; represents the operation time of operation O ij on machine k; P k represents the operation power of machine k; JS IJ represents the workpiece set corresponding to the combined operation O IJ ; OS IJ represents the operation set corresponding to the combined operation O IJ ; MPBPO x represents the x-th forced parallel batch processing operation set; FPBPO x represents the x-th flexible parallel batch processing operation set; MS k represents the startup time of machine k; ME k represents the shutdown time of machine k; Mac ij represents the optional machine set of O ij ; Cap k represents the capacity of machine k; represents the weight of O ij on machine k; represents the no-load power of machine k; E t represents the transfer energy consumption of a single operation of a workpiece; P g represents the fixed power of the workshop;
[0018] Decision variables: C iDenote the completion time of workpiece i; E busy Denote the total energy consumption of machine operation time; E idle Denote the total energy consumption of machine idle time; E tran Denote the total energy consumption of workpiece transfer; E innate Denote the inherent energy consumption of the workshop; Denote operation O IJ The operation on machine k is 1, otherwise 0; Denote operation O ij The operation at position l is 1, otherwise 0; Denote operation O ij The immediate predecessor operation O i(j-1) The operation on machine k is 1, otherwise 0; Denote operation O ij Belonging to MPBPO x The operation in it is 1, otherwise 0; Denote operation O ij Belonging to FPBPO x The operation in it is 1, otherwise 0; S ij Denote operation O ij The start operation time of operation O; S IJ Denote operation O IJ The start operation time of operation O; E IJ Denote operation O IJ The end operation time of operation O; E ij Denote operation O ij The end operation time of operation O; Y IJHG Denote operation O IJ Preceding O HG When operating is 1, otherwise 0; CEF Carbon emission conversion factor.
[0019] Furthermore, the objective function of the generalized job shop scheduling model includes:
[0020] Minimize the makespan C max This objective function is expressed as:
[0021]
[0022] Minimize the total tardiness TD, and this objective function is defined as:
[0023]
[0024] The carbon emission CE is obtained by multiplying the total energy consumption E of the machine operation state busy , the total energy consumption E of the idle state idle , the total energy consumption E of workpiece transfer tran and the inherent energy consumption E of the workshop innate by the conversion factor CEF, and this objective function is expressed as:
[0025] f3 = CE = min{CEF·(E busy +E idle +E tran +E innate )} (3)
[0026] E busy is the total energy consumption consumed by all machines for operating on the workpiece, expressed as:
[0027]
[0028] E idle is the total energy consumption of all machines in the standby state, i.e., no-load operation, expressed as:
[0029]
[0030] E tran is the total energy consumption for workpiece transfer, and this objective is expressed as:
[0031]
[0032] E innate is the common energy consumption of the workshop, and this objective is expressed as:
[0033] E innate = P g C max (7).
[0034] Furthermore, the constraint conditions of the generalized job-shop scheduling model are:
[0035]
[0036] Equation (8) ensures that there will be no interruption in the operation process of the process; Equation (9) ensures that the completion time of any workpiece will not exceed the maximum required completion time; Equation (10) ensures that the start time of the current process for operation is not lower than the completion time of its immediate predecessor process; Equation (11) ensures that the start time of any process for operation does not exceed its end time for operation; Equation (12) ensures that any process only performs one operation on the corresponding machine; Equation (13) ensures that the start time of any process for operation is not negative, and its operation time and end time for operation are both greater than zero; Equation (14) ensures that only one process is in the operation state on one machine at the same time; Equation (15) ensures that the processes belonging to the same mandatory parallel batch processing process set must be batch processed simultaneously; Equation (16) ensures that the processes belonging to the same flexible parallel batch processing process set can be processed separately or randomly combined for batch processing.
[0037] Furthermore, design the coding of machines and processes, specifically:
[0038] The multi-objective generalized job shop scheduling with on-machine parallel batch processing considers operation sequencing and machine selection, selects operation coding, and the coding includes four segments of integer coding sequences; the first segment is the operation coding sequence, each gene is represented by the workpiece serial number, workpiece i, i ∈ [1, N], the occurrence frequency j, j ∈ [1, N i , O ij represents the corresponding operation of the workpiece; for the combined operations coupled by on-machine operations, they are described by the workpiece set, and the corresponding operations take the operations corresponding to each workpiece in the workpiece set; the second segment records the on-machine parallel operation information, the coding of ordinary operations is 0, and the coding of parallel batch processing operations is a number greater than 0; the third segment is the machine coding, and its optional range is [1, M], and the machine coding order corresponds to the operation coding, indicating the operation machine specified by the corresponding operation; the fourth segment is the operation time coding, corresponding to the previous operation coding and machine coding, indicating the operation time required for the corresponding operation to operate on the specified machine.
[0039] Furthermore, in step S2, design the decoding of machines and operations, specifically:
[0040] To obtain a high-quality solution and narrow the search space range, design an active decoding method, and the steps of the active decoding method are as follows:
[0041] S211: Set A O (1, 1:ON), A O (2, 1:ON) store the start and completion times of each operation respectively, ON = |OPlan|; set MP = {MP[1], MP[2],..., MP[M]}, JP = {JP[1], JP[2],..., JP[J]} to record the end times of the immediate predecessors of each machine and workpiece respectively; initialize all elements in MP and JP to 0, k = 1;
[0042] S212: Sequentially read the gene position elements of OPlan[k], 1 ≤ k ≤ |OPlan|, determine its operation O, and obtain the corresponding operation machine m and operation time Pt of O from MPlan and TmPlan respectively O , calculate
[0043] S213: Find the idle time interval [t_s, t_e] of machine m, if max(as O , t_s) + Pt O ≤ t_e, go to step S214, otherwise go to step S215;
[0044] S214: S O = max(as O , t_s), E O = SO +Pt O ; If (S O ==t_s) && (E O ==t_e), delete the interval [t_s, t_e]; If (S O ==t_s) && (E O <t_e), update [t_s, t_e] = [E O , t_e]; If (t_s < S O ) && (E O ==t_e), update [t_s, t_e] = [t_s, S O ; If (t_s < S O ) && (E O <t_e), update [t_s, t_e] = [t_s, S O and insert a new [t_s', t_e'] = [E O , t_e]
[0045] S215: E O = S O +Pt O ; If as O > MP[m], add a new idle interval [t_s', t_e'] = [MP[m], as O on machine m; MP[m] = E O ,
[0046] S216: Update A O (1, k) = S O , A O (2, k) = E O ;
[0047] S217: k = k + 1, if k ≤ |OPlan|, go to step S212, otherwise end.
[0048] Furthermore, in step S2, a hybrid heuristic strategy is used to initialize the population, specifically:
[0049] Design a hybrid initialization method for multi-objective problems. First, randomly generate operation codes, and adopt various heuristic strategies and random selection rules when allocating machines. The steps for initializing population individuals are as follows:
[0050] S221: Input the forced parallel batch processing operation set and the flexible parallel batch processing operation set, initialize the population size NP, define an array NP × 3, and the counter np = 1;
[0051] S222: Generate a new empty process sequence OPlan, coupling sequence PPlan, machine sequence MPlan and operation time sequence TPlan, and set all numbers in [1, N] to unmarked state;
[0052] S223: If all [1, N] are marked, go to step S227; otherwise, randomly select an unmarked number i in [1, N], determine the frequency j of i appearing in OPlan, and if j ≥ N i , mark i and go to step S223, if j<N i , from different workpiece process routes PPlan i ,1≤i≤N to obtain O ij , if O ij For the conventional process, add j to the end of OPlan and add 0 to the end of PPlan; if O ij To force parallel batch process elements, go to step S224. ij For a flexible parallel batch process element, go to step S225;
[0053] S224: Get Corresponding artifact set I x ; If MPBPO x If there is a predecessor parallel batch process in the process, go to step S223, otherwise x Each workpiece number is bound and placed at the end of OPlan, and the corresponding workpiece set is placed at the end of PPlan x Quantity 1, go to step S226;
[0054] S225: Select l x Quantity process composition flexible parallel batch process set Based on Select a machine set ks that satisfies the device capacity constraint if Go to step S225, otherwise obtain Corresponding artifact set I x ;like If there is a predecessor parallel batch process in the process, go to step S223. Otherwise, x Each workpiece number is bound and placed at the end of OPlan, and the corresponding workpiece set is placed at the end of PPlan x Quantity 1 and in FPBPO x The artifacts selected to be placed at the end of OPlan will be deleted;
[0055] S226: Get O ijThe corresponding set of parallel batch processing operations O and the set of precursor operations JP[O] of O that are not placed in the scheduling sequence are obtained. It is determined whether there are any precursor parallel batch processing operations in O. If there are no precursor parallel batch processing operations for the workpiece corresponding to O, the operations in O ij ∈IP[O] corresponding to the workpiece number i are randomly inserted before the workpiece corresponding to O in OPlan; otherwise, they are randomly inserted between the closest precursor parallel batch processing operation to i and the workpiece corresponding to O, and then go to step S223;
[0056] S227: Generate a random number r between (0, 1). If r < 0.25, for the operation O in OPlan[i], where 1 ≤ i ≤ |OPlan|, select as the operation machine according to active decoding; if 0.25 ≤ r < 0.5, select as the operation machine according to active decoding; if 0.5 ≤ r ≤ 0.75, select as the operation machine according to active decoding; if r > 0.75, then randomly select m from Mac O as the operation machine for operation O to form the machine sequence MPlan; determine the operation time of each operation, where the operation time of the parallel batch processing operation is the maximum value of the operation times of each operation in the parallel batch processing operation on the machine, to form the operation time sequence TmPlan; store Pop(np, 1) = OPlan, Pop(np, 2) = MPlan, Pop(np, 3) = TmPlan;
[0057] S228: If np < NP, np = np + 1, and go to step S222; otherwise, end.
[0058] Furthermore, the algorithm introduces crossover and mutation operations that adapt to forced and flexible parallel batch processing, specifically:
[0059] Design sequential crossover one by one. The steps of the sequential crossover operation one by one are as follows:
[0060] S31: Select parent generations P1 and P2, and initialize the child generation CP operation sequence, machine sequence, and operation time sequence to be empty;
[0061] S32: Generate a random number r between (0, 1). If r < 0.5, then select the first gene position element of the P1 operation sequence; otherwise, select the first gene position element of the P2 operation sequence;
[0062] S33: Add the selected element to the end position of the CP operation sequence, and at the same time add the machine and time corresponding to this element to the end of the CP machine sequence and the operation time sequence respectively; remove the gene where the element first appears in the P1 and P2 operation sequences, and delete the machine and time genes corresponding to this element;
[0063] S34: Repeat steps S32 and S33 until the elements in parents P1 and P2 are empty, then the process crossover is completed.
[0064] Furthermore, the crossover and mutation probabilities are adjusted based on reinforcement learning, specifically as follows:
[0065] Set conversion conditions, combine the advantages of the SARSA algorithm and the Q - learning algorithm, and apply SARSA and Q - learning to different stages of the algorithm respectively; in the initial stage of the algorithm, the Q - value is updated based on SARSA and recorded in the Q - value table, and these Q - values will be used as the pre - training process of the Q - learning algorithm; if the conversion condition is met, that is, the number of iterations is greater than the product of the total number of states and the total number of actions, the SARSA algorithm will be converted to the Q - learning algorithm.
[0066] When using reinforcement learning to optimize the crossover and mutation rates and represent the corresponding states, the distribution diversity DV and coverage SC of the solutions are used as key indicators to reflect the quality, coverage range, and balance of the solution set; set four states: S 1 , both the diversity and coverage increase; S 2 , the diversity increases while the coverage decreases; S 3 , the diversity decreases while the coverage increases; S 4 , both the diversity and coverage decrease;
[0067] For each iteration, the agent takes different actions to obtain appropriate crossover and mutation rates. The action set π contains 8 actions. Among them, the crossover probability is between 0.6 - 0.95, the mutation probability is between 0.05 - 0.4, and the intermediate difference is 0.05 for both; the action selection adopts the ε - greedy strategy, as shown in equation (17), where ε is the greedy rate, and r 0-1 is a random value from 0 to 1; when ε ≤ r 0-1 , select the action a corresponding to the maximum value in the Q - value table, otherwise randomly select an action a;
[0068]
[0069] After executing the action, evaluate the quality of the action according to the newly entered state. The effect of the action is positive or negative, corresponding to the corresponding positive and negative reward values respectively; the reward value is set as;
[0070]
[0071] Regard the genetic algorithm part as the environment, and dynamically adjust the crossover and mutation probabilities through the agent. The specific steps are as follows:
[0072] S321: Feed the environmental information back to the agent, and input the fitness value information F of all individuals in the population. iter ;
[0073] S322: Judge the current population iteration times iter. If iter = 1, initialize the Q-value table and define the initial action a. 1 , and go to step S327; if iter > 1, calculate the reward value r. iter ;
[0074] S323: Judge the next state s through the changes in diversity and coverage rate. iter+1 ;
[0075] S324: Judge whether the conversion condition is satisfied. If it is satisfied, update the Q-value table according to the Q-learning formula, and select the next action a through the greedy strategy. iter+1 ; if it is not satisfied, update the Q-value table according to SARSA, and select the next action a through the ε-greedy strategy. iter+1 ;
[0076] S325: Update the current state value s. iter ← s iter+1 , update the current action a. iter ← a iter+1 ;
[0077] S326: Update the fitness value information F of the previous generation. iter-1 ← F iter ;
[0078] S327: The agent makes a decision and updates the crossover and mutation probabilities according to the current action a. iter Update the crossover and mutation probabilities.
[0079] Compared with the prior art, the present invention proposes a knowledge-driven non-dominated sorting genetic algorithm combined with reinforcement learning, which integrates forced and flexible parallel batch tasks, realizes the optimization of multiple objectives such as the makespan, total tardiness, and total carbon emissions, improves the workshop production efficiency, reduces costs, and achieves energy conservation and emission reduction. Brief Description of the Drawings
[0080] Figure 1 It is a flowchart of a multi-objective generalized job shop scheduling method considering same-machine operations.
[0081] Figure 2 It is a coding schematic diagram.
[0082] Figure 3 It is a schematic diagram of OOOX crossover.
[0083] Figure 4 It is a schematic diagram of operation sequence interchange mutation.
[0084] Figure 5 Schematic diagram of single-point mutation for the machine.
[0085] Figure 6 Schematic diagram of recombination mutation for flexible combined processes.
[0086] Figure 7 Schematic diagram of randomly selecting PBPO and moving it forward as much as possible.
[0087] Figure 8 Gantt chart of the CMK05 example scheduling using the multi-objective generalized job shop scheduling method. Specific implementation manners
[0088] The following further describes the multi-objective generalized job shop scheduling method considering same-machine operations of the present invention with reference to the accompanying drawings and specific embodiments.
[0089] Please refer to Figure 1 , the present invention discloses a multi-objective generalized job shop scheduling method considering same-machine operations, including the following steps:
[0090] S1: Establish a generalized job shop scheduling model to describe the generalized job shop scheduling problem considering forced and flexible parallel batch processing, and determine the objective function and constraint conditions of the generalized job shop scheduling model;
[0091] S2: Based on the constraints of the generalized job shop scheduling considering forced and flexible parallel batch processing, design the encoding and decoding of processes and machines, and initialize the population using a hybrid heuristic strategy;
[0092] S3: The algorithm introduces crossover and mutation operations adapted to forced and flexible parallel batch processing, adjusts the crossover and mutation probabilities based on reinforcement learning, and obtains a new generation of population through crossover and mutation for iteration;
[0093] S4: Until the algorithm reaches the maximum number of iterations, the generalized job shop scheduling model outputs the optimization result to obtain the Gantt chart of the generalized job shop scheduling with forced and flexible parallel batch processing.
[0094] Based on the third-generation non-dominated sorting genetic algorithm, the present invention proposes a knowledge-driven non-dominated sorting genetic algorithm enhanced with reinforcement learning (a Knowledge-driven Non-dominated Sorting Genetic Algorithm III enhanced with Reinforcement Learning, KNSGA-III-RL). While optimizing multiple objectives (including makespan, carbon emissions, and total tardiness), it integrates forced and flexible parallel batch processing tasks.
[0095] Step S1, establish a generalized job shop scheduling model to describe the generalized job shop scheduling problem considering forced and flexible parallel batch processing, and determine the objective function and constraints of the generalized job shop scheduling model.
[0096] The generalized job shop scheduling problem considering forced and flexible parallel batch processing is described as follows: N jobs {1, 2, …, N} are processed on M machines {1, 2, …, M}. Each job contains multiple operations, and for each operation, multiple machines can be selected. Moreover, the operation time of each job operation depends on the selected machine. It is allowed to process multiple jobs simultaneously on the same machine, with forced or flexible same-machine operations. Under the constraints of machine availability, operation sequence, and flexible same-machine operations, determine the sequence of operations, the machines used, and the operation time for all operations of all jobs, in order to optimize the makespan, total tardiness, and carbon emissions scheduling metrics, and reduce the conflicts between different objectives.
[0097] According to the actual workshop production situation, the following assumptions also need to be satisfied:
[0098] At the initial moment, all machines are in a usable state, all jobs can also start processing immediately, and there is no priority difference between individual jobs;
[0099] For the operations of the same job, their operation sequence is determined according to the established process route. At the same time point, only one operation of a job can be in the processing process;
[0100] Any machine can only process one (combined) operation at the same moment;
[0101] The time consumed by auxiliary operations such as clamping, disassembling, tool changing, and handling is included in the operation time quota of the operation;
[0102] When a (combined) operation starts on the selected machine, only the availability of the current machine and whether the previous operation of the (combined) operation has been completed need to be considered. The material matching, auxiliary tools, etc. have been prepared before the machine starts the operation;
[0103] During the operation of the operation, ensure the continuous and stable operation of the machine, and the operation process of the entire operation should not be interrupted.
[0104] The generalized job shop scheduling model includes the following mathematical symbols:
[0105] Index variables: i, h represent the job numbers, i, h ∈ {1, 2,..., N}; j, g represent the operation numbers, j, g ∈ {1, 2,..., N i}; k represents the machine number, k ∈ {1, 2,..., M}; x represents the serial number of the parallel batch processing operation set; N represents the total number of workpieces; M represents the total number of machines; N i represents the total number of operations of workpiece i; I, H represent the serial numbers of the workpiece sets corresponding to the combined operations; J, G represent the serial numbers of the operation sets corresponding to the combined operations; U represents an infinitely large real number; N MPBPO represents the number of operations in the forced parallel batch processing operation set; N FPBPO represents the number of operations in the flexible parallel batch processing operation set.
[0106] Parameter variable: O ij represents the j-th operation of workpiece i; O IJ represents the combined operation, O IJ ={O ij , O hg , O pg ,..., O yz}; T i represents the delivery time of workpiece i; represents the operation O IJ 's operation time on machine k; represents the operation O ij 's operation time on machine k; P k represents the operating power of machine k; JS IJ represents the combined operation O IJ corresponding workpiece set; OS IJ represents the combined operation O IJ corresponding operation set; MPBPO x represents the x-th forced parallel batch processing operation set; FPBPO x represents the x-th flexible parallel batch processing operation set; MS k represents the startup time of machine k; ME k represents the shutdown time of machine k; Mac ij represents O ij 's optional machine set; Cap k represents the capacity of machine k; represents O ij 's weight on machine k; represents the no-load power of machine k; E t represents the transfer energy consumption of a single operation of the workpiece; P g represents the fixed power of the workshop.
[0107] Decision variable: C i represents the completion time of workpiece i; E busy represents the total energy consumption of the machine operation time; E idle represents the total energy consumption of the machine no-load time; E tran represents the total transfer energy consumption of the workpiece; Einnate Indicates the inherent energy consumption of the workshop; Indicates process O IJ The operation on machine k is 1, otherwise 0; Indicates process O ij The operation at position l is 1, otherwise 0; Indicates process O ij The immediate predecessor process O of i(j-1) The operation on machine k is 1, otherwise 0; Indicates process O ij Belongs to the operation in MPBPO x The operation is 1, otherwise 0; Indicates process O ij Belongs to the operation in FPBPO x The operation is 1, otherwise 0; S ij Indicates process O ij The start operation time of process O; S IJ Indicates process O IJ The start operation time of process O; E IJ Indicates process O IJ The end operation time of process O; E ij Indicates process O ij The end operation time of process O; Y IJHG Indicates process O IJ Precedes O HG When operating, it is 1, otherwise 0; CEF carbon emission conversion coefficient.
[0108] The objective function of the generalized job shop scheduling model includes:
[0109] Minimize the makespan C max It is an embodiment of the production efficiency of the manufacturing system, and this objective function is expressed as:
[0110]
[0111] Minimizing the total tardiness TD will reduce the workshop delivery delay, thereby improving the workshop's quick response speed and reducing the loss of tardiness compensation. This objective function is defined as:
[0112]
[0113] The scheduling objective based on carbon emissions CE is an important measure for the workshop to achieve energy conservation, emission reduction, and green production. The carbon emissions CE are obtained by multiplying the total energy consumption E busy of the machine's operating state, the total energy consumption E idle of the no-load state, the total energy consumption E tran of the workpiece transfer, and the inherent energy consumption E innate of the workshop by the conversion coefficient CEF. The corresponding objective function is expressed as:
[0114] f3 = CE = min{CEF·(E busy +E idle +E tran +E innate )} (3)
[0115] E busy is the total energy consumption consumed by all machines for operating on the workpiece, expressed as:
[0116]
[0117] E idle is the total energy consumption of all machines in the standby state, i.e., no-load operation, expressed as:
[0118]
[0119] E tran is the total energy consumption for workpiece transfer. It mainly refers to the energy consumption of transferring workpieces by transfer carts, etc. when the adjacent processes of the workpiece need to be switched between different machines. Here, it is assumed that the transfer energy consumption of a single process of all workpieces is equal, and this goal is expressed as:
[0120]
[0121] E innate is the public energy consumption of the workshop, including the energy consumption generated by public facilities such as lighting, ventilation, and air conditioning to maintain the basic production environment of the workshop. This goal is expressed as:
[0122] E innate = P g C max (7).
[0123] The combined process is regarded as a process set composed of multiple conventional processes. When the combined process has only one element, it can be regarded as a conventional process. When establishing the model, they are collectively referred to as processes.
[0124] The constraint conditions of the generalized job shop scheduling model are:
[0125]
[0126] Equation (8) ensures that there will be no interruption in the operation process; Equation (9) ensures that the completion time of any workpiece will not exceed the maximum completion time required; Equation (10) ensures that the start time of the current operation is not lower than the completion time of its immediate predecessor operation(s); Equation (11) ensures that the start time of any operation does not exceed its end time; Equation (12) ensures that any operation is only performed once on the corresponding machine; Equation (13) ensures that the start time of any operation is not negative, and its operation time and end time are both greater than zero; Equation (14) ensures that only one operation is in the working state on a machine at the same time; Equation (15) ensures that the operations belonging to the same set of forced parallel batch processing operations must be batch processed simultaneously; Equation (16) ensures that the operations belonging to the same set of flexible parallel batch processing operations can be processed individually or randomly combined for batch processing.
[0127] Step S2, based on the generalized job shop scheduling constraints considering forced and flexible parallel batch processing, designs the encoding and decoding of operations and machines, and initializes the population using a hybrid heuristic strategy.
[0128] The multi-objective generalized job shop scheduling with identical-machine parallel batch processing considers operation sequencing and machine selection, selects operation encoding, and the encoding includes four segments of integer encoding sequences. The corresponding encoding example is as Figure 2 shown. The first segment is the operation encoding sequence, and each gene is represented by the workpiece number. For workpiece i, i ∈ [1, N], the occurrence frequency is j, j ∈ [1, N i , O ij represents the corresponding operation of the workpiece. For the combined operation with identical-machine operation coupling, it is described by a workpiece set, and the corresponding operation takes the operations corresponding to each workpiece in the workpiece set. For example, the workpieces 2 and 3 in the shaded box in the figure represent operations O 22 and O 33 to construct the combined operation.
[0129] The second segment is used to record the identical-machine parallel operation information. The encoding of ordinary operations is 0, and the encoding of parallel batch processing operations is a number greater than 0. For example, the 12th and 13th encodings are 1, indicating that the corresponding operations O 22 and O 33 are identical-machine parallel operations. The third segment is the machine encoding, and its selectable range is [1, M]. The machine encoding order corresponds to the operation encoding, indicating the operation machine specified by the corresponding (combined) operation. The fourth segment is the operation time encoding, corresponding to the previous operation encoding and machine encoding, indicating the operation time required for the corresponding operation to operate on the specified machine.
[0130] To obtain a high-quality solution and narrow the scope of the search space, an active decoding method is designed. The decoding needs to ensure that the completion time constraints of the machine's predecessor processes and the workpiece's predecessor processes are satisfied for each process. Considering multiple parallel batch processes involved in the problem to be solved makes the constraints considered in decoding more complex. The steps of the active decoding method are as follows:
[0131] S211: Set A O (1,1:ON) and A O (2,1:ON) store the start and completion times of each process respectively, where ON = |OPlan|; set MP = {MP[1], MP[2],..., MP[M]}, JP = {JP[1], JP[2],..., JP[J]} to record the end times of the immediate predecessor processes of each machine and workpiece respectively; initialize all elements in MP and JP to 0, and k = 1;
[0132] S212: Sequentially read the gene position elements of OPlan[k], where 1 ≤ k ≤ |OPlan|, determine its process O, and obtain the corresponding operation machine m and operation time Pt of O from MPlan and TmPlan respectively O , calculate
[0133] S213: Find the idle time interval [t_s, t_e] of machine m. If max(as O , t_s) + Pt O ≤ t_e, go to step S214; otherwise, go to step S215;
[0134] S214: S O = max(as O , t_s), E O = S O + Pt O ; if (S O = t_s) ∧ (E O = t_e), delete the interval [t_s, t_e]; if (S O = t_s) ∧ (E O < t_e), update [t_s, t_e] = [E O , t_e]; if (t_s < S O ) ∧ (E O = t_e), update [t_s, t_e] = [t_s, S O ; if (t_s < S O ) ∧ (E O < t_e), update [t_s, t_e] = [t_s, S O and insert a new [t_s', t_e'] = [E O , t_e],
[0135] S215: E O = S O + Pt O ; If as O > MP[m], add a new idle interval [t_s', t_e'] = [MP[m], as O on machine m; MP[m] = E O ,
[0136] S216: Update A O (1, k) = S O , A O (2, k) = E O ;
[0137] S217: k = k + 1. If k ≤ |OPlan|, go to step S212; otherwise, end.
[0138] To increase the number of non-dominated solutions in the initial population and at the same time take into account the diversity of solutions, the present invention designs a hybrid initialization method for multi-objective problems. First, randomly generate operation codes, and adopt various heuristic strategies and random selection rules according to the characteristics of the scheduling problem when allocating machines. The specific steps for initializing population individuals are as follows:
[0139] S221: Input the forced parallel batch processing operation set and the flexible parallel batch processing operation set, initialize the population size NP, define an array of NP × 3, and the counter np = 1;
[0140] S222: Generate a new empty operation sequence OPlan, coupling sequence PPlan, machine sequence MPlan, and job time sequence TPlan, and set all numbers in [1, N] to the unmarked state;
[0141] S223: If all numbers in [1, N] are marked, go to step S227; otherwise, randomly select an unmarked number i in [1, N], and judge the frequency j of the occurrence of i in OPlan. If j ≥ N i , mark i and go to step S223. If j < N i , obtain O from different workpiece process routes PPlan i , 1 ≤ i ≤ N. If O ij is a regular operation, add j to the end of OPlan and add 0 to the end of PPlan; if O ij is a forced parallel batch processing operation element, go to step S224; if O ij is a flexible parallel batch processing operation element, go to step S225; ij
[0142] S224: Obtain the corresponding workpiece set I x ; If there is a predecessor parallel batch processing operation in the operations of MPBPO x , go to step S223; otherwise, bind and place the workpiece numbers in I x at the end of OPlan, and at the same time place 1 corresponding to the quantity of workpiece set I x at the end of PPlan, and go to step S226;
[0143] S225: Select l quantity of operations in to form a flexible parallel batch processing operation set x and based on select a machine set ks that satisfies the equipment capacity constraint. If , go to step S225; otherwise, obtain the corresponding workpiece set I ; If there is a predecessor parallel batch processing operation in the operations of x , go to step S223; otherwise, bind and place the workpiece numbers in I at the end of OPlan, and at the same time place 1 corresponding to the quantity of workpiece set I x at the end of PPlan, and delete the workpieces selected to be placed at the end of OPlan in FPBPO x ; x
[0144] S226: Obtain the parallel batch processing operation set O corresponding to O ij and the set of predecessor operations JP[O] of O that have not been placed in the scheduling sequence. Judge whether there is a predecessor parallel batch processing operation in O. If there is no predecessor parallel batch processing operation for the workpiece corresponding to O, insert the workpiece numbers i corresponding to O ij ∈IP[O] randomly before the workpiece corresponding to O in OPlan; otherwise, insert it randomly between the closest predecessor parallel batch processing operation to i and the workpiece corresponding to O, and go to step S223;
[0145] S227: Generate a random number r between (0, 1). If r < 0.25, for the operation O in OPlan[i], 1 ≤ i ≤ |OPlan|, select as the operation machine according to active decoding; if 0.25 ≤ r < 0.5, select as the operation machine according to active decoding; if 0.5 ≤ r ≤ 0.75, select as the operation machine according to active decoding; if r > 0.75, then select from Mac ORandomly select m as the operating machine for process O to form the machine sequence MPlan; determine the operation time of each process, where the operation time of the parallel batch processing process is the maximum value of the operation times of each process in the parallel batch processing process on the machine, to form the operation time sequence TmPlan; store Pop(np,1) = OPlan, Pop(np,2) = MPlan, Pop(np,3) = TmPlan;
[0146] S228: If np < NP, np = np + 1, go to step S222; otherwise, end.
[0147] Step S3, the algorithm introduces a crossover and mutation operation that adapts to forced and flexible parallel batch processing, adjusts the crossover and mutation probability based on reinforcement learning, and obtains a new generation of population through crossover and mutation for iteration.
[0148] The crossover operation plays a crucial role in genetic operations. The global search ability of the genetic algorithm actually depends on the effectiveness of the crossover operator. Therefore, the crossover operator has a significant impact on the overall performance of the algorithm. Given the existence of parallel batch processing process constraints in the present invention, the number of parallel batch processing processes covered by each individual is different. Based on this, the conventional crossover method is not applicable to this problem. The present invention designs a one-by-one order crossover (OOOX), and the OOOX operation steps are as follows:
[0149] S311: Select parents P1 and P2, and initialize the offspring CP process sequence, machine sequence, and operation time sequence to be empty;
[0150] S312: Generate a random number r between (0, 1). If r < 0.5, select the first gene position element of the P1 process sequence; otherwise, select the first gene position element of the P2 process sequence;
[0151] S313: Add the selected element to the end position of the CP process sequence, and at the same time classify the corresponding machine and time of this element into the end of the CP machine sequence and the operation time sequence respectively; remove the gene where the element first appears in the P1 and P2 process sequences, and delete the machine and time genes corresponding to this element;
[0152] S314: Repeat steps S312 and S313 until the elements in parents P1 and P2 are empty, then the process crossover is completed.
[0153] The schematic diagram of OOOX crossover is as Figure 3As shown in the figure, the information of two parents P1 and P2 is first shown, then the crossover operation steps are shown, and finally the information of the generated offspring CP is shown. Specifically, when performing the crossover operation, in the first step, with the help of a random number r (r < 0.5), the first gene position element 5 of the P1 process sequence is selected, and the element 5 is added to the end of the offspring CP process sequence. At the same time, the corresponding machine and time of this element are respectively added to the end of the machine sequence and the operation time sequence, and the gene positions of the first occurrence of element 5 in the P1 and P2 process sequences and their corresponding machine and time gene positions are deleted.
[0154] Subsequently, when the randomly generated number r > 0.5 in the third time, the first gene position element 3 of P2 at this time is selected, and the element 3 is added to the end of the offspring CP process sequence. The corresponding machine and time of this element are placed at the end of the offspring CP machine sequence and the operation time sequence. At the same time, the gene positions of the first occurrence of element 3 in the P1 and P2 process sequences and their corresponding machine and time gene positions are deleted; continue to analogize in this way until all the elements in P1 and P2 are processed, at this time the crossover operation is completed and the offspring CP is generated. It can be clearly seen from the finally generated CP that the offspring meets the front and back process constraint conditions of the operation machine, operation time, and batch processing process, which fully shows that OOOX can obtain feasible offspring.
[0155] Mutation operation can increase the diversity of the population. In the later stage of the algorithm search, a good mutation mechanism can enhance the local search ability of the algorithm, so as to search for better solutions. The present invention proposes four mutation methods, such as Figure 5 , Figure 6 , Figure 7 and Figure 8 shown.
[0156] For process mutation, after the mutation operation, the machine sequence and the operation time sequence need to be updated so that the operation machine and operation time corresponding to each process are the same as before the mutation. Figure 4 shows that the elements at the 5th and 9th gene positions of the process sequence are swapped, and the elements at the 15th and 16th combined process gene positions and the 19th process gene position are swapped. Since the element 4 at the 19th gene position before the swap represents process O 44 , and the data 4 represents process O 43 after the swap, it is necessary to update the machine sequence and the operation time sequence to ensure that the operation machine and operation time can correspond to the process.
[0157] For machine mutation, Figure 5 shows that the operation machine 7 of the parallel batch processing process O 24, O 34 is replaced by machine 6. Therefore, the operation time needs to be updated to the parallel batch processing process O 24, O 34The operation time on machine 6. For the recombination mutation of flexible combined processes, as many flexible combined processes as possible are grouped together. Figure 6 shows that the 13th gene position is moved to the 9th gene position, making process O 22, O 33 recombined into a flexible combined process. Randomly select PBPO and try to move it forward as much as possible for mutation. Figure 7 shows that the 12th and 13th combined process gene positions are moved to the 6th gene position. Before moving, the most forward positions of all other processes of the combined process workpiece need to be found to ensure that the processes in the combined process do not exceed the processes before it.
[0158] Based on reinforcement learning, the crossover and mutation probabilities of the genetic algorithm are adjusted. The SARSA algorithm and the Q-learning algorithm are two typical value-based reinforcement learning algorithms, and their goal is to estimate the Q-value of the optimal policy. Q-learning updates the Q-value by selecting the action with the highest expected Q-value in each state and utilizes more accumulated knowledge, so it can obtain stronger final performance. However, in the early learning stage, the effect of Q-learning is low and the quality is poor. In contrast, SARSA has higher learning efficiency and better convergence, but it is prone to falling into local optima and difficult to obtain the global optimal policy.
[0159] To combine the advantages of the two algorithms, SARSA (State-Action-Reward-State-Action) and Q-learning are respectively applied to different stages of the algorithm. In the initial stage of the algorithm, the Q-value is updated based on SARSA and recorded in the Q-value table, and these Q-values will be used as the pre-training process of the Q-learning algorithm. If the conversion condition is met, that is, the number of iterations is greater than the product of the total number of states and the total number of actions, the SARSA algorithm will be converted to the Q-learning algorithm. This conversion condition ensures that the Q-learning algorithm has sufficient learning materials or learning experience, which is specifically measured by the number of non-zero Q-values, and the number of Q-values is closely related to the number of states and actions.
[0160] When using reinforcement learning to optimize the crossover and mutation rates in NSGA-III and represent the corresponding states, the distribution diversity DV and coverage SC of the solutions are usually used as key indicators because they can reflect the quality, coverage range, and balance of the solution set. A total of four states are set: S 1 , both the diversity and coverage increase; S 2 , the diversity increases while the coverage decreases; S 3 , the diversity decreases while the coverage increases; S 4 , both the diversity and coverage decrease.
[0161] For each iteration, the agent will take different actions to obtain appropriate crossover and mutation rates. The action set π contains 8 actions. Among them, the crossover probability is between 0.6 and 0.95, and the mutation probability is between 0.05 and 0.4, with an intermediate difference of 0.05. Action selection adopts the ε-greedy strategy, as shown in equation (17), where ε is called the greed rate, and r 0-1 is a random value from 0 to 1. When ε ≤ r 0-1 , select the action a corresponding to the maximum value in the Q-value table; otherwise, randomly select an action a.
[0162]
[0163] After the action is executed, evaluate the quality of the action based on the newly entered state. The effects generated by the action are positive or negative, corresponding to the corresponding positive and negative reward values. The reward value is set as shown in the following formula.
[0164]
[0165] Based on the above, regard the main part of the genetic algorithm as the environment, and dynamically adjust the crossover and mutation probabilities through the agent. The specific steps are as follows:
[0166] S321: Feed back the environmental information to the agent, and input the fitness value information F of all individuals in the population iter ;
[0167] S322: Judge the current population iteration times iter. If iter = 1, initialize the Q-value table and define the initial action a 1 , and go to step S327; if iter > 1, calculate the reward value r iter ;
[0168] S323: Judge the next state s based on the changes in diversity and coverage iter+1 ;
[0169] S324: Judge whether the conversion condition is satisfied. If it is satisfied, update the Q-value table according to the Q-learning formula and select the next action a through the greedy strategy iter+1 ; if not, update the Q-value table according to SARSA and select the next action a through the ε-greedy strategy iter+1 ;
[0170] S325: Update the current state value s iter ← s iter+1 , and update the current action a iter ← a iter+1 ;
[0171] S326: Update the fitness value information F of the previous generation iter-1 ←F iter ;
[0172] S327: The agent makes a decision and updates the crossover and mutation probabilities according to the current action a iter Update the crossover and mutation probabilities.
[0173] The present invention provides a knowledge-driven hybrid heuristic initialization strategy to ensure the quality of the initial population. Based on the constraints of the same-machine parallel (forced and flexible parallel batch processing) jobs, a new genetic operation method is designed to ensure the feasibility of the newly generated solutions. At the same time, a parameter adaptive adjustment is designed based on reinforcement learning to adapt to the dynamic adjustment of the crossover and mutation probabilities in the environment.
[0174] Taking the MK05 example in the BRdata dataset proposed by Brandim arte as the benchmark example below, the workshop production examples in the electronic product detection workshop are simulated. By selecting the processes of different workpieces, 2 sets of forced parallel batch processing processes and 2 sets of flexible parallel batch processing processes are constructed, so as to form the test example CMK05 applicable to the generalized job shop scheduling problem with parallel batch processing processes. The specific information of the parallel batch processing processes is shown in Table 1. In the electronic product detection workshop, the fixed power of the workshop is 35kW, and the power information of each machine is shown in Table 2. The carbon emission conversion coefficient CEF is set to 0.7559.
[0175] Table 1 Information of parallel batch processing processes
[0176]
[0177] Table 2 Power information of machines
[0178] Machine Operating power / kW Idle power / kW <![CDATA[M 1 > 20 3.45 <![CDATA[M 2 > 15 2.82 <![CDATA[M 3 > 6 0.84 <![CDATA[M 4 > 12 1.58
[0179] All the codes of the present invention are programmed and implemented on MatlabR2021a, and run on a computer configured with a 64-bit Windows 10 operating system, an Intel(R) Core(TM) i5-12600KF CPU@3.70GHz processor and 16G of machine-borne RAM. The population size PS and the maximum number of iterations IterMax of the algorithm are set to 200 and 100 respectively, and the initial crossover rate and mutation rate of KNSGA-III-RL are set to 0.8 and 0.2 respectively.
[0180] Figure 8 To solve the example using the multi-objective generalized job shop scheduling method, one of the obtained scheduling Gantt charts is shown. The forced and flexible same-machine job processes are marked with rectangular boxes in the figure. It shows that the multi-objective generalized job shop scheduling method proposed by the present invention can better solve the low-carbon flexible job shop scheduling problem with same-machine jobs.
[0181] The above description is a detailed description of the preferred and feasible embodiments of the present invention. However, the embodiments are not intended to limit the scope of the patent application of the present invention. Any equivalent changes or modifications made under the technical spirit disclosed by the present invention shall fall within the scope of the patent covered by the present invention.
Claims
1. A multi-objective generalized job shop scheduling method considering same-machine operations, characterized in that: The following steps are involved: S1: Establish a generalized job shop scheduling model, describe the generalized job shop scheduling problem considering mandatory and flexible parallel batch processing, and determine the objective function and constraints of the generalized job shop scheduling model; S2: Based on the generalized job shop scheduling constraints considering mandatory and flexible parallel batch processing, the encoding and decoding of processes and machines are designed, and a hybrid heuristic strategy is used to initialize the population; S3: The algorithm introduces crossover and mutation operations that adapt to mandatory and flexible parallel batch processing, adjusts the crossover and mutation probability based on reinforcement learning, and obtains a new generation of population for iteration through crossover and mutation; S4: Until the algorithm reaches the maximum number of iterations, the generalized job shop scheduling model outputs the optimization result, and the generalized job shop scheduling Gantt chart of mandatory and flexible parallel batch processing is obtained.
2. The multi-objective generalized job shop scheduling method considering same-machine operations according to claim 1 is characterized in that: A generalized job shop scheduling model is established, specifically: The generalized job shop scheduling problem considering mandatory and flexible parallel batch processing is described as follows: N workpieces are operated on M machines, each workpiece contains multiple processes, each process can be selected from multiple machines, and the operation time of each workpiece process depends on the selected machine; multiple jobs can be processed on the same machine at the same time, mandatory same-machine or flexible same-machine operation; under the constraints of machine availability, process sequence and flexible same-machine operation, determine the order of operations, machines used, and operation time of each process of all workpieces, obtain the optimization of maximum completion time, total delay and carbon emission scheduling indicators, and reduce conflicts between different objectives; According to the actual workshop production situation, the following assumptions need to be met: At the initial moment, all machines are in usable state, all workpieces can be processed immediately, and there is no priority difference between workpieces; for each process of the same workpiece, the operation sequence is determined according to the established process route, and at the same time, only one process of a workpiece can be in the process of processing; any machine can only process one process at the same time; the time spent on auxiliary operations is covered by the working time quota of the process; when the process starts on the selected machine, it only needs to consider the availability of the current machine and whether the previous process of the process has been completed; during the operation of the process, the operation flow of the entire process must not be interrupted.
3. The multi-objective generalized job shop scheduling method considering same-machine operations according to claim 1 is characterized in that: The generalized job shop scheduling model includes the following mathematical symbols: Index variables: i, h represent the workpiece number, i, h∈{1,2...,N}; j, g represent the process number, j, g∈{1,2...,N} i }; k represents the machine number, k∈{1,2...,M}; x represents the parallel batch process set number; N represents the total number of workpieces; M represents the total number of machines; N i represents the total number of processes of workpiece i; I and H represent the workpiece set serial numbers corresponding to the combined processes; J and G represent the process set serial numbers corresponding to the combined processes; U represents an infinite real number; N MPBPO Indicates the number of processes in the forced parallel batch process set; N FPBPO It represents the number of processes in the flexible parallel batch process set; Parameter variable: O ij represents the jth process of workpiece i; O IJ Represents the combined process, O IJ = {O ij ,O hg ,O pg ,...,O yz }; T i represents the delivery time of workpiece i; Indicates process O IJ The operation time on machine k; Indicates process O ij The operation time on machine k; P k represents the operating power of machine k; JS IJ Represents the combined process O IJ Corresponding artifact set; OS IJ Represents the combined process O IJ Corresponding process set: MPBPO x Represents the xth forced parallel batch process set; FPBPO x represents the xth flexible parallel batch process set; MS k Indicates the startup time of machine k; ME k represents the downtime of machine k; Mac ij Indicates O ij Optional machine set for Cap k represents the capacity of machine k; Indicates O ij The weight on machine k; represents the no-load power of machine k; E t Represents the transfer energy consumption of a single process of the workpiece; P g Indicates the fixed power of the workshop; Decision variable: C i represents the completion time of workpiece i; E busy Indicates the total energy consumption of the machine during operation; E idle Indicates the total energy consumption of the machine during idle time; E tran Represents the total energy consumption of workpiece transfer; E innate Indicates the inherent energy consumption of the workshop; Indicates process O IJ The job is 1 on machine k, otherwise it is 0; Indicates process O ij The operation at position l is 1, otherwise it is 0; Indicates process O ij The preceding process O i(j-1) The job is 1 on machine k, otherwise it is 0; Indicates process O ij Belongs to MPBPO x The operation in is 1, otherwise it is 0; Indicates process O ij FPBPO x The operation in is 1, otherwise it is 0; S ij Indicates process O ij The start time of the operation; S IJ Indicates process O IJ The start time of the operation; E IJ Indicates process O IJ End time of operation; E ij Indicates process O ij End time of operation; Y IJHG Indicates process O IJ Before O HG It is 1 when operating, otherwise it is 0; CEF carbon emission conversion factor.
4. The multi-objective generalized job shop scheduling method considering same-machine operations according to claim 3 is characterized in that: The objective function of the generalized job shop scheduling model includes: Minimize the maximum completion time C max , the objective function is expressed as: Minimize the total delay TD, the objective function is defined as: Carbon emissions CE is composed of the total energy consumption E of the machine in its operating state. busy , total energy consumption in no-load state E idle , Total energy consumption of workpiece transfer E tran And the inherent energy consumption of the workshop E innate Multiplying by the conversion coefficient CEF, the objective function is expressed as: f3=CE=min{CEF·(E busy +E idle +E tran +E innate )} (3) E busy The total energy consumed by all machines to operate the workpiece is expressed as: E idle It is the total energy consumption of all machines in standby state, i.e. no-load operation, expressed as: E tran is the total energy consumption of workpiece transfer, and the target is expressed as: E innate is the common energy consumption of the workshop, and the target is expressed as: E innate =P g C max (7)。 5. The multi-objective generalized job shop scheduling method considering same-machine operations according to claim 3 is characterized in that: The constraints of the generalized job shop scheduling model are: Formula (8) ensures that there will be no interruption in the process of operation; Formula (9) ensures that the completion time of any workpiece will not exceed the maximum completion time; Formula (10) ensures that the time when the current process starts the operation is not less than the time when its immediate predecessor process completes the operation; Formula (11) ensures that the time when any process starts the operation does not exceed the time when it ends the operation; Formula (12) ensures that any process only performs one operation on the corresponding machine; Formula (13) ensures that the start time of any process is not a negative value, and its operation time and end time are both greater than zero; Formula (14) ensures that only one process is in operation on a machine at the same time; Formula (15) ensures that the processes belonging to the same mandatory parallel batch process set must be batch processed at the same time; Formula (16) ensures that the processes belonging to the same flexible parallel batch process set can be processed individually or randomly combined for batch processing.
6. The multi-objective generalized job shop scheduling method considering same-machine operations according to claim 2 is characterized in that: Design the codes for machines and processes, specifically: Multi-objective generalized job shop scheduling with parallel batch processing on the same machine considers process sorting and machine selection, selects process coding, and the coding includes four segments of integer coding sequences; the first segment is the process coding sequence, each gene is represented by the workpiece number, workpiece i, i∈[1,N], occurrence frequency j, j∈[1,N i ],O ij Represents the process corresponding to the workpiece; the combined process coupled with the same machine operation is described by the workpiece set, and the corresponding process takes the process corresponding to each workpiece in the workpiece set; the second section records the information of parallel operations on the same machine, the ordinary process is coded as 0, and the parallel batch process is coded as a number greater than 0; the third section is the machine code, and its optional range is [1, M]. The machine code sequence corresponds to the process code, indicating the machine specified for the corresponding process; the fourth section is the operation time code, which corresponds to the previous process code and machine code, and indicates the operation time required for the corresponding process to operate on the specified machine.
7. The multi-objective generalized job shop scheduling method considering same-machine operations according to claim 2 is characterized in that: In step S2, the decoding of the design machine and process is specifically as follows: In order to obtain high-quality solutions and narrow the scope of the search space, an active decoding method is designed. The steps of the active decoding method are as follows: S211: Setting A O (1,1:ON), A O (2,1:ON) stores the start and completion time of each process respectively, ON = |OPlan|; set MP = {MP[1], MP[2], ..., MP[|M|]}, JP = {JP[1], JP[2], ..., JP[|J|]} to record the end time of the previous process of each machine and workpiece respectively; initialize all elements in MP and JP to 0, k = 1; S212: Sequentially read OPlan[k], 1≤k≤|OPlan| gene bit elements, determine its process O, and obtain the corresponding operation machine m and operation time Pt from MPlan and TmPlan respectively. O , calculate as O =max{JP[i]}, S213: Find the idle time interval [t_s, t_e] of machine m. If max(as O ,t_s)+Pt O ≤t_e, go to step S214, otherwise go to step S215; S214:S O =max(as O ,t_s),E O =S O +Pt O ; If (S O =t_s)∧(E O =t_e), delete the interval [t_s,t_e]; if (S O =t_s)∧(E O <t_e), update [t_s,t_e]=[E O ,t_e]; if (t_s<S O )∧(E O =t_e), update [t_s,t_e] = [t_s,S O ]; if (t_s<S O )∧(E O <t_e), update [t_s, t_e] = [t_s, S O ] and insert the new [t_s',t_e']=[E O ,t_e],JP[i]=E O , S215: E O =S O +Pt O ; if as O >MP[m], add a new idle interval [t_s',t_e'] on machine m = [MP[m],as O ];MP[m]=E O , JP[i]=E O , S216: Update A O (1,k)=S O ,A O (2,k)=E O ; S217: k=k+1, if k≤|OPlan|, go to step S212, otherwise end.
8. The multi-objective generalized job shop scheduling method considering same-machine operations according to claim 7 is characterized in that: In step S2, a hybrid heuristic strategy is used to initialize the population, specifically: A hybrid initialization method for multi-objective problems is designed. First, the process code is randomly generated. When allocating machines, a variety of heuristic strategies and random selection rules are adopted. The steps for initializing the population individuals are as follows: S221: input the forced parallel batch process set and the flexible parallel batch process set, initialize the population size NP, define the array NP×3, and the counter np=1; S222: Generate a new empty process sequence OPlan, coupling sequence PPlan, machine sequence MPlan and operation time sequence TPlan, and set all numbers in [1, N] to unmarked state; S223: If [1, N] are all marked, go to step S227; Conversely, randomly select an unmarked number i in [1, N] and determine the frequency j of i appearing in OPlan. If j ≥ N i , mark i and go to step S223, if j<N i , from different workpiece process routes PPlan i ,1≤i≤N to obtain O ij , if O ij For the conventional process, add j to the end of OPlan and add 0 to the end of PPlan; if O ij To force parallel batch process elements, go to step S224. ij For a flexible parallel batch process element, go to step S225; S224: Get O ij ∈MPBPO x , Corresponding artifact set I x ; If MPBPO x If there is a predecessor parallel batch process in the process, go to step S223, otherwise x Each workpiece number is bound and placed at the end of OPlan, and the corresponding workpiece set is placed at the end of PPlan x Quantity 1, go to step S226; S225: In FPBPO x , Select l x Quantity process composition flexible parallel batch process set Based on Select a machine set ks that satisfies the device capacity constraint if Go to step S225, otherwise obtain Corresponding artifact set I x ;like If there is a predecessor parallel batch process in the process, go to step S223. Otherwise, x Each workpiece number is bound and placed at the end of OPlan, and the corresponding workpiece set is placed at the end of PPlan x Quantity 1 and in FPBPO x The artifacts selected to be placed at the end of OPlan will be deleted; S226: Get O ij The corresponding parallel batch process set O and the predecessor process set JP[O] of O that is not put into the scheduling sequence are judged whether there is a predecessor parallel batch process in O. If there is no predecessor parallel batch process corresponding to O, O is placed in turn. ij ∈IP[O] randomly inserts the corresponding workpiece number i before the workpiece corresponding to O in OPlan; otherwise, randomly inserts it between the predecessor parallel batch process closest to i and the workpiece corresponding to O, and then goes to step S223; S227: Generate a random number r between (0,1). If r<0.25, for process O in OPlan[i], 1≤i≤|OPlan|, select according to active decoding k1,k2,...,k i ∈Mac O As a working machine; If 0.25≤r<0.5; select according to active decoding k1,k2,...,k i ∈Mac O As a working machine; If 0.5≤r≤0.75; select according to active decoding k1,k2,...,k i ∈Mac O As a working machine; if r>0.75, then from Mac O Randomly select m as the operating machine of process O to form a machine sequence MPlan; determine the operating time of each process, where the operating time of the parallel batch process is the maximum operating time of each process in the parallel batch process on the machine, forming an operating time sequence TmPlan; store Pop(np,1)=OPlan, Pop(np,2)=MPlan, Pop(np,3)=TmPlan; S228: If np<NP, np=np+1, go to step S222; otherwise end.
9. The multi-objective generalized job shop scheduling method considering same-machine operations according to claim 1, characterized in that: The algorithm introduces crossover mutation operations that adapt to mandatory and flexible parallel batch processing, specifically: Design a sequential crossover. The steps for sequential crossover are as follows: S311: Select the parent generations P1 and P2, and initialize the child generation CP process sequence, machine sequence, and operation time sequence to empty; S312: Generate a random number r between (0,1). If r<0.5, select the first gene position element of the P1 process sequence; otherwise, select the first gene position element of the P2 process sequence. S313: Add the selected element to the end of the CP process sequence, and place the machine and time corresponding to the element at the end of the CP machine sequence and the operation time sequence respectively; remove the gene where the element first appears in the P1 and P2 process sequences, and delete the machine and time genes corresponding to the element; S314: Repeat step S312 and step S313 until the elements in the parent generations P1 and P2 are empty and the process crossover is completed.
10. The multi-objective generalized job shop scheduling method considering same-machine operations according to claim 1, characterized in that: Adjust the crossover mutation probability based on reinforcement learning, specifically: Set the conversion conditions, combine the advantages of SARSA algorithm and Q-learning algorithm, and apply SARSA and Q-learning to different stages of the algorithm respectively; In the initial stage of the algorithm, the Q value is updated based on SARSA and recorded in the Q value table. These Q values will be used as the pre-training process of the Q-learning algorithm. If the conversion condition is met and the number of iterations is greater than the product of the total number of states and the total number of actions, the SARSA algorithm will be converted to the Q-learning algorithm. When using reinforcement learning to optimize the crossover and mutation rates and represent the corresponding states, the solution distribution diversity DV and coverage SC are used as key indicators to reflect the quality, coverage and balance of the solution set; four states are set: S1, both diversity and coverage increase; S2, diversity increases and coverage decreases; S3, diversity decreases and coverage increases; S4, both diversity and coverage decrease; For each iteration, the agent takes different actions to obtain the appropriate crossover rate and mutation rate. The action set π contains 8 actions, among which the crossover probability is between 0.6-0.95 and the mutation probability is between 0.05-0.4, and the intermediate difference is 0.
05. The action selection adopts the ε-greedy strategy, as shown in formula (17), where ε is the greedy rate and r 0-1 is a random value between 0 and 1; when ε≤r 0-1 When , select the action a corresponding to the maximum value in the Q value table, otherwise randomly select an action a; After the action is executed, the quality of the action is evaluated according to the latest state. The effect of the action is positive or negative, corresponding to the corresponding positive and negative rewards. The reward value is set to; The genetic algorithm is considered as the environment, and the crossover probability is dynamically adjusted through the agent. The specific steps are as follows: S321: Feedback the environmental information to the agent, and input the fitness value information F of all individuals in the population iter ; S322: Determine the number of iterations of the current population iter. If iter=1, initialize the Q value table, define the initial action a1, and go to step S327; if iter>1, calculate the reward value r iter ; S323: Determine the next state based on the diversity and coverage changes iter+1 ; S324: Determine whether the transition condition is met. If so, update the Q value table according to the Q-learning formula and select the next action a through the greedy strategy iter+1 If not satisfied, update the Q value table according to SARSA and select the next action a through the ε-greedy strategy iter+1 ; S325: Update current state value s iter ←s iter+1 , update the current action a iter ←a iter+1 ; S326: Update the previous generation fitness value information F iter-1 ←F iter ; S327: The agent makes a decision based on the current action a iter Update the crossover mutation probability.
Citation Information
Cited By
Flexible same-machine parallel batch optimization scheduling method with multi-cavity vulcanization constraint in rubber tire production workshop
CN121303473A