Optimization method for distributed heterogeneous hybrid workshop scheduling problem
By establishing a hybrid integer linear programming model and a super-heuristic algorithm for Q learning, the distributed heterogeneous hybrid workshop scheduling problem is optimized, and the problem of difficult to achieve energy saving and multi-objective optimization in the existing technology is solved, and the effect of reducing electricity prices and total completion time is achieved.
Patent Information
- Application Number
- CN202510201939.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-05-30
Smart Images

Figure CN120065944A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of distributed hybrid flow shop scheduling, and particularly relates to an optimization method for a distributed heterogeneous hybrid shop scheduling problem. Background Art
[0002] Scheduling plays a core role in manufacturing systems. By reasonably allocating resources, it realizes an effective balance of multiple conflicting goals and meets the needs of different decision-makers within a specified time. As an advanced manufacturing system, the Distributed Heterogeneous Hybrid Flow Shop Schedule Problem (DHHFSP) has received increasing attention. Existing technologies for solving the multi-objective DHHFSP include: using diverse neighborhood search strategies and iterative greedy methods to solve the bi-objective problem with makespan and total energy consumption, achieving good performance; using the shuffled frog leaping algorithm to solve the bi-objective problem with makespan and total tardiness objectives; considering heterogeneous factories with different numbers of parallel machines and using a dual-population cooperative memetic algorithm, considering collaborative iterative initialization, and using a Pareto-based hybrid iterative greedy algorithm to improve production efficiency while helping to reduce energy consumption and improve the sustainability of the production system.
[0003] However, the existing technologies have not yet solved the energy-saving DHHFSP for the following reasons: 1. Since the workshops are involved in different regions and policies, they are geographically dispersed, and the situations of different workshops are also different. The equipment is aging, and the environments where the equipment is located are very different. The processing methods and processes in different workshops may change, and the differences between workshops increase the difficulty of the problem. 2. In the production of manufacturing enterprises, often not only the processing time is considered as a single goal. Multiple goals such as energy consumption, labor cost, and equipment wear are also considered simultaneously. Multi-objective problems are widespread. Manufacturing enterprises need to consider both user needs and production costs, which undoubtedly makes the problem more complex. There is relatively little research on optimizing the multi-objective DHHFSP at present. 3. For distributed heterogeneous problems, the solution space of problems with large amounts of data is much larger than that of single-objective problems, and it is difficult to find the global optimal point of the problem. 4. There are various random dynamic events in the actual production scheduling of enterprises, such as production equipment failures, shortages of energy and materials, changes in processing time, insertions of emergency jobs, changes in job due dates, and interferences from uncertain factors such as logistics and warehousing. Dynamic adjustments need to be made to dynamic events. The adjustment of dynamic events not only needs to consider the scheduling scheme of the workshop but also needs to be optimized from a global perspective, especially the flexibility of the distributed manufacturing system needs to be improved. Therefore, finding an efficient and energy-saving method to solve DHHFSP has become an urgent problem to be solved at present. Summary of the Invention
[0004] To solve the above problems, the present invention provides an optimization method for the distributed heterogeneous hybrid shop scheduling problem.
[0005] To implement the above technology, the specific steps are as follows:
[0006] S1. Establish a mixed-integer linear programming (MILP) model for the distributed heterogeneous hybrid shop scheduling problem (DHHFSP) and set constraints for the model;
[0007] The mixed-integer linear programming model includes two objective functions, namely: f with the goal of minimizing the total completion time 1 and f with the goal of minimizing the total electricity price 2 , that is:
[0008]
[0009] In the formula, C max represents the total completion time; TEC represents the total electricity price;
[0010] Among them, the expressions of f 1 and f 2 are as follows:
[0011] f 1 = C max = max{Ct j,s}, 1 ≤ j ≤ n, 1 ≤ s ≤ S (2)
[0012]
[0013] In the formula, Ct j,s represents the completion time of workpiece j in stage s; n represents the number of workpieces; S represents the number of stages; S represents the stage set, S = (1, 2,..., S); F represents the number of factories; f represents the factory index, f = {1, 2,..., F}; m represents the number of machines; i represents the machine index, i = {1, 2,..., m}; T represents the set of time periods; t represents the time period index, t = {1, 2,..., T}; EC f,i represents the total energy consumption generated on machine i in factory f; TEC run represents the electricity price during the operation of the machine; TEC idle represents the electricity price during the idle process of the machine; TEC off / on represents the electricity price during the startup and shutdown process of the machine; h f (t) represents the correspondence between the time-of-use electricity price and the time interval, and the expression is as follows:
[0014]
[0015] In the formula,φ f Indicates the number of time-of-use electricity price stages for factory f; Indicates a specific time point; Indicates at the time point To the time point The corresponding electricity price within;
[0016] The electricity price TEC during the machine operation run The expression is as follows:
[0017]
[0018] In the formula, Indicates the energy consumption during the machine operation in factory f;
[0019] The electricity price TEC during the machine idle process idle The expression is as follows:
[0020]
[0021] In the formula, Indicates the energy consumption during the machine idle process in factory f;
[0022] The electricity price TEC during the machine startup and shutdown process off / on The expression is as follows:
[0023]
[0024] In the formula, Indicates the energy consumption during the machine startup and shutdown process in factory f;
[0025] Furthermore, the constraints on the mixed-integer linear programming model include:
[0026] Each workpiece can only be assigned to one factory, and can only be assigned to one factory. The expression is as follows:
[0027]
[0028] In the formula, J represents the set of workpieces, J = (1, 2,..., n); x j,f Indicates that if workpiece j is processed in factory f, x j,f = 1, otherwise x j,f = 0;
[0029] Each factory should contain at least one workpiece. The expression is as follows:
[0030]
[0031] In the formula, F represents the set of factories, F = (1, 2,..., F);
[0032] Each workpiece must go through all processes, and each process can only be placed on one machine for processing. The expression is as follows:
[0033]
[0034] Wherein, represents the number of machines in stage s at factory f; S represents the set of stages, S = (1, 2,..., S); y i,j,s,f represents that if workpiece j is processed on machine i in stage s at factory f, y i,j,s,f = 1, otherwise y i,j,s,f = 0;
[0035] The start time of the workpiece is greater than or equal to 0. The expression is as follows:
[0036]
[0037] Wherein, St i,j,s,f represents the start time of workpiece j on machine i in stage s at factory f; M represents the set of machines, M = (1, 2,..., m);
[0038] The completion time of a workpiece is greater than the start time, and the start time of the next workpiece is greater than the completion time of the previous workpiece. The expression is as follows:
[0039]
[0040] Wherein, C i,j,s,f represents the completion time of workpiece j on machine i in stage s at factory f;
[0041] In the same stage, a workpiece cannot be processed by multiple machines simultaneously, and one machine can only process one workpiece at a time. The expression is as follows:
[0042]
[0043] Wherein, Ct j,s represents the completion time of workpiece j in stage s; t j,s represents the processing time of workpiece j in stage s; Ct j-1,s represents the completion time of workpiece j - 1 in stage s; t j-1,s represents the processing time of workpiece j - 1 in stage s; Ct j,s+1 represents the completion time of workpiece j in stage s + 1; t j,s+1 represents the processing time of workpiece j in stage s + 1; Ct j′,s represents the completion time of workpiece j' in stage s; t j′,sDenote the processing time of workpiece \(j'\) at stage \(s\); \(L\) represents a positive number approaching infinity; \(z\) j,j′,s,f Denote that if workpiece \(j\) is processed before \(j'\) at stage \(s\) in factory \(f\), \(z\) j,j′,s,f \( = 1\), otherwise \(z\) j,j′,s,f \( = 0\); \(z\) j',j,s,t Denote that if workpiece \(j'\) is processed before \(j\) at stage \(s\) in factory \(f\), \(z\) j',j,s,t \( = 1\), otherwise \(z\) j',j,s,t \( = 0\); \(y\) i,j',s,t Denote that if workpiece \(j'\) is processed on machine \(i\) at stage \(s\) in factory \(f\), \(y\) i,j',s,t \( = 1\), otherwise \(y\) i,j,s,f \( = 0\); \(j\neq j'\);
[0044] Then, the first, second, and third decision variables are determined by Equations (4) - (13), and the expressions are as follows:
[0045]
[0046] S2. Use the hyper - heuristic algorithm of Q - learning to solve the optimization scheduling problem of the mixed - integer linear programming (MILP) model for the distributed heterogeneous hybrid - shop scheduling problem (DHHFSP). The steps are as follows:
[0047] S2.1. Establish the Q - learning model, and the expression is as follows:
[0048]
[0049] In the formula, \(s\) t Denotes the current state; \(a\) t Denotes the current action; \(s'\) t+1 Denotes the state after executing action \(a\) t ; \(A\) represents the set of all actions; \(a\) represents an action, \(a\in A\); \(r\) t+1 Denotes the instantaneous reinforcement signal; \(\lambda\in[0,1]\) represents the dynamic learning rate, that is, the update amplitude for each update; \(\gamma\in[0,1]\) represents the discount factor; \(Q\) t \((s\) t , a\) t ) represents the Q - value;
[0050] The core of Q - learning is the Q - value (Q Value), which represents the expected return of taking a certain action in a specific state; the Q - value is stored and updated through the Q - table (Q - Table), and the Q - table uses the Q - learning model to calculate the cumulative reward value;
[0051] S2.2. Set the key parameters of the Q - learning model and set the termination condition;
[0052] Among them, the key parameters include: population size ps, initial learning rate λ, discount factor γ, initial greed rate ε; the termination condition is the running time of 500 seconds;
[0053] The Q-learning model is set as follows: The five key parameters are set at five levels, and experimental data are obtained through different groups for analysis of better parameter groups; different values of the four key parameters are combined, and the sensitivity of the parameters is analyzed by using the full-factor experimental design method through the analysis of variance (ANOVA) technique. Each parameter group is independently repeated 10 times, and different group settings and results are shown in Table 1;
[0054] Table 1: Different group settings and results
[0055]
[0056] The best values in the table are marked in bold;
[0057] The best parameter combination obtained is: ps = 40, λ = 0.3, γ = 0.7, ε = 0.4;
[0058] S2.3. Initialize the population and Q-table by defining five low-level individuals (LLH);
[0059] The five low-level heuristic (LLH) rules include:
[0060] (1) Insertion of critical workpiece in front: Randomly select a critical workpiece from the critical workpieces in the workpiece set of a critical factory in the factory set, and insert this workpiece in front of other workpieces. If either the minimized completion time (f 1 ) or the total electricity price (f 2 ) or both are optimized, the insertion operation is completed; otherwise, continue to select the next position;
[0061] (2) Insertion of critical workpiece at the back: Randomly select a critical workpiece from the critical workpieces in the workpiece set of a critical factory in the factory set, and insert this workpiece behind other workpieces. If either the minimized completion time (f 1 ) or the total electricity price (f 2 ) or both are optimized, the insertion operation is completed; otherwise, continue to select the next position;
[0062] (3) Insertion of non-critical workpiece in front: Randomly select a workpiece from the workpieces in the workpiece set of a non-critical factory in the factory set, and insert this workpiece in front of the workpieces of other non-critical factories. If either the minimized earliest completion time (f 1 ) or the total electricity price (f 2 ) or both are optimized, the insertion operation is completed; otherwise, continue to select the next position;
[0063] (4) Post-insertion of non-critical workpieces: Randomly select a workpiece from the non-critical workpieces in the factory set, and insert this workpiece after the workpieces in other non-critical factories. If one or both of the minimization of the earliest completion time (f 1 ) and the total electricity price (f 2 ) are optimized, the insertion operation is completed; otherwise, continue to select the next position;
[0064] (5) Exchange of critical workpieces: Randomly select a critical workpiece from the critical path, and then exchange this workpiece with other workpieces. If one or both of the minimization of the earliest completion time (f 1 ) and the total electricity price (f 2 ) are optimized, the exchange operation is completed; otherwise, continue to select the next position;
[0065] The critical path is: In the hybrid flow shop scheduling problem, finding the critical path is an important step in optimizing the scheduling; the critical path is the path that determines the minimization of the earliest completion time in the scheduling scheme, and it represents the longest time path from the start to the completion of the task; the processing delay of each task on the critical path will directly lead to the delay of the completion time;
[0066] The critical factory is: The factory with the largest overall completion time; the critical factory is the factory that plays a leading role in the total scheduling time, and its optimization has the greatest impact on the overall performance;
[0067] The reasons for selecting the above LLH are as follows:
[0068] There are pre-insertion and post-insertion operations for critical factories and non-critical factories. The operations of critical factories have a significant impact on changing the minimization of the earliest completion time (f 1 ), while the operations of non-critical factories have a significant impact on changing the total electricity price (f 2 ). However, if only pre-insertion and post-insertion operations are performed, the search space is not large enough. Therefore, the exchange of workpieces is added to expand the search space, which is beneficial to jumping out of the local optimum;
[0069] Define the current state s t as the fitness expression of multiple objectives in the distributed heterogeneous hybrid shop scheduling problem (DHHFSP) as follows:
[0070]
[0071] In the formula, z′ min and z′ max represent the minimum and maximum objective values; o represents the number of objectives. In the present invention, there are two objectives (f 1 and f 2) Then \(o = 2\); since the two objectives have different orders of magnitude, the objective values need to be normalized to \([0, 1]\). When \(q = 1\), \(f\) q \((\pi)\) represents the earliest completion time for minimizing the first objective of the current individual, \(z'\) min,q is the value of the earliest completion time for minimizing the smallest in the population, \(z'\) max,q represents the value of the earliest completion time for minimizing the largest in the population; when \(q = 2\), \(f\) q \((\pi)\) represents the total electricity price of the second objective of the current individual, \(z'\) min,q is the value of the total electricity price that is the smallest in the population, \(z'\) max,q represents the value of the largest TEC in the population;
[0072] Define the selection of LLH as an action. During the Q-table update process, there are 5 choices from the current action to the next action. Therefore, a \(5\times5\) probability matrix is used as the Q-table of the algorithm;
[0073]
[0074] In the formula: gen represents the number of generations of algorithm operation. When initializing, gen = 1; the element \(p\) xy \((gen)(x = 1,\cdots,5,y = 1\cdots5)\) represents the probability that the next second action of the first action \(x\) is \(y\) (here the action is defined as the operation performed after selecting LLH), \(p\) ij (1) all take the value of \(1 / 5\); the Q-table update formula is shown in Equation (14), where \(n\) is the number of jobs, \(Q\) t (s t ,a t ) refers to the Q-value of taking action \(a\) t at state \(s\) t , \(\max Q(s\) t+1 ,a) represents the maximum Q-value at state \(s\) t+1 when all actions are executed; \(\lambda\in[0,1]\) is the dynamic learning rate;
[0075] At the initial iteration, the Q-values corresponding to all state-actions are set to zero, so as to randomly select the initial state and action;
[0076] The initial population number is equal to 40 set in Table 1. 40 populations need to be generated. A hybrid initialization strategy is used to generate the population, that is, randomly generate 70% (28) individuals, and then use the NEH algorithm to generate 30% (12) individuals;
[0077] The NEH algorithm decodes while encoding. Add jobs to the sequence one by one. The added job tries all positions in the existing sequence and records the earliest completion time for minimizing the partial sequence (\(f\) 1 ) total electricity price (\(f\)2 ), calculate the fitness of all partial sequences using Equation (15); after all positions have been tried for insertion, find the position θ with the minimum fitness, and insert the workpiece into position θ until all workpieces have been inserted and the generated individuals are added to the population;
[0078] After generating 40 populations, perform the initialization operation on the initial population, as shown in Algorithm 1;
[0079]
[0080]
[0081] S2.4. Select 5 high-level populations (HLI) from the population generated by low-level individuals (LLH) to update the Q-table, and select the highest Q value from the updated Q-table using the greedy strategy;
[0082] The selection method is as follows: select the top 5 high-level individual populations (HLI) with the best performance by calculating the contribution rate of the population generated by low-level individuals (LLH); where the contribution rate is the average fitness of all individuals in the population (calculate the individual fitness using Equation (15));
[0083] The method for updating the Q-table is as follows: select the high-level individuals with the best performance through the contribution rate, and update the Q-table using Equation (14);
[0084] The iterative formula for the greedy rate ε is as follows:
[0085]
[0086] In the formula, ε initial is the initial value; ε final is the final value; gen is the number of generations of the algorithm running, and gen l is the current generation of running;
[0087] The initial greedy rate ε is set to 0.4;
[0088] The greedy strategy formula is:
[0089]
[0090] In the formula, ε t is the iterative greedy rate, and the iterative formula is Equation (17);
[0091] Through the continuous update of the greedy rate, the probability of Q-learning to select the best action also changes slowly. In the early stage of the algorithm operation, the algorithm obtains more search space, so the value of ε is very large, and there will be more randomly selected actions in Equation (18). Similarly, in the later stage of the algorithm, the algorithm selects some more valuable actions to find a smaller total completion time and total electricity price, so the value of ε slowly becomes smaller, making the action with the largest Q value in the Q table selected in Equation (18) to operate on the population;
[0092] Check whether the number of iterations has reached the maximum number of iterations. If not satisfied, go to S2.3; if satisfied, end;
[0093] The number of iterations is set to the maximum number of iterations of 100. If satisfied, output the Pareto front; otherwise, go to S2.3 and iterate repeatedly until the maximum number of iterations is satisfied;
[0094] S2.5. After meeting the termination condition, execute the right-shift strategy to reduce the total electricity price;
[0095] The specific right-shift operation is as follows:
[0096] While keeping the maximum completion time unchanged, transfer the production in the peak electricity price interval to the relatively low electricity price interval through the right-shift operation to save the electricity cost; limited by the characteristics of the DHHFSP problem, the right-shift operation needs to be adjusted sequentially from the last stage forward; perform the right-shift operation sequentially from the first factory to the last factory; sequentially search for the machine idle time from the last stage to the first stage in the first factory, shift the workpiece during the idle time period, and calculate the energy consumption of the current workpiece at the current stage, compare the original energy consumption with the energy consumption after the right-shift; if the original energy consumption is greater than the energy consumption after the right-shift, it is judged that the workpiece right-shift will be in the low electricity price interval; then perform the right-shift operation, otherwise, do not perform the right-shift operation;
[0097] The right-shift strategy is shown in Algorithm 2;
[0098]
[0099]
[0100] S3. Verify the accuracy of the model;
[0101] Specifically, verify the effectiveness of the comparison verification method through experiments with four existing algorithms for solving DHHFSP; the four algorithms are: Non-dominated Sorting Genetic Algorithm II (NSGA-II), Multi-Objective Evolutionary Algorithm Based on Decomposition (MOEA / D), Multi-Dimensional Differential Evolution (MDDE), and Knowledge-Based Cooperative Algorithm (KCA) for comparison.
[0102] Advantages of the present invention:
[0103] Compared with other comparison algorithms, the present invention can find a solution that effectively optimizes both minimizing the total completion time and minimizing the total electricity price, demonstrating its superiority in the distributed hybrid flow shop scheduling problem based on time-of-use electricity price.
[0104] Based on the characteristics of production scheduling, the present invention uses rule-based scheduling as the core. Under the time-of-use electricity price strategy, by encoding the resource allocation of processing factories, the equipment selection of workpieces, and the processing order, the complex DHHFSP is simplified to the exchange of individual sequences in the population, greatly reducing the solution complexity;
[0105] The present invention optimizes the scheduling scheme; the scheduling model is based on the total completion time of workpieces on each device, and at the same time combines a right-shift energy-saving strategy to ensure that the electricity price is reduced without affecting production efficiency, thus providing effective guidance for actual production. Through experiments, it is verified that the proposed hyper-heuristic algorithm based on reinforcement learning has the advantages of high solution efficiency and good solution results compared with other advanced algorithms in solving the DHHFSP with the goal of minimizing the total completion time and minimizing the total electricity price. Description of the drawings
[0106] Figure 1 It is a flow chart of the present invention;
[0107] Figure 2 It is a schematic diagram of the critical path of the present invention;
[0108] Figure 3 It is a schematic diagram of the right-shift operation of the present invention; among them, part a is a schematic diagram of the right-shift operation of factory 1; part b is a schematic diagram of the right-shift operation of factory 1;
[0109] Figure 4 It is a time-of-use electricity price scheme corresponding to formula (19) of the embodiment of the present invention; among them, part a is a schematic diagram of stage 1 and stage 2 of the machine in factory 1; part b is a schematic diagram of energy consumption and time-of-use electricity price;
[0110] Figure 5 It is a time-of-use electricity price scheme corresponding to formula (20) of the embodiment of the present invention; among them, part a is a schematic diagram of stage 1 and stage 2 of the machine in factory 1; part b is a schematic diagram of energy consumption and time-of-use electricity price;
[0111] Figure 6 It is a comparison chart of the pareto front results of the time-of-use electricity price scheme corresponding to formula (19) of the embodiment of the present invention with the fast elitist multi-objective genetic algorithm (NSGA-II), the multi-objective evolutionary algorithm based on decomposition (MOEA / D), the differential evolution algorithm (MDDE), and the knowledge-based collaborative algorithm (KCA);
[0112] Figure 7 This is a comparison graph of the pareto frontiers of the time-of-use electricity price plan corresponding to formula (20) of the embodiments of the present invention with respect to the fast elitist multi-objective genetic algorithm (NSGA-II), the multi-objective evolutionary algorithm based on decomposition (MOEA / D), the differential evolution algorithm (MDDE), and the knowledge-based collaborative algorithm (KCA). Specific embodiments
[0113] The present invention will be further described in detail below in conjunction with specific embodiments.
[0114] As Figure 1 shown, an optimization method for a distributed heterogeneous hybrid shop scheduling problem includes the following steps:
[0115] S1. Establish a mixed-integer linear programming (MILP) model for the distributed heterogeneous hybrid shop scheduling problem (DHHFSP) and set constraints for the model;
[0116] The mixed-integer linear programming model includes two objective functions, namely: f 1 aiming to minimize the total completion time and f 2 aiming to minimize the total electricity price, that is:
[0117]
[0118] In the formula, C max represents the total completion time; TEC represents the total electricity price;
[0119] Among them, the expressions of f 1 and f 2 are as follows:
[0120] f 1 = C max = max{Ct j,s}, 1 ≤ j ≤ n, 1 ≤ s ≤ S (2)
[0121]
[0122] In the formula, Ct j,s represents the completion time of workpiece j in stage s; n represents the number of workpieces; S represents the number of stages; S represents the stage set, S = (1, 2,..., S); F represents the number of factories; f represents the factory index, f = {1, 2,..., F}; m represents the number of machines; i represents the machine index, i = {1, 2,..., m}; T represents the set of time periods; t represents the time period index, t = {1, 2,..., T}; EC f,i represents the total energy consumption generated on machine i in factory f; TEC run represents the electricity price during the operation of the machine; TEC idleRepresents the electricity price during the machine idle process; TEC off / on Represents the electricity price during the machine startup and shutdown process; h f (t) represents the correspondence between the time-of-use electricity price and the time interval, and the expression is as follows:
[0123]
[0124] In the formula, φ f Represents the number of time-of-use electricity price stages for factory f; Represents a specific time point; Represents at the time point To the time point The corresponding electricity price within;
[0125] The electricity price TEC during the machine operation process run The expression is as follows:
[0126]
[0127] In the formula, Represents the energy consumption during the machine operation process in factory f;
[0128] The electricity price TEC during the machine idle process idle The expression is as follows:
[0129]
[0130] In the formula, Represents the energy consumption during the machine idle process in factory f;
[0131] The electricity price TEC during the machine startup and shutdown process off / on The expression is as follows:
[0132]
[0133] In the formula, Represents the energy consumption during the machine startup and shutdown process in factory f;
[0134] Furthermore, the constraints on the mixed-integer linear programming model include:
[0135] Each workpiece can only be assigned to one factory, and can only be assigned to one factory, and the expression is as follows:
[0136]
[0137] In the formula, J represents the workpiece set, J = (1, 2,..., n); x j,f Represents if workpiece j is processed in factory f, x j,f = 1, otherwise xj,f = 0;
[0138] Each factory contains at least one workpiece, and the expression is as follows:
[0139]
[0140] In the formula, F represents the set of factories, F = (1, 2,..., F);
[0141] Each workpiece must go through all processes, and each process can only be placed on one machine for processing. The expression is as follows:
[0142]
[0143] In the formula, represents the number of machines in stage s on factory f; S represents the set of stages, S = (1, 2,..., S); y i,j,s,f represents that if workpiece j is processed on machine i in stage s on factory f, y i,j,s,f = 1, otherwise y i,j,s,f = 0;
[0144] The start time of the workpiece is greater than or equal to 0. The expression is as follows:
[0145]
[0146] In the formula, St i,j,s,f represents the start time of workpiece j on machine i in stage s on factory f; M represents the set of machines, M = (1, 2,..., m);
[0147] The completion time of a workpiece is greater than the start time, and the start time of the next workpiece is greater than the completion time of the previous workpiece. The expression is as follows:
[0148]
[0149] In the formula, C i,j,s,f represents the completion time of workpiece j on machine i in stage s on factory f;
[0150] In the same stage, a workpiece cannot be processed by multiple machines simultaneously, and one machine can only process one workpiece at a time. The expression is as follows:
[0151]
[0152] In the formula, Ct j,s represents the completion time of workpiece j in stage s; t j,s represents the processing time of workpiece j in stage s; Ct j-1,s represents the completion time of workpiece j - 1 in stage s; tj-1,s Denotes the processing time of workpiece j-1 at stage s; Ct j,s+1 Denotes the completion time of workpiece j at stage s+1; t j,s+1 Denotes the processing time of workpiece j at stage s+1; Ct j′,s Denotes the completion time of workpiece j′ at stage s; t j′,s Denotes the processing time of workpiece j′ at stage s; L represents a positive number that is infinitely large; z j,j′,s,f Denotes if workpiece j is processed before j′ at stage s in factory f z j,j′,s,f =1, otherwise z j,j′,s,f =0; z j',j,s,t Denotes if workpiece j′ is processed before j at stage s in factory f z j',j,s,t =1, otherwise z j',j,s,t =0; y i,j',s,t Denotes if workpiece j′ is processed on machine i at stage s in factory f, y i,j',s,t =1, otherwise y i,j,s,f =0; j≠j′;
[0153] Then there are the first, second, and third decision variables, and the expressions are as follows:
[0154]
[0155] S2. Use the hyper-heuristic algorithm of Q-learning to solve the optimal scheduling problem of the mixed-integer linear programming (MILP) model for the distributed heterogeneous hybrid job shop scheduling problem (DHHFSP), and the steps are as follows:
[0156] S2.1. Establish the Q-learning model, and the expression is as follows:
[0157]
[0158] In the formula, s t Denotes the current state; a t Denotes the current action; s t+1 Denotes the state after executing action a t ; A denotes the total action set; a denotes an action, a∈A; r t+1 Denotes the instantaneous reinforcement signal; λ∈[0,1] denotes the dynamic learning rate, that is, the update amplitude of each update; γ∈[0,1] denotes the discount factor; Q t (s t ,a t ) denotes the Q value;
[0159] The core of Q-learning is the Q value, which represents the expected return of taking a certain action in a specific state; the Q value is stored and updated through a Q-table, and the Q-table uses the Q-learning model to calculate the cumulative reward value;
[0160] S2.2. Set the key parameters of the Q-learning model and set the termination conditions;
[0161] Among them, the key parameters include: population size ps, initial learning rate λ, discount factor γ, initial greed rate ε; the termination condition is the running time of 500 seconds;
[0162] The method of setting the Q-learning model is as follows: Set the key parameters at 5 levels, obtain experimental data through different groups, and analyze for better parameter groups; Combine different values of the 4 key parameters, and use the full-factor experimental design method to analyze the sensitivity of the parameters through the analysis of variance (ANOVA) technique. Each parameter group runs independently and repeatedly 10 times. The different group settings and results are shown in Table 1;
[0163] Table 1: Different group settings and results
[0164]
[0165]
[0166] The best values in the table are marked in bold;
[0167] The best parameter combination obtained is: ps = 40, λ = 0.3, γ = 0.7, ε = 0.4;
[0168] S2.3. Define 5 low-level individuals (LLH) to initialize the population and the Q-table;
[0169] The 5 low-level heuristic (LLH) rules include:
[0170] (1). Key workpiece pre-insertion: Randomly select a key workpiece from the key workpieces in the workpiece set of a key factory in the factory set, and insert this workpiece before other workpieces. If the makespan (f 1 ) and the total electricity price (f 2 ) are optimized for one or both of them, the insertion operation is completed; otherwise, continue to select the next position;
[0171] (2). Key workpiece post-insertion: Randomly select a key workpiece from the key workpieces in the workpiece set of a key factory in the factory set, and insert this workpiece after other workpieces. If the makespan (f 1 ) and the total electricity price (f 2 ) are optimized for one or both of them, the insertion operation is completed; otherwise, continue to select the next position;
[0172] (3) Insertion of non-critical workpiece: Randomly select a workpiece from the non-critical workpieces in the factory set and insert this workpiece before the workpieces in other non-critical factories. If the earliest completion time (f 1 ) and the total electricity price (f 2 ) are optimized either one or both, the insertion operation is completed; otherwise, continue to select the next position.
[0173] (4) Post-insertion of non-critical workpiece: Randomly select a workpiece from the non-critical workpieces in the factory set and insert this workpiece after the workpieces in other non-critical factories. If the earliest completion time (f 1 ) and the total electricity price (f 2 ) are optimized either one or both, the insertion operation is completed; otherwise, continue to select the next position.
[0174] (5) Exchange of critical workpiece: Randomly select a critical workpiece from the critical path and then exchange this workpiece with other workpieces. If the earliest completion time (f 1 ) and the total electricity price (f 2 ) are optimized either one or both, the exchange operation is completed; otherwise, continue to select the next position.
[0175] Critical path: In the hybrid flow shop scheduling problem, finding the critical path is an important step in optimizing the scheduling. The critical path is the path that determines the minimum earliest completion time in the scheduling plan, which represents the longest time path from the start to the completion of the task. The processing delay of each task on the critical path will directly lead to the delay of the completion time. The determination of the critical path is as Figure 2 shown.
[0176] Critical factory: The factory with the largest overall completion time; the factory that plays a leading role in the total scheduling time, and its optimization has the greatest impact on the overall performance.
[0177] The reasons for selecting the above LLH are as follows:
[0178] There are pre-insertion and post-insertion operations for critical factories and non-critical factories. The operations of critical factories have a significant impact on changing the earliest completion time (f 1 ), while the operations of non-critical factories have a significant impact on changing the total electricity price (f 2 ). However, if only pre-insertion and post-insertion operations are performed, the search space is not large enough. Therefore, workpiece exchange is added to expand the search space, which is beneficial to jumping out of the local optimum.
[0179] The current state s tThe fitness expression for multiple objectives in the distributed heterogeneous hybrid job shop scheduling problem (DHHFSP) is defined as follows:
[0180]
[0181] In the formula, z′ min and z′ max represent the minimum and maximum objective values; o represents the number of objectives. In this invention, there are two objectives (f 1 and f 2 ), so o = 2; because the two objectives have different orders of magnitude, the objective values need to be normalized to [0, 1]. When q = 1, f q (π) represents the minimum makespan of the first objective of the current individual, z′ min,q is the minimum value of the minimum makespan in the population, and z′ max,q represents the maximum value of the minimum makespan in the population; when q = 2, f q (π) represents the total electricity cost of the second objective of the current individual, z′ min,q is the minimum value of the total electricity cost in the population, and z′ max,q represents the maximum value of the TEC in the population;
[0182] Define the selection of LLH as an action. During the Q-table update process, there are 5 choices from the current action to the next action. Therefore, a 5×5 probability matrix is used as the Q-table of the algorithm;
[0183]
[0184] In the formula: gen represents the number of generations of the algorithm running. When initialized, gen = 1; the element p xy (gen)(x = 1,..., 5, y = 1...5) represents the probability that the next second action of the first action x is y (here the action is defined as the operation performed after selecting LLH), and p ij (1) all take the value of 1 / 5; the Q-table update formula is shown in Equation (14), where n is the number of jobs, and Q t (s t , a t ) refers to the Q value of taking action a t at state s t , maxQ(s t+1 , a) represents the maximum Q value at state s t+1 when all actions are executed, and λ ∈ [0, 1] is the dynamic learning rate;
[0185] At the initial iteration, the Q values corresponding to all state-actions are set to zero, so as to randomly select the initial state and action;
[0186] The initial population size is equal to 40 set in Table 1. It is necessary to generate 40 populations. A hybrid initialization strategy is adopted to generate the populations, that is, randomly generate 70% (28) individuals, and then use the NEH algorithm to generate 30% (12) individuals;
[0187] The NEH algorithm decodes while encoding. Add jobs to the sequence one by one. The added job tries all positions in the existing sequence and records the minimized earliest completion time of the partial sequence under the tried position (f 1 ) total electricity price (f 2 ), and calculate the fitness of all partial sequences by using Equation (15); after all positions have been tried for insertion, find the position θ with the minimum fitness, and insert the job into position θ until all jobs have been inserted and the generated individuals are added to the population;
[0188] After generating 40 populations, perform the initialization operation on the initial population, as shown in Algorithm 1;
[0189]
[0190] S2.4. Select 5 high-level populations (HLI) from the populations generated by low-level individuals (LLH) to update the Q table, and select the highest Q value from the updated Q table using the greedy strategy;
[0191] The selection method is: select the top 5 high-level individual populations (HLI) optimally by calculating the contribution rate of the populations generated by low-level individuals (LLH); among them, the contribution rate is the average value of the fitness of all individuals in the population (calculate the individual fitness fitness by Equation (15));
[0192] The method of updating the Q table is: select the optimal high-level individuals by the contribution rate and update the Q table by Equation (14);
[0193] The iterative formula of the greedy rate ε is as follows:
[0194]
[0195] In the formula, ε initial is the initial value; ε final is the final value; gen is the number of generations of algorithm operation, gen l is the current running generation;
[0196] The initial greedy rate ε is set to 0.4;
[0197] The greedy strategy formula is:
[0198]
[0199] where ε t is the iterative greedy rate, and the iterative formula is Equation (17);
[0200] By continuously updating the greedy rate, the probability of Q-learning to select the best action is also slowly changing. In the early stage of the algorithm operation, the algorithm obtains more search space. Therefore, the value of ε is very large, and there will be more randomly selected actions in Equation (18). Similarly, in the later stage of the algorithm, the algorithm selects some more valuable actions to find a smaller total completion time and total electricity price. Therefore, the value of ε slowly becomes smaller, so that the action with the largest Q value in the Q table is selected in Equation (18) to operate on the population;
[0201] Check whether the number of iterations has reached the maximum number of iterations. If not, go to S2.3; if satisfied, end;
[0202] The number of iterations is set to the maximum number of iterations of 100. If satisfied, output the Pareto front; otherwise, go to S2.3 and iterate repeatedly until the maximum number of iterations is satisfied;
[0203] S2.5. After satisfying the termination condition, execute the right shift strategy to reduce the total electricity price;
[0204] The specific right shift operation is as follows:
[0205] While keeping the maximum completion time unchanged, transfer the production in the peak electricity price interval to the relatively low electricity price interval through the right shift operation to save the electricity cost; restricted by the characteristics of the DHHFSP problem, the right shift operation needs to be adjusted sequentially from the last stage forward; perform the right shift operation sequentially from the first factory to the last factory; sequentially search for the machine idle time from the last stage to the first stage in the first factory, and shift the workpiece during the idle time period, and calculate the energy consumption of the current workpiece at the current stage, and compare the original energy consumption with the energy consumption after the right shift; if the original energy consumption is greater than the energy consumption after the right shift, it is judged that the workpiece right shift will be in the low electricity price interval; then perform the right shift operation, otherwise, do not perform the right shift operation; the schematic diagram of the right shift operation is as Figure 3 shown. It can be seen from Figure 3 parts a and b in that after the right shift, the operation time of machine 3 in factory 1 and factory 2 is shifted to the right as a whole, which moves the start time of the machine backward and shortens the startup time and standby time, thereby reducing the energy consumption, achieving the purpose of energy saving, and the TEC of machine 3 calculated by Algorithm 2 also decreases, achieving the dual advantages of reducing energy consumption and saving electricity costs;
[0206] The right shift strategy is shown in Algorithm 2;
[0207]
[0208] S3. Verify the accuracy of the model;
[0209] Specifically, experiments are conducted to verify the effectiveness of the comparison verification method with four existing algorithms for solving DHHFSP. The four algorithms are: Non-dominated Sorting Genetic Algorithm II (NSGA-II), Multi-Objective Evolutionary Algorithm Based on Decomposition (MOEA / D), Multi-Differential Evolution (MDDE), and Knowledge-Based Cooperative Algorithm (KCA) for comparison.
[0210] The present invention uses 45 test cases to verify the scientificity and effectiveness of the above-mentioned energy-saving distributed assembly hybrid flow shop scheduling method.
[0211] The number of machines, processing energy consumption, and average processing time of each workpiece in different stages are obtained through randomly generated data. The number of workpieces n = {20, 50, 100}, the number of factories F = {2, 3, 4, 5, 6}, and the number of processing stages s = {2, 5, 8}. The instances are described as n×F×S, with a total of 45 instances. The processing time of each workpiece in each stage is randomly generated in [1, 99]. The processing power of each machine is randomly generated in [20, 50]. The idle power is randomly generated from the uniform distribution [1.0, 2.0]. The processing energy consumption is the processing power multiplied by the processing time, the idle energy consumption is the idle power multiplied by the idle time, and the start-up and shutdown energy consumption is randomly generated in [1.0, 10.0]. Regarding the time-of-use electricity price, each factory randomly selects one Figure 4 、 Figure 5 and the time-of-use electricity price schemes shown in formulas (19) and (20). The total energy consumption (TEC) can be obtained by multiplying the energy consumption by the electricity price at the corresponding time. In most parts of China, the 24 hours of a day are divided into three or five time periods, namely, peak period, mid-peak period, and off-peak period, and each time period has a different electricity price. The peak period electricity price is the highest, the mid-peak electricity price is the second highest, and the off-peak period electricity price is the lowest.
[0212] From Figure 4 and Figure 5 it can be seen that for two factories in different regions, the electricity price strategies are different. Workpieces {1, 2, 3, 6, 9} are processed in factory 1, and workpieces {4, 5, 7, 8, 10} are processed in factory 2. It can be seen that each workpiece has to go through two stages, and there are more than one machine to choose from in each stage, which conforms to the model of the DHHFSP problem. At the same time, the time-of-use electricity price strategy is considered, where the time-of-use electricity price model h f (t) is as follows:
[0213]
[0214] Compared with the fast elitist multi-objective genetic algorithm (NSGA-II), the multi-objective evolutionary algorithm based on decomposition (MOEA / D), the differential evolution algorithm (MDDE), and the knowledge-based collaborative algorithm (KCA), the Pareto frontiers of the experiments are as Figure 6 , Figure 7 shown.
[0215] Figure 6 shows the Pareto frontiers distribution of different algorithms in terms of makespan (completion time) and TEC (total energy consumption) under the conditions of 2 factories, 20 jobs, and 5 stages.
[0216] Figure 7 shows the Pareto frontiers distribution of different algorithms in terms of makespan (completion time) and TEC (total energy consumption) under the conditions of 2 factories, 50 jobs, and 2 stages.
[0217] Figure 6 and Figure 7 show:
[0218] In this invention, QLLHEA (diamond), NSGA-II (circle), MOEA / D (diamond), MDDE (pentagram), KCA (triangle);
[0219] Analyze the effectiveness of QLLHEA:
[0220] 1. Distribution trend of the Pareto front:
[0221] In the two figures, the Pareto front (diamond line) generated by QLLHEA is generally closer to the lower left. In the Pareto diagram, the abscissa is makespan with the unit of hour; the ordinate is TEC with the unit of yuan. The closer to the lower left, the smaller the makespan and TEC of the solution obtained by the algorithm, which means it performs better when optimizing makespan and TEC simultaneously.
[0222] QLLHEA can find better solutions, achieving a better balance between completion time and energy consumption, and covering a better area of the solution space.
[0223] 2. Diversity and uniformity of solutions:
[0224] Uniformity of solution distribution: The solutions of QLLHEA are more evenly distributed on the Pareto front, which means a large distribution span on the entire coordinate axis, indicating more non-dominated solutions and providing more and better solutions. Especially in two test problems with different scales (20 jobs and 50 jobs), good solution diversity is demonstrated.
[0225] The uniform distribution of the solution set indicates that QLLHEA can provide diverse choices among different optimization objectives, making it more adaptable and performing better in different cases (small-scale and large-scale cases). This shows that the algorithm can find better solutions in both small-scale and large-scale cases, reflecting its stronger adaptability.
[0226] 3. Comparison with other algorithms:
[0227] Comparison with the Non-dominated Sorting Genetic Algorithm II (NSGA-II), the Multi-Objective Evolutionary Algorithm Based on Decomposition (MOEA / D), the Multi-Dimensional Differential Evolution (MDDE), and the Knowledge-based Cooperative Algorithm (KCA): Most of the solution sets of NSGA-II and MOEA / D are located in the upper right corner of the Pareto front, indicating that they cannot find a good balance between the makespan and energy consumption simultaneously.
[0228] Significant advantages: The solution set of QLLHEA performs better in both metrics, with a smaller makespan and a smaller TEC. Compared with other algorithms, it can significantly reduce energy consumption and makespan.
[0229] 4. Performance stability:
[0230] In test instances of different scales, QLLHEA has always been able to find solutions with a smaller makespan and a smaller TEC. For other algorithms (such as MOEA / D and MDDE), the quality of the solutions decreases significantly when the scale is large, while the solutions of QLLHEA are still excellent.
[0231] In summary, QLHHEA has obvious advantages. As the number of jobs, stages, and factories increases, the difficulty of solving the problem increases, and QLLHEA demonstrates stronger search and optimization capabilities.
Claims
1. An optimization method for distributed heterogeneous hybrid workshop scheduling problem, characterized in that: The following steps are involved: S1. Establish a mixed integer linear programming model for the distributed heterogeneous hybrid workshop scheduling problem and set constraints on the model; The mixed integer linear programming model contains two objective functions: f1 with the goal of minimizing the total completion time and f2 with the goal of minimizing the total electricity price; S2. Use the Q-learning hyper-heuristic algorithm to solve the optimization scheduling problem of the mixed integer linear programming model of the distributed heterogeneous hybrid workshop scheduling problem and obtain the optimization solution; S3. Verify the accuracy of the model.
2. The optimization method for distributed heterogeneous hybrid workshop scheduling problem according to claim 1 is characterized in that: The expressions of f1 with the goal of minimizing the total completion time and f2 with the goal of minimizing the total electricity price are as follows: f1=C max =max{Ct j,s },1≤j≤n,1≤s≤S In the formula, Ct j,s represents the completion time of job j at stage s; n represents the number of jobs; S represents the number of stages; S represents the stage set, S = (1, 2, ..., S); F represents the number of factories; f represents the index of the factory, f = {1, 2, ..., F}; m represents the number of machines; i represents the machine index, i = {1, 2, ..., m}; T represents the time period set; t represents the time period index, t = {1, 2, ..., T}; EC f,i represents the total energy consumption of machine i in factory f; TEC run Indicates the electricity price during machine operation; TEC idle Indicates the electricity price when the machine is idle; TEC off / on Indicates the electricity price during the machine startup and shutdown process; h f (t) represents the corresponding relationship between time-of-use electricity price and time interval; The setting of constraints on the model includes: Each artifact can be assigned to one plant and one plant only; Each factory contains at least one artifact; Each workpiece must go through all the processes, and each process can only be placed on one machine for processing; The start time of the workpiece is greater than or equal to 0; The completion time of a workpiece is greater than the start time, and the start time of the next workpiece is greater than the completion time of the previous workpiece; At the same stage, a workpiece cannot be processed by multiple machines at the same time, and one machine can only process one workpiece at a time; A first decision variable, a second decision variable, and a third decision variable are determined.
3. The optimization method for distributed heterogeneous hybrid workshop scheduling problem according to claim 1 is characterized in that: The steps of using the Q-learning hyper-heuristic algorithm to solve the optimization scheduling problem of the mixed integer linear programming model of the distributed heterogeneous hybrid workshop scheduling problem and obtain the optimization solution are as follows: S2.
1. Establish a Q-learning model; S2.2, set the key parameters of the Q learning model and set the termination conditions; Among them, the key parameters include: population size ps, initial learning rate λ, discount factor γ, initial greedy rate ε; S2.3, define 5 low-level individuals to initialize the population and Q table; Among them, the five low-level heuristic rules include: 1) Forward insertion of key workpieces; 2) Insert key workpieces later; 3) Non-critical workpieces are inserted forward; 4) Non-critical workpieces are inserted later; 5) Exchange of key workpieces; S2.4, select 5 high-level populations from the population generated by low-level individuals to update the Q table, and use the greedy strategy to select the highest Q value from the updated Q table; The selection method is: select the top 5 best high-level individual populations by calculating the contribution rate of the population generated by low-level individuals; among them, the contribution rate is the average fitness of all individuals in the population, and the fitness expression is as follows: In the formula, z′ min and z′ max represents the minimum and maximum target values; o represents the number of targets. The present invention has two targets f1 and f2, so o = 2; because the two targets have different orders of magnitude, the target values are normalized to [0, 1]. When q = 1, f q (π) represents the earliest completion time of the first goal of the current individual, z′ min,q is the smallest value in the population that minimizes the earliest completion time, z′ max,q represents the value of the largest minimum earliest completion time in the population; when q = 2, f q (π) represents the second target total electricity price of the current individual, z′ min,q is the minimum total electricity price in the population, z′ max,q It represents the maximum TEC value in the population; The Q-table is updated by selecting the best high-level individuals through contribution rate to update the Q-table; S2.
5. After the termination condition is met, the right shift strategy is executed to reduce the total electricity price.
4. The optimization method for distributed heterogeneous hybrid workshop scheduling problem according to claim 3 is characterized in that: In the above-mentioned definition of 5 low-level individuals to initialize the population and Q table, the Q table adopts a 5×5 probability matrix; The population initialization adopts a mixed initialization strategy to generate the population, that is, 70% of the individuals are randomly generated, and then the NEH algorithm is used to generate 30% of the individuals.
5. The optimization method for distributed heterogeneous hybrid workshop scheduling problem according to claim 3 is characterized in that: The accuracy of the verification model is verified through experiments by comparing it with four existing algorithms for solving distributed heterogeneous hybrid workshop scheduling problems to verify the effectiveness of the verification method; the four algorithms are: fast elite multi-objective genetic algorithm, decomposition-based multi-objective evolutionary algorithm, differential evolution algorithm, and knowledge-based collaborative algorithm.
Citation Information
Cited By
Full-active-scheduling multi-workshop combined scheduling method and system and storage medium
CN122219381A