An energy-saving job shop scheduling system for batch flow jobs
Patent Information
- Application Number
- CN202610670705.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-15
- Publication Date
- 2026-09-22
AI Technical Summary
[0006]然而,针对带可变子批的节能批量流柔性作业车间混排调度问题,现有方法在兼顾总延迟与总能耗双目标优化的同时,尚难以实现子批划分、工序排序与机器分配的协同优化,且协同进化算法与强化学习的深度融合策略仍有待进一步探索
[0047](1)本发明定义了带可变子批的节能批量流柔性作业车间混排调度问题的整数规划模型,并提出了一种将能耗和完工时间相结合的评价准则。
Smart Images

Figure CN122797987A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of multi-variety, small-batch production scheduling technology in the manufacturing industry, and particularly relates to an energy-saving workshop scheduling system for batch flow operations. Background Technology
[0002] In industrial production, with the increasing demand for personalized customization and the widespread adoption of multi-variety, small-batch production models, order density has significantly increased. Therefore, to meet the huge market demand, most enterprises adopt diversified methods to improve production capacity. However, traditional job shop scheduling methods are no longer sufficient to meet the needs of personalized production. Batch flow scheduling, as an effective strategy, is widely used to cope with high-density orders and improve equipment utilization and production efficiency. In batch flow scheduling, each workpiece is divided into multiple sub-batches, allowing overlapping processing between different processes to shorten completion time. Depending on the sub-batch division method, the batch flow scheduling problem can be divided into equal-quantity batch scheduling, uniform sub-batch scheduling, and variable sub-batch scheduling. Equal-quantity batch scheduling and uniform sub-batch scheduling limit the improvement of production efficiency. In contrast, variable sub-batch scheduling, by allowing the sub-batch size to change flexibly between different processes or within the same process, can achieve higher resource utilization and production efficiency, and therefore has received widespread attention from academia and industry. On the other hand, based on whether other products are allowed to be inserted into the processing of the same product during the processing, batch flow scheduling can be divided into mixed-batch scheduling and non-mixed-batch scheduling problems. Batch flow scheduling technology is widely used in production planning and scheduling optimization. It can effectively shorten the production cycle, improve equipment utilization, and has a significant effect on improving product quality and reducing production costs.
[0003] Customers are increasingly focusing on the reliability of on-time product delivery, emphasizing that workpieces should be completed and delivered on schedule. Delivery delays not only affect company reputation but also reduce customer satisfaction. Total Tardiness (TTD), as a key objective for optimizing scheduling, is used to measure the efficiency and timeliness of the production or task completion process. Therefore, using TTD as an optimization target more closely aligns with the needs of actual production processes.
[0004] Due to the high complexity of LSFJSP-VI, global search requires more computational resources to fully explore the target space. Traditional methods such as cooperative search and meme algorithms typically perform global and local searches simultaneously within the main population, which reduces population diversity to some extent. Inspired by the divide-and-conquer strategy, researchers have proposed a Co-evolution Algorithm (CEA) to solve NP-hard problems by dividing computational resources into multiple subproblems. The CEA is inspired by the mutual adaptation and co-evolution among species in nature. Unlike traditional evolutionary algorithms that use a single population to evolve independently, the CEA decomposes complex problems into several interrelated subproblems, maintaining an independent subpopulation for each subproblem. These subpopulations co-evolve through information exchange and collaborative evaluation. This "divide and conquer" strategy effectively reduces the search space complexity of high-dimensional problems while preserving the coupling relationships between subproblems.
[0005] Reinforcement Learning (RL) is an important branch of machine learning. Its core idea is that an agent continuously interacts with its environment, selects actions based on its current state, and continuously optimizes its decision-making strategy based on reward signals from the environment, ultimately maximizing long-term cumulative rewards. Unlike supervised learning, which relies on labeled data, and unsupervised learning, which mines the inherent structure of data, reinforcement learning employs a "trial and error" mechanism, autonomously learning the optimal strategy in a dynamic and unknown environment. Therefore, it has significant advantages in sequential decision-making, adaptive control, and other fields. The basic framework of reinforcement learning consists of five core elements: state space, action space, state transition probabilities, reward function, and discount factor. Its learning process is typically modeled as a Markov Decision Process (MDP). Q-learning, as one of the most classic model-free algorithms in reinforcement learning, does not require knowledge of the state transition probabilities of the environment. It approximates the optimal strategy by iteratively updating the state-action value function. In recent years, Q-learning has been widely used to solve various combinatorial optimization problems due to its simple structure, theoretically guaranteed convergence, and lack of need for an environment model. In scheduling problems, Q-learning has been used to dynamically select scheduling rules, adaptively adjust search parameters, and guide the selection of local search operators, effectively improving the algorithm's adaptability to dynamic changes in the problem. Combining Q-learning with metaheuristic algorithms can achieve adaptive adjustment of algorithm behavior without increasing computational overhead, becoming one of the current research hotspots in the field of intelligent optimization.
[0006] However, for the problem of mixed scheduling in flexible job shops with variable sub-batches, existing methods, while achieving dual objectives of optimizing total delay and total energy consumption, still struggle to achieve coordinated optimization of sub-batch partitioning, process sequencing, and machine allocation. Furthermore, the deep integration strategy of co-evolutionary algorithms and reinforcement learning requires further exploration. Therefore, there is an urgent need for an efficient scheduling method that can simultaneously optimize total delay and total energy consumption, support variable sub-batch partitioning and mixed processing, and integrate co-evolutionary and reinforcement learning strategies. This would address the shortcomings of existing technologies in the coordinated optimization of sub-batch partitioning, process sequencing, and machine allocation, as well as their weak adaptive adjustment capabilities. Summary of the Invention
[0007] The purpose of this invention is to provide an energy-efficient shop floor scheduling system for batch flow operations, which aims to simultaneously minimize total tardiness (TTD) and total energy consumption (TEC) to solve the problem of mixed scheduling in flexible shop floors with variable sub-batches. This scheduling system can optimize the operating efficiency and performance of mixed scheduling systems for flexible shop floors with variable sub-batches.
[0008] To achieve the above objectives, the technical solution of the present invention is as follows:
[0009] An energy-saving workshop scheduling system for batch workflow operations is characterized by comprising a population initialization module, a co-evolutionary search module, and a local reinforcement module;
[0010] The population initialization module is used to generate an initial population with high quality and diversity; wherein, the scheduling solution is represented by three vectors, namely Layer1, Layer2, and Layer3; Layer1 specifies the number of sub-batches into which each operation of the job is divided. Layer2 defines the size of each sub-batch. Layer3 describes the mixed sub-batch processing sequence on the machine;
[0011] The population initialization module includes:
[0012] The first initialization method is used to determine the number of sub-batches for each process in each workpiece, where the number of sub-batches is limited to a predefined range. Within this interval, each candidate value is used as a corresponding weight and normalized to obtain a probability distribution. A roulette wheel selection mechanism is constructed based on the cumulative probability, and the number of sub-batches is determined by random numbers. A second initialization method is used to determine the size of each sub-batch, and a method that simultaneously includes uniform and non-uniform sub-batch sizes is designed; a uniform random perturbation factor is introduced. To control the division of the batch size of the sub-batch, and the batch size of each sub-batch is determined by formula (1), which provides flexibility in adjusting the size of the subtask; in order to further ensure the feasibility of the solution, a repair mechanism is applied during the adjustment process, and the batch size that has reached the minimum size of 1 cannot be further reduced, thereby avoiding the generation of invalid (zero or negative) batch sizes.
[0013]
[0014] in, Let i be the batch size of workpiece i. Let i be the number of sub-batches of workpiece i in process j, and This represents the batch size of each sub-batch;
[0015] The third initialization method is used to determine the mixed sub-batch processing sequence on all machines. This initialization method clarifies the relationship between processing and delivery, enabling scheduling based on the urgency of delivery, thereby promoting on-time delivery. The urgency of delivery depends on the delivery time, the number of jobs, and the processing time. Scheduling rules based on the urgency of delivery are developed. First, jobs are sorted according to importance and deadline, generating priority index and deadline index. These selection probabilities are combined using an addition operator to form a comprehensive probability, which is used to arrange the sub-batches of each job in order. Job priority and sub-batch priority are calculated by formula (2) and formula (3) respectively to accurately determine the job priority and the urgency of sub-batch delivery.
[0016] in, For the batch size of each workpiece, The batch size for each sub-batch, Let be the processing time of process j for workpiece i on machine k. Let be the deadline for workpiece i. and Processing priority.
[0017] The co-evolutionary search module is used to promote the effective evolution of different sub-problems and the efficient collaboration between sub-problems, including an information storage mechanism and an evolutionary operator selection strategy;
[0018] The information storage mechanism is used to obtain accurate sub-block partitioning and mixed sub-batch processing sequences. This mechanism constructs three sets: a sub-batch quantity set, a sub-batch size set, and a mixed sub-batch processing sequence set. These sets store the sub-batch quantity, sub-block size, and mixed sub-batch processing sequence of the non-dominated solution set. After each iteration, the information storage mechanism selects appropriate sub-batch quantity, sub-batch size, and mixed sub-batch processing sequence from the non-dominated solution set based on target value comparisons, ensuring the capture of high-quality solutions. Four evolutionary operators, EO1, EO2, EO3, and EO4, are designed for a specific problem to determine and identify appropriate sub-batch quantity, sub-batch size, and mixed sub-batch processing sequence. The encoding structures extracted from the sub-batch quantity set, sub-block size set, and mixed sub-batch processing sequence set are denoted by SNS, SSS, and ISS, respectively. The specific design of these four evolutionary operators is described below:
[0019] EO1: The ISS is aligned with the current mixed sub-batch processing sequence to identify reference sub-batches: any sub-batch whose workpiece index and operation index match the ISS is considered a reference sub-batch; if no reference sub-batch exists in the sequence, an intermediate sub-batch is selected as the reference. Subsequently, for every interval between two adjacent reference sub-batches, the first sub-batch in that interval is moved to the end, and the remaining sub-batches are moved forward one position in sequence; if the interval contains only one sub-batch, it remains unchanged; intervals before the first reference sub-batch and after the last reference sub-batch are also treated in the same way. This exchange operation will disrupt the processing order of the operations. Finally, a repair operation ensures that all sequences and constraints are satisfied.
[0020] EO2: Randomly select two indices a and b, where a∈{1,…,N} represents the index of a certain workpiece in the current solution, and b∈{1,…,M} represents the index of a certain operation in workpiece a; after determining a and b, replace the corresponding fragments of operation b in Layer1 and Layer2 of operation a with the corresponding fragments in SNS and SSS, respectively.
[0021] EO3: Randomly select a workpiece index a∈{1,…,N} from the current individual; then, replace the corresponding fragments of all processes of the selected workpiece a in Layer1 and Layer2 with the corresponding fragments in SNS and SSS;
[0022] EO4: Randomly select a process index b∈{1,…,M} from the current individual, and then replace the corresponding fragments of all workpieces in Layer1 and Layer2 with the corresponding fragments in SNS and SSS.
[0023] 4. An energy-saving workshop scheduling system for batch flow operations according to claim 3, characterized in that the Evolutionary Operator Selection Strategy (EOSS) is used to solve the problem that the four evolutionary operators have different effects at different iteration stages; EOS converts the historical success rate and failure rate of each evolutionary operator as well as the success rate and failure rate in the current iteration into a quantitative knowledge score, as shown in formula (4); the knowledge score reflects the overall utility of each operator and is normalized by formula (5) to form a probability distribution; then, the roulette wheel method is used to select the most promising operator; EOS can dynamically adjust its selection probability according to the performance of the operator, thereby enhancing the algorithm's ability to develop in effective search directions;
[0024]
[0025] in, and They represent evolution operators respectively. The historical number of successes and failures when generating nondominated solutions in previous iterations; and They represent evolution operators respectively. The historical number of successes and failures when generating non-dominated solutions in the current iteration; 'a' is the success factor, used to increase the selection probability of the operator; 'b' is the failure factor, used to decrease the selection probability of the operator; 'ε' is used to prevent the operator weight from becoming 0, thus ensuring that all operators have a chance to be selected in the initial stage of the algorithm. This represents the knowledge value of the x-th evolution operator; The normalized knowledge value represents the selection probability of the x-th evolutionary operator. When an evolutionary operator produces a new solution, its corresponding knowledge value and selection probability will be updated.
[0026] The local enhancement module is a local search based on problem characteristics. Based on the characteristics of sub-batch adjustment, critical paths, and critical blocks, six local search operators are designed to reduce TEC and TTD. These six local search operators are denoted as LSOi, i∈[1,6], and are as follows:
[0027] LSO1: Randomly select a key sub-batch and map it to the corresponding sub-batch in Layer 2. Let z represent the current batch size of this sub-batch. This represents the total batch size of the job, with the batch size of sub-batches perturbed by h, where... ;
[0028] LSO2: Select the critical sub-batch with the largest batch size and divide it into two equal-sized sub-batches.
[0029] LSO3: Exchanges an urgent sub-batch with its preceding sub-batch on the same machine, thereby reducing its dwell time in the system and further shortening the waiting time of subsequent sub-batch.
[0030] LSO4: On the critical path, split the longest delayed sub-batch and insert it into an earlier available time period to reduce TTD;
[0031] LSO5: N6 is a robust neighborhood structure designed for the job shop scheduling problem. Its goal is to optimize scheduling by changing the order of operations in the critical block. N6 reduces completion time through the following steps: In the critical path, randomly select a non-last operation from the leading segment π1 and move it to the end of π1; randomly select a pair of operations in the middle segment π2, move one to the beginning of π2 and the other to the end of π2; randomly select a non-first operation from the tail segment πn and move it to the beginning of πn.
[0032] LSO6 is an improvement on the Deadline Neighbor Relation (DNR) method. DNR refers to the situation where two adjacent sub-batches S1 and S2 constitute a critical sub-batch pair on the same machine: when the completion time of S1 is earlier than its delivery time, while the completion time of S2 is later than its delivery time; if identified as a critical sub-batch pair, the positions of these two sub-batches are swapped to reduce TTD.
[0033] 5. An energy-saving workshop scheduling system for batch flow operations according to claim 5, characterized in that it further includes a Q-learning operator selection mechanism, which uses Q-learning to update the solutions in the Pareto optimal set, constructs operator selection strategies for different solutions through Q-learning, and uses a 4×3 index Q-table to assist in selecting appropriate local search operators. The update of the Q-table is shown in formula (6).
[0034]
[0035] Where reward and α represent the action reward and learning rate, respectively. and These correspond to the state and action of the new solution, respectively; then, the following Q-learning mechanism is used to select appropriate operators for different solutions;
[0036] This mechanism uses the solutions in each generation of the Pareto optimal set as the agent. By defining the state space, reward function, and action set, and combining the ε-greedy policy, it achieves adaptive selection of the operator. Let π represent the current solution and π* represent the newly generated solution; the specific definitions are as follows:
[0037] Agent: the solutions in the Pareto optimal set of each generation act as agents;
[0038] State: the state is defined by comparing the changes of π and π* on TEC and TTD, so as to help the agent adapt to different optimization scenarios: state 0 represents TEC(π)<TEC(π*) and TTD(π)<TTD(π*); state 1 represents TEC(π)<TEC(π*) and TTD(π)>TTD(π*); state 2 represents TEC(π)>TEC(π*) and TTD(π)<TTD(π*); state 4 represents TEC(π)>TEC(π*) and TTD(π)>TTD(π*);
[0039] Reward: the reward value is proportional to the improvement amplitude; therefore, the reward value is calculated by comparing the improvement of solution π and solution π*, as shown in formula (7); this formula measures the relative reduction of TEC and TTD, and selects the smaller improvement value as the final reward to ensure stability;
[0040] Action: the local search operation is regarded as an action, and a plurality of local search operators are combined to form an action; an action is selected from the action set as a combined local search operator to optimize the target value; according to the design, 3 actions are defined in total, forming an action set: Action={A1, A2, A3}. A1 comprises LSO1 and LSO2, which is used for the sub-lot adjustment strategy; A2 comprises LSO3 and LSO4, which targets sub-lots on the critical path; A3 comprises LSO5 and LSO6, which is used for critical block operation;
[0041] Under state s, an ε-greedy strategy is used to select an action, as shown in formula (8):
[0042]
[0043] Wherein, a random number rand is generated, when rand<ε, a greedy strategy is adopted in the k-th row of the Q-table to select the action with the maximum Q value; otherwise, an action is randomly selected; once the action is selected, π will execute all search operators in the action, and select the optimal operator for further search through a greedy strategy.
[0044] In a second aspect, the present invention provides a computer-readable storage medium, which contains a computer program that can implement the above method steps when processed by a CPU.
[0045] This invention designs a solver for a flexible job shop scheduling system with variable sub-batches and energy-saving batch flow. The method is based on a knowledge-driven co-evolutionary algorithm and aims to minimize total delay and total energy consumption to solve the flexible job shop scheduling problem with variable sub-batches and energy-saving batch flow. First, an integer programming model for the flexible job shop scheduling problem with variable sub-batches and energy-saving batch flow is defined, and an evaluation criterion combining energy consumption and completion time is proposed. Second, a high-quality initial population is generated using a knowledge-driven initialization method based on an encoding structure. During the co-evolutionary search process, the population evolves through evolutionary operators over one generation. An evolutionary operator selection strategy is proposed, which integrates feedback information from the history and current performance of evolutionary operators and transforms this information into quantitative knowledge values. Based on these knowledge values, the operator selection probability is dynamically adjusted to match the solution with better evolutionary operators with different characteristics. Finally, Q-learning is used to select the optimal search operator for non-dominated solutions and find more potential non-dominated solutions locally.
[0046] The present invention has the following beneficial effects:
[0047] (1) This invention defines an integer programming model for the mixed scheduling problem of flexible workshop with variable sub-batches and proposes an evaluation criterion that combines energy consumption and completion time.
[0048] (2) This invention proposes an evolutionary operator selection strategy and transforms the success rate and failure rate of the operator in the historical iteration and the current iteration into quantitative knowledge to guide the algorithm to perform self-learning operator selection.
[0049] (3) The present invention designs a learning engine mechanism based on Q-learning to select the optimal local search operator for non-dominated solutions, thereby reducing invalid search in KMCEA.
[0050] (4) The present invention is simple in logic, easy to implement and easy to extend, and can extend the optimizer to meet most scheduling problems in the current field of intelligent manufacturing production. Attached Figure Description
[0051] Figure 1 Gantt chart for the mixed scheduling problem of flexible job workshops with variable sub-batches and energy-saving batch flow;
[0052] Figure 2 This is a schematic diagram of EO1 in an embodiment of the present invention;
[0053] Figure 3 This is a schematic diagram of EO2 in an embodiment of the present invention;
[0054] Figure 4 This is a schematic diagram of EO3 in an embodiment of the present invention;
[0055] Figure 5 This is a schematic diagram of EO4 in an embodiment of the present invention;
[0056] Figure 6 This is a schematic diagram of the encoding in an embodiment of the present invention;
[0057] Figure 7 This is a schematic diagram of LSO6 in an embodiment of the present invention;
[0058] Figure 8 This is a flowchart of the algorithm in an embodiment of the present invention. Detailed Implementation
[0059] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0060] This invention studies the Energy-efficient Lot-streaming Flexible Job Shop Scheduling Problem with Variable Sublots and Intermingling Sublots (LSFJSP-VI), with objectives including total tardiness (TTD) and total energy consumption (TEC). The Gantt chart for LSFJSP-VI is shown below. Figure 1 As shown.
[0061] An energy-saving workshop scheduling system for batch workflow operations includes:
[0062] 1. Population initialization module
[0063] An initialization method designed specifically for the problem characteristics can obtain an initial population with high quality and diversity. The scheduling solution is represented by three vectors, including Layer1, Layer2, and Layer3, as follows: Figure 6 As shown, Layer 1 specifies the number of sub-batches each operation of the job is divided into. Layer 2 defines the size of each sub-batch. Layer 3 describes the sequence of mixed sub-batches processed on the machine. Based on the coding structure, three knowledge-based initialization methods were identified to generate the complete solution.
[0064] The first initialization method aims to determine the number of sub-batches for each operation within each workpiece. The number of sub-batches is not fixed but limited to a predefined range [L, R]. Within this range, each candidate value is treated as a corresponding weight and normalized to obtain a probability distribution. Subsequently, a roulette wheel selection mechanism is constructed based on the cumulative probability, and the final number of sub-batches is determined according to the position of the random number within the cumulative probability range. The pseudocode for determining the number of sub-batches is shown in Algorithm 1.
[0065]
[0066] The second initialization method is used to determine the size of each sub-batch. When some tasks are divided into uniformly sized sub-batches, the start and finish times of each sub-batch tend to be consistent, thereby minimizing the waiting time between stages. However, in actual scheduling, the next sub-batch does not start until one sub-batch finishes processing, which leads to idle machine time due to resource waiting or higher-priority sub-batches. Some workpieces are divided into highly non-uniformly sized batches, allowing smaller sub-batches to fill the idle machine time, thereby effectively improving machine utilization to solve this problem. Based on this knowledge, a method that simultaneously includes uniform and non-uniform sub-batch sizes is designed. A uniform random perturbation factor Δs is introduced to control the division of the batch size of the sub-batches, and the batch size of each sub-batch is determined by formula (1). Formula (1) provides flexibility in adjusting the size of subtasks. To further ensure the feasibility of the solution, a repair mechanism is applied during the adjustment process. The batch size that has reached the minimum size of 1 cannot be further reduced, thereby avoiding the generation of invalid (zero or negative) batch sizes. The pseudocode for determining the sub-batch size is shown in Algorithm 2.
[0067]
[0068]
[0069] The third initialization method aims to determine the mixed sub-batch processing sequence on all machines. This initialization method clarifies the relationship between processing and delivery, enabling scheduling based on the urgency of delivery, thereby promoting on-time delivery. The urgency of delivery depends on the delivery time, the number of jobs, and the processing time. Scheduling rules based on delivery urgency are developed. First, jobs are sorted according to importance and deadline, generating priority indices and deadline indices. These selection probabilities are combined using an addition operator to form a composite probability used to sequentially arrange the sub-batches of each job. Job priorities and sub-batch priorities are calculated using formulas (2) and (3), respectively, to accurately determine the job priorities and the urgency of sub-batch delivery. The pseudocode for determining the mixed sub-batch processing sequence is shown in Algorithm 3.
[0070] in, For the batch size of each workpiece, The batch size for each sub-batch, Let be the processing time of process j for workpiece i on machine k. Let be the deadline for workpiece i. and Processing priority.
[0071]
[0072]
[0073] 2. Co-evolutionary search module
[0074] In evolutionary search, a comprehensive exploration of the search space is crucial. Therefore, a co-evolutionary search is introduced to promote the efficient evolution of different subproblems and achieve efficient collaboration among them. To further enhance the coordination among subproblems, an information storage mechanism is designed to obtain accurate sub-block partitioning and mixed sub-batch processing sequences. The information storage mechanism constructs three sets: a set of sub-batch numbers, a set of sub-batch sizes, and a set of mixed sub-batch processing sequences. These sets are used to store the sub-batch numbers, sub-block sizes, and mixed sub-batch processing sequences of the non-dominated solution set. After each iteration, the information storage mechanism selects appropriate sub-batch numbers, sub-batch sizes, and mixed sub-batch processing sequences from the non-dominated solution set based on target value comparisons, ensuring the capture of high-quality solutions. The information storage mechanism provides guidance for co-evolutionary search, enhances global search capabilities, and incorporates feedback from historical solutions. Evolutionary operations are designed for specific problems to enrich the search patterns. Four evolutionary operators, including EO1, EO2, EO3, and EO4, are designed to determine and identify appropriate sub-batch numbers, sub-batch sizes, and mixed sub-batch processing sequences. Suppose that the encoding structures extracted from the set of sub-batch quantities, the set of sub-block sizes, and the set of mixed sub-batch processing sequences are represented by SNS, SSS, and ISS, respectively. The specific designs of these four evolutionary operators are as follows:
[0075] EO1: The ISS is aligned with the current mixed sub-batch processing sequence to identify reference sub-batches: any sub-batch whose job index and operation index match the ISS is considered a reference sub-batch. If no reference sub-batch exists in the sequence, an intermediate sub-batch is selected as the reference. Subsequently, for every interval between two adjacent reference sub-batches, the first sub-batch in that interval is moved to the end, and the remaining sub-batches are moved forward one position in sequence; if the interval contains only one sub-batch, it remains unchanged. Intervals before the first reference sub-batch and after the last reference sub-batch are treated in the same way. This exchange operation shuffles the processing order of operations. Finally, a repair operation ensures that all sequences and constraints are satisfied. An example of EO1 is shown below. Figure 2 As shown.
[0076] EO2: Randomly select two indices a and b, where a∈{1,…,N} represents the index of a certain workpiece in the current solution, and b∈{1,…,M} represents the index of a certain operation in workpiece a; after determining a and b, replace the corresponding fragments of operation b in Layer1 and Layer2 of operation a with the corresponding fragments in SNS and SSS, respectively.
[0077] EO3: Randomly select a workpiece index a∈{1,…,N} from the current individual; then, replace the corresponding fragments of all processes of the selected workpiece a in Layer1 and Layer2 with the corresponding fragments in SNS and SSS;
[0078] EO4: Randomly select a process index b∈{1,…,M} from the current individual, and then replace the corresponding fragments of all workpieces in Layer1 and Layer2 with the corresponding fragments in SNS and SSS.
[0079] In KMCEA, an Evolutionary Operator Selection Strategy (EOSS) is proposed to address the issue that four evolutionary operators have varying impacts at different iteration stages. EOSS converts the historical success and failure rates of each evolutionary operator, as well as the success and failure rates in the current iteration, into a quantitative knowledge score, as shown in Equation (4). The knowledge score reflects the overall utility of each operator and is normalized using Equation (5) to form a probability distribution. Subsequently, a roulette wheel method is used to select the most promising operator. EOSS can dynamically adjust the selection probability based on the operator's performance, thereby enhancing the algorithm's ability to develop effective search directions. The pseudocode for EOSS is shown in Algorithm 4.
[0080] in, and They represent evolution operators respectively. The historical number of successes and failures when generating nondominated solutions in previous iterations; and They represent evolution operators respectively. The historical number of successes and failures when generating non-dominated solutions in the current iteration; 'a' is the success factor, used to increase the selection probability of the operator; 'b' is the failure factor, used to decrease the selection probability of the operator; 'ε' is used to prevent the operator weight from becoming 0, thus ensuring that all operators have a chance to be selected in the initial stage of the algorithm. This represents the knowledge value of the x-th evolution operator; The normalized knowledge value represents the selection probability of the x-th evolutionary operator. When an evolutionary operator produces a new solution, its corresponding knowledge value and selection probability will be updated.
[0081]
[0082] 3. Local reinforcement module
[0083] Local search operators based on problem characteristics can significantly reduce the randomness of local searches and improve search efficiency. In the LSJSP-VI problem, optimization performance is significantly affected by both sub-batch partitioning and the staggered processing sequence of sub-batches. The maximum completion time is determined by the critical path, which is defined as the longest path from the start of processing of the first sub-batch to the completion of processing of the last sub-batch. A critical block refers to adjacent nodes on the same machine on the critical path, which are partitioned into the same critical block. A critical sub-batch is a sub-batch located on the critical path. In the critical path, the sub-batch with the longest processing time is defined as the urgent sub-batch. Therefore, based on the characteristics of sub-batch adjustment, critical path, and critical block, six local search operators are designed to reduce TEC and TTD. These six local search operators are denoted as LSOi, i∈[1,6], as follows.
[0084] LSO1: Randomly select a key sub-batch and map it to the corresponding sub-batch in Layer 2. Let z represent the current batch size of this sub-batch. This represents the total batch size of the job, with the batch size of sub-batches perturbed by h, where... ;
[0085] LSO2: Select the critical sub-batch with the largest batch size and divide it into two equal-sized sub-batches.
[0086] LSO3: Exchanges an urgent sub-batch with its preceding sub-batch on the same machine, thereby reducing its dwell time in the system and further shortening the waiting time of subsequent sub-batch.
[0087] LSO4: On the critical path, split the longest delayed sub-batch and insert it into an earlier available time period to reduce TTD;
[0088] LSO5: N6 is a robust neighborhood structure designed for the job shop scheduling problem. Its goal is to optimize scheduling by changing the order of operations in the critical block. N6 reduces completion time through the following steps: In the critical path, randomly select a non-last operation from the leading segment π1 and move it to the end of π1; randomly select a pair of operations in the middle segment π2, move one to the beginning of π2 and the other to the end of π2; randomly select a non-first operation from the tail segment πn and move it to the beginning of πn.
[0089] LSO6 is an improvement on the Deadline Neighbor Relation (DNR) method. DNR refers to the situation where two adjacent sub-batches S1 and S2 constitute a critical sub-batch pair on the same machine: when the completion time of S1 is earlier than its delivery time, while the completion time of S2 is later than its delivery time; if identified as a critical sub-batch pair, the positions of these two sub-batches are swapped to reduce TTD.
[0090] 6. An energy-saving workshop scheduling system for batch flow operations according to claim 5, characterized in that it further includes a Q-learning operator selection mechanism, which uses Q-learning to update the solutions in the Pareto optimal set, constructs operator selection strategies for different solutions through Q-learning, and uses a 4×3 index Q-table to assist in selecting appropriate local search operators. The update of the Q-table is shown in formula (6).
[0091]
[0092] Where reward and α represent the action reward and learning rate, respectively. and These correspond to the state and action of the new solution, respectively; then, the following Q-learning mechanism is used to select appropriate operators for different solutions;
[0093] This mechanism uses the solutions in each generation of the Pareto optimal set as the agent. By defining the state space, reward function, and action set, and combining the ε-greedy policy, it achieves adaptive selection of the operator. Let π represent the current solution and π* represent the newly generated solution; the specific definitions are as follows:
[0094] Agent: The solution in the Pareto optimal set of each generation serves as the agent;
[0095] State: The state is defined by comparing the changes of π and π* on TEC and TTD, so as to help the agent adapt to different optimization situations: state 0 represents TEC(π)<TEC(π*) and TTD(π)<TTD(π*); state 1 represents TEC(π)<TEC(π*) and TTD(π)>TTD(π*); state 2 represents TEC(π)>TEC(π*) and TTD(π)<TTD(π*); state 4 represents TEC(π)>TEC(π*) and TTD(π)>TTD(π*);
[0096] Reward: The reward value is proportional to the improvement amplitude; therefore, the reward value is calculated by comparing the improvement of solutions π and π*, as shown in formula (7); this formula measures the relative reduction of TEC and TTD, and selects the smaller improvement value as the final reward to ensure stability;
[0097]
[0098] Action: Local search operations are regarded as an action, and multiple local search operators are combined to form an action; an action is selected from the action set as a combined local search operator to optimize the objective value. According to the design, 3 actions are defined in total, forming the action set: Action={A1,A2,A3}. A1 includes LSO1 and LSO2, which are used for the sublot adjustment strategy; A2 includes LSO3 and LSO4, which target sublots on the critical path; A3 includes LSO5 and LSO6, which are used for critical block operations;
[0099] In state s, the ε-greedy strategy is used to select an action, as shown in formula (8):
[0100]
[0101] Wherein, a random number rand is generated. When rand<ε, the greedy strategy is adopted in the k-th row of the Q-table to select the action with the maximum Q value; otherwise, an action is selected randomly. Once an action is selected, π will execute all search operators in the action, and select the optimal operator for further search through the greedy strategy. This greedy strategy will select the local search operator that can maximize the improvement degree of the objective function. The pseudo-code of local reinforcement is shown in Algorithm 5.
[0102]
[0103] 4. System Flow
[0104] First, KMCEA uses a knowledge-driven initialization method based on encoding structure to generate the initial population. Then, it saves the non-dominated solution set and records the corresponding optimal encoding structure. Second, during the co-evolutionary search process, the population undergoes one generation of evolution through evolutionary operators. An evolutionary operator selection strategy is proposed, which integrates feedback information from the history and current performance of evolutionary operators and transforms this information into quantitative knowledge values. Based on these knowledge values, the operator selection probability is dynamically adjusted to match the solution with better evolutionary operators with different characteristics. Furthermore, KMCEA employs Q-learning to select the optimal search operator for non-dominated solutions and finds more potential non-dominated solutions locally. KMCEA iterates until the termination condition is met, finally outputting the Pareto optimal set as the final scheduling solution for the LSJSP-VI problem. The flowchart of KMCEA is as follows. Figure 8 As shown.
[0105] The following numerical example illustrates the calculation method of TEC and TTD in LSJSP-VI. Figure 1 This demonstrates an LSJSP-VI instance containing three machines and two jobs. Each job contains three operations to be processed. Taking operation 1 of job1 as an example, its setup time on machine M3 is 3 seconds, the processing time for a batch is 3 seconds, and the batch size of job1 is 7. The time to delivery and completion (TTD) are determined by the delivery date and completion time: the delay time for job1 is TTD1 = max{0, 75-94.5} = 0; the delay time for job2 is TTD2 = max{0, 106-150} = 0, and the TTD is TTD1 + TTD2 = 0 + 0 = 0.
[0106] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.
Claims
1. An energy-saving workshop scheduling system for batch workflow operations, characterized in that, It includes a population initialization module, a co-evolutionary search module, and a local reinforcement module; The population initialization module is used to generate an initial population with high quality and diversity; wherein, the scheduling solution is represented by three vectors, namely Layer1, Layer2 and Layer3; Layer1 specifies the number of sub-batches into which each operation of the job is divided, Layer2 defines the size of each sub-batch, and Layer3 describes the mixed sub-batch processing sequence on the machine; The co-evolutionary search module is used to promote the effective evolution of different sub-problems and the efficient collaboration between sub-problems, including an information storage mechanism and an evolutionary operator selection strategy; The local enhancement module is a local search based on problem characteristics. According to the characteristics of sub-batch adjustment, critical path and critical block, six local search operators are designed to reduce TEC and TTD.
2. The energy-saving workshop scheduling system for batch workflow operations according to claim 1, characterized in that, The population initialization module includes: The first initialization method is used to determine the number of sub-batches for each process in each workpiece, where the number of sub-batches is limited to a predefined range. Within this interval, each candidate value is used as a corresponding weight and normalized to obtain a probability distribution. A roulette wheel selection mechanism is constructed based on the cumulative probability, and the number of sub-batches is determined by random numbers. A second initialization method is used to determine the size of each sub-batch, and a method that simultaneously includes uniform and non-uniform sub-batch sizes is designed; a uniform random perturbation factor is introduced. To control the division of the batch size of the sub-batch, and the batch size of each sub-batch is determined by formula (1), which provides flexibility in adjusting the size of the subtask; in order to further ensure the feasibility of the solution, a repair mechanism is applied during the adjustment process, and the batch size that has reached the minimum size of 1 cannot be further reduced, thereby avoiding the generation of invalid (zero or negative) batch sizes. in, Let i be the batch size of workpiece i. Let i be the number of sub-batches of workpiece i in process j, and This represents the batch size of each sub-batch; The third initialization method is used to determine the mixed sub-batch processing sequence on all machines. This initialization method clarifies the relationship between processing and delivery, enabling scheduling based on the urgency of delivery, thereby promoting on-time delivery. The urgency of delivery depends on the delivery time, the number of jobs, and the processing time. Scheduling rules based on the urgency of delivery are developed. First, jobs are sorted according to importance and deadline, generating priority index and deadline index. These selection probabilities are combined using an addition operator to form a comprehensive probability, which is used to arrange the sub-batches of each job in order. Job priority and sub-batch priority are calculated by formula (2) and formula (3) respectively to accurately determine the job priority and the urgency of sub-batch delivery. in, For the batch size of each workpiece, The batch size for each sub-batch, Let be the processing time of process j for workpiece i on machine k. Let be the deadline for workpiece i. and Processing priority.
3. The energy-saving workshop scheduling system for batch workflow operations according to claim 2, characterized in that, The information storage mechanism is used to obtain accurate sub-block partitioning and mixed sub-batch processing sequences. This mechanism constructs three sets: a sub-batch quantity set, a sub-batch size set, and a mixed sub-batch processing sequence set. These sets store the sub-batch quantity, sub-block size, and mixed sub-batch processing sequence of the non-dominated solution set. After each iteration, the information storage mechanism selects appropriate sub-batch quantity, sub-batch size, and mixed sub-batch processing sequence from the non-dominated solution set based on target value comparisons, ensuring the capture of high-quality solutions. Four evolutionary operators, EO1, EO2, EO3, and EO4, are designed for a specific problem to determine and identify appropriate sub-batch quantity, sub-batch size, and mixed sub-batch processing sequence. The encoding structures extracted from the sub-batch quantity set, sub-block size set, and mixed sub-batch processing sequence set are denoted by SNS, SSS, and ISS, respectively. The specific design of these four evolutionary operators is described below: EO1: The ISS is aligned with the current mixed sub-batch processing sequence to identify the reference sub-batch: any sub-batch whose job index and operation index match the ISS is considered the reference sub-batch; If there is no reference sub-batch in the sequence, the middle sub-batch is selected as the reference. Then, for each interval between two adjacent reference sub-batches, the first sub-batch in the interval is moved to the end, and the remaining sub-batches are moved forward one position in turn. If the interval contains only one sub-batch, it remains unchanged; the intervals before the first reference sub-batch and after the last reference sub-batch are also treated in the same way. This exchange operation will disrupt the processing order of the steps; finally, a repair operation is used to ensure that all sequences and constraints are satisfied. EO2: Randomly select two indices a and b, where a∈{1,…,N} represents the index of a certain workpiece in the current solution, and b∈{1,…,M} represents the index of a certain operation in workpiece a; after determining a and b, replace the corresponding fragments of operation b in Layer1 and Layer2 of operation a with the corresponding fragments in SNS and SSS, respectively. EO3: Randomly select a workpiece index a∈{1,…,N} from the current individual; Subsequently, replace the corresponding segments of all processes of the selected workpiece a in Layer1 and Layer2 with the corresponding segments in SNS and SSS; EO4: Randomly select a process index b∈{1,…,M} from the current individual, and then replace the corresponding fragments of all workpieces in Layer1 and Layer2 with the corresponding fragments in SNS and SSS.
4. The energy-saving workshop scheduling system for batch workflow operations according to claim 3, characterized in that, The Evolutionary Operator Selection Strategy (EOSS) is used to address the problem that the four evolutionary operators have different effects at different iteration stages. EEOSS converts the historical success rate and failure rate of each evolutionary operator as well as the success rate and failure rate in the current iteration into a quantitative knowledge score, as shown in formula (4). The knowledge score reflects the overall utility of each operator and is normalized by formula (5) to form a probability distribution. Then, the roulette wheel method is used to select the most promising operator. EOSS can dynamically adjust the selection probability of the operator according to its performance, thereby enhancing the algorithm's ability to develop in effective search directions. in, and They represent evolution operators respectively. The historical number of successes and failures when generating nondominated solutions in previous iterations; and They represent evolution operators respectively. The historical number of successes and failures when generating non-dominated solutions in the current iteration; 'a' is the success factor, used to increase the selection probability of the operator; 'b' is the failure factor, used to decrease the selection probability of the operator; 'ε' is used to prevent the operator weight from becoming 0, thus ensuring that all operators have a chance to be selected in the initial stage of the algorithm. This represents the knowledge value of the x-th evolution operator; The normalized knowledge value represents the selection probability of the x-th evolutionary operator. When an evolutionary operator produces a new solution, its corresponding knowledge value and selection probability will be updated.
5. An energy-saving workshop scheduling system for batch workflow operations according to claim 4, characterized in that, The six local search operators are denoted as LSOi, i∈[1,6], and are as follows: LSO1: Randomly select a key sub-batch and map it to the corresponding sub-batch in Layer 2. Let z represent the current batch size of this sub-batch. This represents the total batch size of the job, with the batch size of sub-batches perturbed by h, where... ; LSO2: Select the critical sub-batch with the largest batch size and divide it into two equal-sized sub-batches. LSO3: Exchanges an urgent sub-batch with its preceding sub-batch on the same machine, thereby reducing its dwell time in the system and further shortening the waiting time of subsequent sub-batch. LSO4: On the critical path, split the longest delayed sub-batch and insert it into an earlier available time period to reduce TTD; LSO5: N6 is a robust neighborhood structure designed for the job shop scheduling problem. Its goal is to optimize scheduling by changing the order of operations in the critical block. N6 reduces completion time through the following steps: In the critical path, randomly select a non-last operation from the leading segment π1 and move it to the end of π1; randomly select a pair of operations in the middle segment π2, move one to the beginning of π2 and the other to the end of π2; randomly select a non-first operation from the tail segment πn and move it to the beginning of πn. LSO6: an improvement to the Deadline Neighbor Relation (DNR) method, where DNR refers to a situation where two adjacent sublots S1 and S2 on the same machine form a critical sublot pair: when the completion time of S1 is earlier than its due date, while the completion time of S2 is later than its due date; if the sublots are identified as a critical pair, the positions of these two sublots are swapped to reduce TTD.
6. An energy-saving workshop scheduling system for batch workflow operations according to claim 5, characterized in that, It also includes a Q-learning operator selection mechanism, which uses Q-learning to update solutions in the Pareto optimal set, constructs operator selection strategies for different solutions through Q-learning, and uses a 4×3 index Q-table to assist in selecting appropriate local search operators. The update of the Q-table is shown in formula (6), Where reward and α represent the action reward and learning rate, respectively. and These correspond to the state and action of the new solution, respectively; then, the following Q-learning mechanism is used to select appropriate operators for different solutions; This mechanism takes solutions in each generation of the Pareto optimal set as agents, realizes adaptive selection of operators by defining the state space, reward function and action set, combined with the ε-greedy strategy. Let π represent the current solution, and π* represent the newly generated solution; the specific definitions are as follows: Agent: solutions in each generation of the Pareto optimal set act as agents; State: The state is defined by comparing the changes of TEC and TTD between π and π*, so as to help the agent adapt to different optimization scenarios: State 0 represents TEC(π)<TEC(π*) and TTD(π)<TTD(π*); State 1 represents TEC(π)<TEC(π*) and TTD(π)>TTD(π*); State 2 represents TEC(π)>TEC(π*) and TTD(π)<TTD(π*); State 4 represents TEC(π)>TEC(π*) and TTD(π)>TTD(π*); Reward: The reward value is proportional to the improvement amplitude; therefore, the reward value is calculated by comparing the improvement of solution π and π*, as shown in formula (7); this formula measures the relative reduction of TEC and TTD, and selects the smaller improvement value as the final reward to ensure stability; Action: Local search operations are regarded as an action, and multiple local search operators are combined to form an action; an action is selected from the action set as a combined local search operator to optimize the objective value; according to the design, 3 actions are defined in total, forming the action set: Action={A1,A2,A3}. A1 includes LSO1 and LSO2, which are used for sublot adjustment strategies; A2 includes LSO3 and LSO4, targeting sublots on the critical path; A3 includes LSO5 and LSO6, which are used for critical block operations; Under state s, the ε-greedy strategy is used to select an action, as shown in formula (8): Wherein, a random number rand is generated. When rand<ε, the greedy strategy is used to select the action with the maximum Q value in the k-th row of the Q-table; otherwise, an action is selected randomly. Once an action is selected, π will execute all search operators in the action, and select the optimal operator for further search through the greedy strategy.