Knowledge and data hybrid-driven intelligent scheduling method under random arrival of batch

By introducing a knowledge and data-driven intelligent scheduling method with mixed batch arrival in the scheduling system, and using dual-deep Q network optimization scheduling strategy, the problems of untimely scheduling response and insufficient policy adaptability in dynamic batch environments are solved, significantly improving scheduling performance and flexibility.

CN120218566AActive Publication Date: 2025-06-27LIAOCHENG UNIV
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202510671789.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-06-27
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

Existing scheduling methods respond in dynamic batch environments to problems such as untimely scheduling response, lack of effective knowledge guidance and insufficient strategy adaptability.

Method used

The intelligent scheduling method driven by knowledge and data is adopted under batch random arrival. By defining the policy network, state space and action space, and using a dual-deep Q network for training, combining batch operation selection strategies, greedy batch operation splitting strategies, machine allocation strategies, and position-based scheduling strategies, optimized scheduling solutions.

Benefits of technology

It improves the flexibility, adaptability and overall scheduling performance of the scheduling system, and can more effectively deal with random arrival problems in dynamic batch environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218566A_ABST
    Figure CN120218566A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of hybrid flow shop scheduling, in particular to a knowledge and data hybrid-driven intelligent scheduling method under random arrival of batches. The method comprises the following steps: initializing batches in a to-be-scheduled batch set into batch operations, selecting one batch operation by utilizing a batch operation selection strategy, performing processing sequence adjustment on the batch operations on a selected machine based on a position insertion strategy, and removing processed batch operations from an unscheduled batch operation set and an available batch operation set. A strategy based on key sub-batch operation is executed on key sub-batch operation, randomly arriving batches are converted into batch operation to be added into an unscheduled batch operation set, a strategy network is trained based on the operation process by using a double-depth Q network, and an intelligent scheduling scheme under random arrival of the batches is obtained by using the trained strategy network. And terminating the scheduling. The workshop scheduling method has the positive effects of realizing high efficiency, intelligence and flexibility of workshop scheduling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of hybrid flow shop scheduling, and in particular, belongs to an intelligent scheduling method driven by a hybrid of knowledge and data under random arrival of batches. Background Art

[0002] With the increasing of personalized customer demands, enterprises must quickly respond to market demands. In this context, the arrival of new batches is highly random, and these batches can be converted into one or more batches. To adapt to the dynamic environmental changes brought by these new batches, the scheduling must be adjusted in real time according to their arrival times. However, the problem of random arrival of batches has not been fully solved in existing research. Therefore, there is an urgent need to explore flexible and advanced scheduling methods to ensure that new batches can be inserted into the production plan in a timely manner while minimizing the delay of all batches.

[0003] Currently, the solution methods for scheduling problems can be mainly divided into three categories: mathematical programming models, heuristic algorithms, and meta-heuristic algorithms. Although mathematical programming models can obtain optimal solutions, the computational complexity of mathematical programming models will increase sharply with the growth of the problem scale, resulting in a significant extension of the solution time and an increase in resource consumption, which is unacceptable for a dynamic environment that requires real-time scheduling. Although meta-heuristic algorithms are more efficient than mathematical programming in solving large-scale problems, it may still take a long time to find the desired solution, making it difficult to meet the real-time scheduling requirements. Heuristic algorithms provide an effective alternative. Heuristic algorithms emphasize quick decision-making. Such algorithms utilize domain knowledge and heuristic rules and can quickly generate high-quality scheduling solutions, thus being particularly suitable for dynamic scheduling environments. However, heuristic algorithms usually rely on static production data, and with the increase of the problem scale, the performance of heuristic algorithms may decline. In addition, heuristic algorithms are often designed for specific scenarios and are difficult to adapt to the changes in dynamic environments.

[0004] In summary, there are still certain limitations in the research on intelligent scheduling driven by a hybrid of knowledge and data under random arrival of batches. To better meet the actual needs of real production scenarios, it is necessary to conduct in-depth research on the intelligent scheduling problem driven by a hybrid of knowledge and data under random arrival of batches to find more practical solutions. Summary of the Invention

[0005] The present invention provides an intelligent scheduling method driven by a hybrid of knowledge and data under random arrival of batches, which solves the technical problems of the existing scheduling methods such as untimely scheduling response, lack of effective knowledge guidance, and insufficient strategy adaptability in dealing with dynamic batch environments, so as to achieve the purpose of improving the flexibility, adaptive ability, and overall scheduling performance of the scheduling system.

[0006] The intelligent scheduling method driven by the hybrid of knowledge and data under the random arrival of batches provided by the present invention is characterized by including the following steps: Step 1, define a policy network, a state space, and an action space, and control the policy network to select an action according to the current state. Each action corresponds to one of the rule combinations of the batch operation selection policy and the machine allocation policy. Step 2, obtain the set of batches to be scheduled at the initial moment, and initialize the batches in the set of batches to be scheduled as batch operations. Batch operations refer to the processing operations performed by batches at each production stage. The batch operations of all production stages are initialized as an unscheduled batch operation set, and the batch operations of the first production stage are initialized as an available batch operation set. Step 3, select a batch operation from the available batch operation set using the batch operation selection policy. If the selected batch operation is a batch operation of the first production stage, it means that the batch has not been split yet. Execute the greedy-based batch operation splitting policy to split the batch operation into several sub-batch operations, and then execute the machine allocation policy to allocate a machine for the selected batch operation and perform processing on the selected machine. If the selected batch operation is not a batch operation of the first production stage, use the splitting result of the batch operation of the first production stage, execute the machine allocation policy to allocate a machine for the selected batch operation, and perform processing on the selected machine. Step 4, execute the position-based insertion policy to adjust the processing order of the batch operations on the selected machine, and the selected machine completes the processing according to the adjusted batch operation order. Step 5, remove the batch operations that have completed processing from the unscheduled batch operation set and the available batch operation set. If the removed batch operation is not a batch operation of the last production stage, add the batch operation of the next production stage of the removed batch operation to the available batch operation set. Step 6, after the batch operations of all production stages have completed processing, find the critical sub-batch operation from all sub-batch operations. The critical sub-batch operation refers to the sub-batch operation with the least total processing time from the first production stage to the last production stage. Execute the policy based on the critical sub-batch operation to update the entire processing process. Step 7, if there are batches that arrive randomly from the start to the current moment, convert the randomly arrived batches into batch operations and add them to the unscheduled batch operation set, and add the batch operations of the first production stage to the available batch operation set. Step 8, when the unscheduled batch operation set is empty, the batch operations of all batches are scheduled. Step 9, use the double deep Q network to train the policy network based on Steps 2 - 8, obtain the intelligent scheduling scheme under the random arrival of batches using the trained policy network, and terminate the scheduling.

[0007] Furthermore, the state space is composed of the general attributes of the overall scheduling state of the workshop and the specific attributes of each production stage. The general attributes have 9 dimensions, and the specific attributes of each production stage have 18 dimensions. That is, if the workshop contains m production stages, the total dimension of the state space is 9 + 18 × m; The action space corresponds to all rule combinations of the batch operation selection strategy and the machine allocation strategy. There are eight rules for the batch operation selection strategy and five rules for the machine allocation strategy. There are 40 rule combinations of the batch operation selection strategy and the machine allocation strategy, which are used as the output of the policy network.

[0008] Further, the batch operation selection strategy includes, The earliest delivery date priority rule gives priority to batch operations with the earliest delivery date for scheduling; The minimum slack time priority rule gives priority to scheduling the batch operation with the minimum slack time. The slack time is defined as the batch delivery time minus the remaining processing time. The longest processing time priority rule gives priority to the batch operations with the longest processing time for scheduling; The shortest processing time priority rule gives priority to batch operations with the shortest processing time; The shortest remaining processing time priority rule gives priority to the job with the shortest remaining processing time. The remaining processing time refers to the total time required for subsequent production stages after the current batch operation is completed; The longest remaining processing time priority rule gives priority to the batch operations with the longest remaining processing time; The earliest start time priority rule sorts batch operations according to their earliest start time, and schedules batch operations with the highest order first. The latest start time priority rule sorts batch operations according to their latest start time, and prioritizes batch operations with higher order.

[0009] Further, the machine allocation strategy includes, The shortest waiting time priority rule gives priority to the machine that has the shortest waiting time for the current batch operation; The load balancing priority rule gives priority to the machine with the smallest current load. The load is defined as the number of batch operations assigned to the machine. The earliest idle time priority rule gives priority to the machine that becomes idle the earliest after completing the currently assigned batch operation; The longest idle time priority rule: when there are idle machines, the machine with the longest idle time is given priority; The earliest completion time is given priority. The machine with the earliest completion time is selected first, and the current batch operation is assigned to the machine that completes the current batch operation the earliest.

[0010] Furthermore, the greedy batch operation splitting strategy includes the following operation process. Obtain the size of the batch operation to be split, set the maximum number of sub-batch operations that can be split, set the initial size of each sub-batch operation to zero, and at the same time set the remaining batch operation size to be allocated. Enter a loop, select a unit 1 from the remaining unallocated batch operation size, and sequentially allocate the selected unit 1 to each sub-batch operation. Calculate the processing time corresponding to each sub-batch operation, compare the processing times of all sub-batch operations after allocating the selected unit 1, select the sub-batch operation with the minimum processing time as the final allocation object of the unit 1, update the size of the sub-batch operation, subtract 1 from the total remaining unallocated batch operation amount, and then perform the next loop until the remaining unallocated total is 0, completing the greedy splitting process of the batch operation. Furthermore, the position-based insertion strategy includes the following operation process.

[0011] Define parameters: represents the batch set, represents the total number of batches, , represents a certain batch; represents the set of production stages, represents the total number of production stages, , represents a certain production stage; represents the batch operation of batch ; Determine the positions on the machine that can be used as insertion points according to the selected batch operation, the selected machine, and the batches that the machine has currently completed scheduling. Sequentially traverse each insertable position and judge whether inserting the selected batch operation into the current position meets the feasibility conditions. The feasibility conditions are shown in Formulas 1 and 2. Formula 1 Wherein, represents the length of the idle time period of the current machine, represents the size of the batch operation ; represents the unit processing time of the batch operation in the production stage ; represents the machine setup time of the batch operation in the production stage ; Formula 1 ensures that the idle time of the selected machine at a certain position can accommodate the minimum processing time required for the selected batch operation. Meeting Formula 1 means that the idle time of the selected machine at a certain position can accommodate the selected batch operation. ​ Formula 2 Wherein, represents the latest time of the machine idle time; represents the maximum number of sub - batch operations; represents the batch operation of the maximum number of sub - batch operations at the end time in the production stage ; represents the time when a certain production stage transfers to the next production stage; represents the batch operation of the maximum number of sub - batch operations size; Formula 2 ensures that the latest time point of the idle time of the selected machine at a certain position can meet the processing requirements of the last sub - batch operation of the selected batch operation. Meeting Formula 2 means that moving the selected batch operation to a certain position of the selected machine will not affect the arrangement of subsequent batch operations; If the insertion position meets the above feasibility conditions, then calculate the scheduling time of each sub - batch operation of the selected batch operation. Select the larger value between the sum of the end time of the previous batch operation after completion in the current production stage and the machine setup time, and the sum of the end time of the previous batch operation after processing completion in the previous production stage and the transfer time as the start time of the first sub - batch. Select the larger value between the end time of the previous sub - batch operation of the selected batch operation and the sum of the end time of the sub - batch operation in the previous production stage and the transfer time as the start time of the subsequent sub - batches. The end time of each sub - batch operation is the sum of the start time and the processing time of the sub - batch operation. The processing time of the sub - batch operation is the unit processing time multiplied by the size of the sub - batch operation. After calculating the time of all sub - batch operations, judge whether the insertion of the selected batch operation causes delay to the subsequent sub - batch operations on the selected machine. If there is no delay, insert the selected batch operation into the current position and stop the position - based insertion strategy. Otherwise, skip the current position and continue to try the next insertion point until all positions are traversed or a feasible insertion position is found.

[0012] Furthermore, the strategy based on the critical sub - batch operation includes the following operation process, Calculate the reduced size of the critical sub - batch operation in each production stage except the first production stage. Traverse all production stages except the first production stage and calculate the reduced size of the sub - batch operation of the batch operation in the production stage of the batch operation , , The calculation method is as follows, Wherein, the symbol represents rounding down; Representing batch operations Sub-batch operations At the production stage Start time; Representing batch operations Sub-batch operations At the production stage End time; Select the minimum value from the reduced sizes of all production stages calculated, remove the selected minimum reduced size from the critical sub-batch operations, merge it into the previous sub-batch operation of the critical sub-batch operations, complete the adjustment of the sizes of the two sub-batch operations, and update the scheduling process of the two sub-batch operations at each production stage.

[0013] Furthermore, the method for training the policy network using the double deep Q-network based on Steps 2 - 8 is as follows. The double deep Q-network adopts a double-network structure, including a policy network and a target network, which are used for action selection and Q-value evaluation respectively. The double deep Q-network combines the experience replay mechanism for stable training. The specific process is as follows. Construct a training dataset, which contains several scheduling instances with a fixed number of production stages, and set the maximum number of training rounds. At the same time, initialize the current training round to the first round. Initialize two neural networks, namely the policy network and the target network, and establish an experience replay buffer. At the beginning of each round of training, randomly select an instance from the training dataset for training. During the training process, judge whether the maximum number of training rounds has been reached. If not, continue to execute this round of training. If so, output the trained policy network. In each round of training, the policy network obtains the current shop floor state and checks whether the current scheduling task is completed. If it is completed, increment the current training round by one and start the next round. Otherwise, continue to execute the subsequent operations. The policy network adopts an ε-greedy strategy to select actions, randomly selects an action with probability ε; selects the current optimal action with the remaining probability 1 - ε. After executing the action, the policy network obtains a new state and the corresponding reward value. The current state, the selected action, the reward value, and the next state are used as an experience data and stored in the experience replay buffer. The policy network randomly samples a small batch of experience data from the experience replay buffer and calculates the target Q-value y for each experience data. The calculation method is as follows.

[0014] Among them, r represents the reward value obtained after executing action a in the current state s; γ is the discount factor, which is used to measure the importance of future rewards; s′ is the next state reached after executing the action; Represents the action value estimated by the target network for state s′ and all actions a′; Represents the action value with the largest estimated value among all actions in state s'. Calculate the loss between the Q value obtained by the policy network and the target Q value, and perform gradient calculation on the parameters in the policy network through the backpropagation algorithm to update the parameters in the policy network. After reaching the preset number of training rounds, synchronize the parameters of the target network to the parameters of the policy network. The training process is repeated continuously until the set maximum number of training rounds is reached, and output the trained policy network. Furthermore, the current state is input into the trained policy network to obtain the Q values corresponding to all actions. An action refers to the rule combination of batch operation selection strategy and machine allocation strategy. Select the rule combination with the largest Q value as the current action, that is, determine the rule combination of batch operation selection strategy and machine allocation strategy in the entire scheduling process, execute the scheduling process to obtain a scheduling plan, and complete the scheduling.

[0015] The present invention proposes an intelligent scheduling method driven by a hybrid of knowledge and data under random batch arrivals, constructs a configurable constructive integrated solution framework, aims to address the problems of uncertain batch arrival times and complex resource allocation in actual production. The configurable constructive integrated solution framework integrates multiple scheduling strategies, including batch operation selection strategy, greedy batch operation splitting strategy, machine allocation strategy, position-based insertion scheduling strategy, and key sub-batch operation-based optimization strategy. On this basis, the present invention further introduces a double deep Q network as a reinforcement learning model, and iteratively trains the policy network through the continuous interaction between the policy network and the scheduling environment. The double deep Q network can effectively handle the problem of optimal action selection in a high-dimensional state space and avoid the phenomenon of overestimation of Q values that may occur in traditional Q learning. During the training process, the policy network selects an action combination with a higher Q value according to the current state, thereby promoting the entire scheduling process to evolve towards a better direction. As the training progresses, the policy network can gradually learn a more efficient and flexible scheduling method, significantly improving the scheduling performance and response ability of the hybrid flow shop under dynamic batch scenarios. The configurable constructive integrated solution framework introduced with the double deep Q network combines a hybrid intelligent optimization mechanism of knowledge and data, has good generalization and adaptability, is applicable to batch scheduling optimization problems in a variety of complex manufacturing environments, and has the positive effects of realizing efficient, intelligent, and flexible shop scheduling. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 Is the implementation flowchart of the present invention; Figure 2 Is the flowchart of training using the double deep Q network of the present invention; Figure 3This is the data distribution diagram obtained from the experiments of the present invention and SPTCH_UBLSS, NEHCH_UBLSS, MCH_UBLSS, and JCH_UBLSS; Figure 4 This is the mean and confidence interval diagram of the data obtained from the experiments of the present invention and SPTCH_UBLSS, NEHCH_UBLSS, MCH_UBLSS, and JCH_UBLSS; Figure 5 This is the batch interaction diagram obtained from the experiments of the present invention and SPTCH_UBLSS, NEHCH_UBLSS, MCH_UBLSS, and JCH_UBLSS; Figure 6 This is the stage interaction diagram obtained from the experiments of the present invention and SPTCH_UBLSS, NEHCH_UBLSS, MCH_UBLSS, and JCH_UBLSS; Figure 7 This is the data distribution diagram obtained from the experiments of the present invention and CVND, GAS2C1, and GAS2C2; Figure 8 This is the mean and confidence interval diagram of the data obtained from the experiments of the present invention and CVND, GAS2C1, and GAS2C2. Detailed implementation manners

[0017] As Figure 1 - Figure 2 shown, the intelligent scheduling method driven by the mixture of knowledge and data under the random arrival of batches provided by the present invention is mainly implemented through the following steps.

[0018] Step 1: Define the policy network, state space, and action space. According to the current state, control the policy network to select an action, and each action corresponds to one of the rule combinations of the batch operation selection policy and the machine allocation policy. Specifically, the state space is composed of the general attributes of the overall workshop scheduling state and the specific attributes of each production stage. The general attributes are 9-dimensional in total, and the specific attribute of each production stage is 18-dimensional. That is, if the workshop contains m production stages, the total dimension of the state space is 9 + 18×m. The action space corresponds to all the rule combinations of the batch operation selection policy and the machine allocation policy. There are eight rules for the batch operation selection policy and five rules for the machine allocation policy. There are 40 rule combinations between the batch operation selection policy and the machine allocation policy, which are used as the output of the policy network. Among them, the rules of the batch operation selection policy and the machine allocation policy are as follows.

[0019] The batch operation selection policy includes: The earliest due date first rule, which preferentially selects the batch operation with the earliest due date for scheduling; The earliest due date (EDD) rule gives priority to scheduling the batch operation with the minimum slack time. The slack time is defined as the due date of the batch minus the remaining processing time. The longest processing time (LPT) rule gives priority to scheduling the batch operation with the longest processing time. The shortest processing time (SPT) rule gives priority to the batch operation with the shortest processing time. The shortest remaining processing time (SRPT) rule gives priority to the job with the shortest remaining processing time. The remaining processing time refers to the total processing time required in the subsequent production stages after the current batch operation is completed. The longest remaining processing time (LRPT) rule gives priority to scheduling the batch operation with the longest remaining processing time. The earliest start time (EST) rule sorts according to the earliest start time of each batch operation and gives priority to scheduling the batch operations ranked higher. The latest start time (LST) rule sorts according to the latest start time of each batch operation and gives priority to scheduling the batch operations ranked higher.

[0020] Machine allocation strategies include: The shortest waiting time (SWT) rule gives priority to selecting the machine with the shortest waiting time for the current batch operation. The load balancing (LB) rule gives priority to selecting the machine with the minimum current load. The load is defined as the number of batch operations assigned to the machine. The earliest available time (EAT) rule gives priority to selecting the machine that will be available earliest after completing the currently assigned batch operations. The longest available time (LAT) rule, when there are available machines, gives priority to selecting the machine with the longest available time. The earliest completion time (ECT) rule gives priority to selecting the machine with the earliest completion time and assigns the current batch operation to the machine that will complete the current batch operation earliest.

[0021] Step 2: Obtain the set of batches to be scheduled at the initial moment, and initialize the batches in the set of batches to be scheduled as batch operations. A batch operation refers to the processing operation of a batch in each production stage. The batch operations of all production stages are initialized as an unscheduled batch operation set, and the batch operations of the first production stage are initialized as an available batch operation set.

[0022] Step 3: Select a batch operation from the set of available batch operations using the batch operation selection strategy. If the selected batch operation is a batch operation in the first production stage, it means the batch has not been split yet. Execute the greedy batch operation splitting strategy to split the batch operation into several sub-batch operations, and then execute the machine allocation strategy to allocate a machine for the selected batch operation and process it on the selected machine. If the selected batch operation is not a batch operation in the first production stage, use the splitting result of the batch operation in the first production stage, execute the machine allocation strategy to allocate a machine for the selected batch operation, and process it on the selected machine. The operation process of the greedy batch operation splitting strategy is as follows. Obtain the size of the batch operation to be split, and set the maximum number of sub-batch operations that can be split. Set the initial size of each sub-batch operation to zero, and at the same time set the remaining size of the batch operation to be allocated. Enter a loop, select a unit of 1 from the remaining unallocated batch operation size, and sequentially allocate the selected unit of 1 to each sub-batch operation. Calculate the processing time corresponding to each sub-batch operation, compare the processing times of all sub-batch operations after allocating the selected unit of 1, select the sub-batch operation with the minimum processing time as the final allocation object for the unit of 1, update the size of the sub-batch operation, subtract 1 from the total remaining unallocated batch operation amount, and then perform the next loop until the remaining unallocated total is 0, completing the greedy splitting process of the batch operation.

[0023] Step 4: Execute the position-based insertion strategy to adjust the processing order of the batch operations on the selected machine. The selected machine completes the processing according to the adjusted batch operation order. The position-based insertion strategy includes the following operation process.

[0024] Define parameters: represents the batch set, , represents the total number of batches, represents a certain batch; represents the set of production stages, , represents the total number of production stages, represents a certain production stage; represents the batch 's batch operation; According to the selected batch operation, the selected machine, and the batches that the machine has currently completed scheduling, determine the positions on the machine that can be used as insertion points. Sequentially traverse each insertable position, and judge whether inserting the selected batch operation into the current position meets the feasibility conditions. The feasibility conditions are shown in Formulas 1 and 2, Formula 1 Among them, represents the length of the idle time period of the current machine, represents the batch operation The size of; Indicates batch operation In the production stage The unit processing time of; Indicates batch operation In the production stage The machine setup time of; Formula 1 ensures that the idle time of the selected machine at a certain position can accommodate the minimum processing time required for the selected batch operation. Meeting Formula 1 means that the idle time of the selected machine at a certain position can accommodate the selected batch operation; Formula 2 Wherein, Indicates the latest time of the machine idle time; Indicates the maximum number of sub - batch operations; Indicates batch operation The maximum sub - batch operation of; In the production stage The end time on; Indicates a certain production stage The time to transfer to the next production stage; Indicates batch operation The maximum sub - batch operation of; The size of; Formula 2 ensures that the latest time point of the idle time of the selected machine at a certain position can meet the processing requirements of the last sub - batch operation of the selected batch operation. Meeting Formula 2 means that moving the selected batch operation to a certain position of the selected machine will not affect the arrangement of subsequent batch operations; If the insertion position meets the above feasibility conditions, then calculate the scheduling time of each sub - batch operation of the selected batch operation. From the sum of the end time of the previous batch operation after completion in the current production stage and the machine setup time, and the sum of the end time of the previous batch operation after processing in the previous production stage and the transfer time, select the larger value as the start time of the first sub - batch. From the end time of the previous sub - batch operation of the selected batch operation, and the sum of the end time of the sub - batch operation in the previous production stage and the transfer time, select the larger value as the start time of the subsequent sub - batches. The end time of each sub - batch operation is the sum of the start time and the processing time of the sub - batch operation. The processing time of the sub - batch operation is the unit processing time multiplied by the size of the sub - batch operation. After calculating the time for all sub - batch operations, determine whether inserting the selected batch operation will cause delays to the subsequent sub - batch operations on the selected machine. If there is no delay, insert the selected batch operation into the current position and stop the position - based insertion strategy. Otherwise, skip the current position and continue to try the next insertion point until all positions are traversed or a feasible insertion position is found.

[0025] Step 5: Remove the batch operations that have completed processing from the unscheduled batch operation set and the available batch operation set. If the removed batch operation is not the batch operation of the last production stage, add the batch operation of the next production stage of the removed batch operation to the available batch operation set.

[0026] Step 6: After the batch operations of all production stages have completed processing, find the critical sub-batch operation from all sub-batch operations. The critical sub-batch operation refers to the sub-batch operation with the least total processing time from the first production stage to the last production stage. Execute the strategy based on the critical sub-batch operation to update the entire processing process. Among them, the strategy based on the critical sub-batch operation includes the following operation process.

[0027] Calculate the reduction size of the critical sub-batch operation in each production stage except the first production stage. Traverse all production stages except the first production stage and calculate the reduction size of the sub-batch operation of the batch operation in the production stage . The calculation method is as follows. Among them, the symbol represents rounding down; represents the start time of the sub-batch operation of the batch operation in the production stage ; represents the end time of the sub-batch operation of the batch operation in the production stage ; Select the minimum value from the calculated reduction sizes of all production stages, remove the selected minimum reduction size from the critical sub-batch operation, merge it into the previous sub-batch operation of the critical sub-batch operation, complete the adjustment of the sizes of the two sub-batch operations, and update the scheduling process of the two sub-batch operations in each production stage.

[0028] Step 7: If there are randomly arriving batches from the start to the current moment, convert the randomly arriving batches into batch operations and add them to the unscheduled batch operation set, and add the batch operations of the first production stage to the available batch operation set.

[0029] Step 8: When the unscheduled batch operation set is empty, the batch operations of all batches are completed scheduling.

[0030] Step 9: Use the double deep Q-network to train the policy network based on Steps 2 - 8, and obtain an intelligent scheduling scheme under batch random arrivals using the trained policy network, then terminate the scheduling. The method of using the double deep Q-network to train the policy network based on Steps 2 - 8 is as follows.

[0031] The double deep Q-network adopts a double-network structure, including a policy network and a target network, which are used for action selection and Q-value evaluation respectively. The double deep Q-network combines an experience replay mechanism for stable training. The specific process is as follows: Construct a training dataset, which contains several scheduling instances with a fixed number of production stages. Set the maximum number of training rounds, and initialize the current training round as the first round. Initialize two neural networks, namely the policy network and the target network, and establish an experience replay buffer. At the beginning of each round of training, randomly select an instance from the training dataset for training. During the training process, judge whether the maximum number of training rounds has been reached. If not, continue to execute this round of training; if so, output the trained policy network. In each round of training, the policy network obtains the current shop floor state and checks whether the current scheduling task is completed. If it is completed, increment the current training round by one and start the next round; otherwise, continue to execute the subsequent operations. The policy network uses an ε-greedy strategy to select actions, randomly selects an action with probability ε, and selects the current optimal action with the remaining probability 1 - ε. After executing the action, the policy network obtains a new state and the corresponding reward value. The current state, the selected action, the reward value, and the next state are used as an experience data and stored in the experience replay buffer. The policy network randomly samples a small batch of experience data from the experience replay buffer and calculates the target Q-value y for each piece of experience data. The calculation method is as follows.

[0032] where r represents the reward value obtained after executing action a in the current state s; γ is the discount factor, which is used to measure the importance of future rewards; s′ is the next state reached after executing the action; represents the action value estimated by the target network for state s′ and all actions a′; represents the maximum estimated action value among all actions in state s′; Calculate the loss between the Q-value obtained by the policy network and the target Q-value, and calculate the gradient of the parameters in the policy network through the backpropagation algorithm to update the parameters in the policy network. After reaching the preset number of training rounds, synchronize the parameters of the target network to the parameters of the policy network. The training process is repeated continuously until the set maximum number of training rounds is reached, and then output the trained policy network; The current state is input into the trained policy network to obtain the Q-values corresponding to all actions. An action refers to the combined rules of the batch operation selection policy and the machine allocation policy. The rule combination with the maximum Q-value is selected as the current action, that is, the rule combinations of the batch operation selection policy and the machine allocation policy in the entire scheduling process are determined. The scheduling process is executed to obtain a scheduling plan, and the scheduling is completed.

[0033] To better prove the effectiveness of the present invention, the following further describes and explains the present invention through experimental analysis of a series of examples of the present invention.

[0034] The test data considers four different numbers of production stages, which are taken from the set {3, 5, 8, 10}. The four different numbers of production stages correspond to four different models. In addition, to make the model applicable to different batch sizes, in each production stage, the number of batches is selected from the set {10, 20, 40, 60, 80, 100}. For each model, a total of 1200 different training instances are generated, with 200 instances corresponding to each batch size. Similarly, for each model, a total of 240 test instances are generated, with 40 instances corresponding to each batch size.

[0035] To simulate the arrival of batches in a dynamic environment, for each instance, the release time is generated as follows: for each batch, there is a 50% probability that its release time is 0, indicating that it can be put into production at the initial moment; there is another 50% probability that its release time is randomly generated from the uniform distribution U[1, 100] (seconds). The unit quantity of each batch is selected from the uniform distribution U[50, 100]; the unit processing time is randomly generated from the uniform distribution U[1, 10]; the machine setup time and transfer time are respectively drawn from the uniform distributions U[50, 100] and U[10, 20] (seconds). In addition, the maximum number of sub-batch operations is set to 5. Regarding the due date of each batch, it is selected from the uniform distribution where a represents the minimum completion time of the largest batch, b represents the sum of the minimum completion times of all batches, represents the due date density parameter, which is set to 0.5 in the present invention. represents the release time of the i-th batch. Among them, the calculation methods of a and b are as follows:

[0036] In the double deep Q-network, the adopted deep Q-network consists of four fully connected layers: the input layer contains 256 units for receiving a state vector with the same dimension as the state; followed by two hidden layers, each containing 128 and 64 units respectively, and using the ReLU activation function; the output layer corresponds to 40 actions in the action space and outputs their respective Q-values. The training settings of the policy network are as follows: the maximum number of training epochs is 1000, the discount factor γ is set to 0.99 to balance short-term and long-term rewards; an attenuation factor of 0.8 is applied to the rewards; the Adam optimizer with a learning rate of 1e-4 is used to update the weights of the policy network; the buffer size of the experience replay mechanism is 5000 transition records, and 32 samples are randomly drawn from it each time for network update. The exploration of the policy network starts from 1.0 and gradually decays with a coefficient of 0.995 until it reaches the minimum value of 0.01. To ensure the stability of the training process, the target network is updated every 100 training epochs.

[0037] To evaluate the performance of each comparative algorithm, the present invention uses the average percentage deviation (APD) value as the performance metric for each instance.

[0038] The uniform-based batch splitting strategy (UBLSS) is introduced into the shortest processing time-based construction heuristic (SPTCH), the NEH rule-based construction heuristic (NEHCH), the memory mechanism-based construction heuristic (MCH), and the Johnson rule-based construction heuristic (JCH) respectively to form four comparative algorithms, namely SPTCH_UBLSS, NEHCH_UBLSS, MCH_UBLSS, and JCH_UBLSS. The four comparative algorithms are compared with the knowledge and data hybrid-driven intelligent scheduling method (CCESF_DDQN) under random batch arrivals of the present invention. The experimental results obtained are shown in Table 1: Table 1 Experimental result table of the present invention and construction heuristic algorithms

[0039] Each number of production stages and number of batches in Table 1 constitutes a test problem. For example, 3 production stages and 10 batches will form a test problem 3×10. In the algorithm data in the table, 0.076(0.106), 0.076 represents the APD value obtained by the algorithm, and 0.106 represents the running time of the test problem, with the unit of seconds. It can be seen from the results in Table 1 that the CCESF_DDQN of the present invention has obtained the minimum APD value in all problems, showing the optimal scheduling performance. Through further statistical analysis, it can be concluded that when dealing with problems of different numbers of production stages and batches, the present invention has good stability and low time consumption while ensuring the solution quality, and the comprehensive performance is superior to other comparative algorithms.

[0040] From Figure 3 the violin plots, it can be seen that the data distribution range of CCESF_DDQN of the present invention is relatively small, and the values are mainly concentrated in the lower interval, especially close to zero. In contrast, the data distributions of the other four constructive heuristic algorithms are relatively similar and the range is wider, with greater volatility. In addition, the median of CCESF_DDQN of the present invention is close to zero and the interquartile range is small, indicating that its volatility is low, while the medians of the other comparative algorithms are generally high, close to or exceeding 0.8. The above results show that CCESF_DDQN of the present invention performs excellently in all problems and has strong stability.

[0041] It can be seen from Figure 4 that CCESF_DDQN of the present invention is significantly superior to the other four comparative algorithms both in terms of the mean and the data distribution interval. Therefore, it can be concluded that CCESF_DDQN of the present invention has more advantages than the constructive heuristic algorithms when solving the intelligent scheduling problem with knowledge and data mixed drive under batch random arrival.

[0042] In addition, from Table 1, Figure 5 , Figure 6 it can be seen that as the number of production stages and batches increases, the time consumption of all algorithms shows an upward trend. In contrast, SPTCH_UBLSS has the least time consumption; but according to the analysis of its APD value index, its performance is the worst. CCESF_DDQN of the present invention ranks second lowest in terms of time consumption among all algorithms, while its APD value shows the best performance. Generally speaking, CCESF_DDQN of the present invention shows excellent performance in terms of both effectiveness and efficiency.

[0043] CCESF_DDQN of the present invention is compared with three meta-heuristic algorithms, including cooperative variable neighborhood descent algorithm (CVND), single-point crossover genetic algorithm (GAS2C1) and two-point crossover genetic algorithm (GAS2C2), where the termination time is set to t×n×m (milliseconds), where t is taken from 20, 40 and 80 (milliseconds). When t takes 80 milliseconds, CVND, GAS2C1 and GAS2C2 have basically converged on most test instances. Tables 2, 3 and 4 respectively show the APD values and running times of the test problems of all algorithms on 24 different test problems: Table 2 Experimental result table of the present invention and CVND

[0044] Table 3 Experimental result table of the present invention and GAS2C1

[0045] Table 4 Experimental result table of the present invention and GAS2C2 In Table 2, CVND_t20 represents that CVND is experimented under the condition that t takes 20 milliseconds, and the same applies to Tables 3 and 4. It can be seen from Table 2 that compared with CVND, the CCESF_DDQN of the present invention has achieved the best performance in 23 out of 24 test problems. It can be seen from Table 3 that compared with GAS2C1, the CCESF_DDQN of the present invention has achieved the best performance in 22 out of 24 test problems. It can be seen from Table 4 that compared with GAS2C2, the CCESF_DDQN of the present invention has achieved the best performance in 17 out of 24 test problems. For the problems with 3 production stages, CVND obtained the minimum APD value in the instances with batch numbers of 10, 40, and 60. When the batch numbers were 40 and 60, CVND had reached the convergence state. In these problems, the CCESF_DDQN of the present invention performed sub-optimally. In the problems with more than 3 production stages, GAS2C1 achieved the best performance in 4 test problems with 5 production stages and batch numbers of 20 or 100, 8 production stages and batch number of 60, and 10 production stages and batch number of 20. The CCESF_DDQN of the present invention was second. Except for the above instances that were better than the CCESF_DDQN of the present invention, the CCESF_DDQN of the present invention performed best in most of the other instances. Therefore, the comprehensive performance of the present invention is the best.

[0046] Figure 7 As shown, the data distribution of the CCESF_DDQN of the present invention is more concentrated as a whole, indicating that its performance in multiple test problems is more consistent. Although other meta-heuristic algorithms as comparison algorithms also show a trend of narrowing distribution when the termination time increases, indicating an improvement in the solution accuracy, their distribution ranges are still significantly wider than that of the CCESF_DDQN. By analyzing Figure 7 the widths of the violin plots obtained by each algorithm, it can be found that the CCESF_DDQN has a higher frequency in the lower APD value range, indicating that it has strong effectiveness in optimizing problems. At the same time, the median of the CCESF_DDQN is significantly lower than that of other meta-heuristic algorithms, further indicating that the CCESF_DDQN of the present invention performs better in most cases. From Figure 8 it can be seen that the CCESF_DDQN of the present invention has the best performance in terms of APD value and the smallest error, indicating that the stability and diversity of the scheduling scheme obtained by the present invention are significantly better than those of other comparison algorithms. Through the above analysis, it can be concluded that the CCESF_DDQN of the present invention performs best in terms of accuracy and stability when solving the intelligent scheduling problem of knowledge and data hybrid drive under random arrival of batches.

Claims

1. An intelligent scheduling method driven by a hybrid of knowledge and data under random batch arrival, characterized in that including the following steps, Step 1: Define a policy network, a state space, and an action space. According to the current state, control the policy network to select an action, where each action corresponds to one of the rule combinations of the batch operation selection policy and the machine allocation policy; Step 2: At the initial moment, obtain the set of batches to be scheduled, and initialize the batches in the set of batches to be scheduled as batch operations. Batch operations refer to the processing operations performed by batches in each production stage. The batch operations of all production stages are initially set as the set of unscheduled batch operations, and the batch operations of the first production stage are initially set as the set of available batch operations; Step 3: Select a batch operation from the set of available batch operations using the batch operation selection policy. If the selected batch operation is a batch operation of the first production stage, it means that the batch has not been split yet. Execute the greedy-based batch operation splitting policy to split the batch operation into several sub-batch operations, and then execute the machine allocation policy to allocate a machine for the selected batch operation and perform processing on the selected machine. If the selected batch operation is not a batch operation of the first production stage, use the splitting result of the batch operation of the first production stage, execute the machine allocation policy to allocate a machine for the selected batch operation, and perform processing on the selected machine; Step 4: Execute the position-based insertion policy to adjust the processing order of the batch operations on the selected machine, and the selected machine completes processing according to the adjusted batch operation order; Step 5: Remove the batch operations that have completed processing from the set of unscheduled batch operations and the set of available batch operations. If the removed batch operation is not a batch operation of the last production stage, add the batch operations of the next production stage of the removed batch operation to the set of available batch operations; Step 6: After the batch operations of all production stages have completed processing, find the critical sub-batch operation from all sub-batch operations. The critical sub-batch operation refers to the sub-batch operation with the least total processing time from the first production stage to the last production stage. Execute the policy based on the critical sub-batch operation to update the entire processing process; Step 7: If there are randomly arriving batches from the start to the current moment, convert the randomly arriving batches into batch operations and add them to the set of unscheduled batch operations, and add the batch operations of the first production stage to the set of available batch operations; Step 8: When the set of unscheduled batch operations is empty, the batch operations of all batches are scheduled; Step 9: Use the double deep Q-network to train the policy network based on Steps 2 - 8, and obtain an intelligent scheduling scheme under randomly arriving batches using the trained policy network, and terminate the scheduling.

2. The intelligent scheduling method driven by the mixture of knowledge and data with random arrival of batches according to claim 1, wherein The state space consists of the general attributes of the overall workshop scheduling state and the specific attributes of each production stage. There are 9 general attributes in total, and the specific attribute of each production stage is 18 dimensions. That is, if the workshop contains m production stages, the total dimension of the state space is 9 + 18×m; The action space corresponds to all the rule combinations of the batch operation selection strategy and the machine allocation strategy. There are eight rules for the batch operation selection strategy and five rules for the machine allocation strategy. There are 40 rule combinations between the batch operation selection strategy and the machine allocation strategy, which are used as the output of the policy network.

3. The intelligent scheduling method driven by the mixture of knowledge and data with random arrival of batches according to claim 2, characterized in that, The batch operation selection strategy includes the earliest due date first rule, which preferentially selects the batch operation with the earliest due date for scheduling; the shortest slack time first rule, which preferentially selects the batch operation with the shortest slack time for scheduling. The slack time is defined as the batch due date minus the remaining processing time; the longest processing time first rule, which preferentially selects the batch operation with the longest processing time for scheduling; the shortest processing time first rule, which preferentially selects the batch operation with the shortest processing time; the shortest remaining processing time first rule, which preferentially selects the job with the shortest remaining processing time. The remaining processing time refers to the total processing time required in the subsequent production stage after the current batch operation is completed; the longest remaining processing time first rule, which preferentially schedules the batch operation with the longest remaining processing time; the earliest start time first rule, which sorts according to the earliest start time of each batch operation and preferentially schedules the batch operations ranked higher; the latest start time first rule, which sorts according to the latest start time of each batch operation and preferentially schedules the batch operations ranked higher.

4. The intelligent scheduling method driven by the hybrid of knowledge and data with random arrival of batches according to claim 3, characterized in that The machine allocation strategy includes the shortest waiting time first rule, which preferentially selects the machine with the shortest waiting time for the current batch operation; the load balancing first rule, which preferentially selects the machine with the smallest current load. The load is defined as the number of batch operations assigned to the machine; the earliest idle time first rule, which preferentially selects the machine that will be idle earliest after completing the currently assigned batch operations; the longest idle time first rule, which preferentially selects the machine with the longest idle time when there are idle machines; the earliest completion time first, which preferentially selects the machine with the earliest completion time and assigns the current batch operation to the machine that will complete the current batch operation earliest.

5. The intelligent scheduling method driven by the mixture of next knowledge and data with random arrival of batches according to claim 4, characterized in that, The greedy batch operation splitting strategy includes the following operation process Obtain the size of the batch operation to be split, set the maximum number of sub-batch operations that can be split, set the initial size of each sub-batch operation to zero, and at the same time set the size of the remaining batch operation to be allocated. Enter the loop, select unit 1 from the remaining unallocated batch operation size, and sequentially assign the selected unit 1 to each sub-batch operation. Calculate the processing time corresponding to each sub-batch operation, compare the processing times of all sub-batch operations after the unit 1 is allocated, select the sub-batch operation with the smallest processing time as the final allocation object of unit 1, update the size of the sub-batch operation, subtract 1 from the total amount of the remaining unallocated batch operation, and then perform the next loop until the remaining unallocated total is 0, completing the greedy splitting process of the batch operation.

6. The intelligent scheduling method driven by hybrid knowledge and data with random batch arrival according to claim 5, characterized in that The position-based insertion strategy includes the following operation process Define parameters: Represents the batch set, , Represents the total number of batches, Represents a certain batch; Represents the set of production stages, , Represents the total number of production stages, Represents a certain production stage; Represents the batch batch operation of; Based on the selected batch operation, the selected machine, and the batches that the machine has currently completed scheduling, determine the positions on the machine that can serve as insertion points. Traverse each insertable position in sequence, and judge whether inserting the selected batch operation into the current position meets the feasibility conditions. The feasibility conditions are shown in Formulas 1 and 2. Formula 1 Among them, represents the length of the idle time period of the current machine, represents the batch operation size; represents the batch operation in the production stage unit processing time; represents the batch operation in the production stage machine setup time; Formula 1 ensures that the idle time of the selected machine at a certain position can accommodate the minimum processing time required for the selected batch operation. Satisfying Formula 1 means that the idle time of the selected machine at a certain position can accommodate the selected batch operation; Formula 2 Among them, represents the latest time of the machine idle time; represents the maximum number of sub-batch operations; represents the batch operation of the maximum number of sub-batch operations at the end time in the production stage ; represents the time when a certain production stage transfers to the next production stage; represents the batch operation of the maximum number of sub-batch operations size; Formula 2 ensures that the latest time point of the idle time of the selected machine at a certain position can meet the processing requirements of the last sub-batch operation of the selected batch operation. Satisfying Formula 2 means that moving the selected batch operation to a certain position of the selected machine for processing will not affect the arrangement of subsequent batch operations; If the insertion position meets the above feasibility conditions, calculate the scheduling times of the sub-batch operations of the selected batch operation. Among the sum of the end time after the previous batch operation is completed in the current production stage and the machine setup time, and the sum of the end time after the previous batch operation is processed and completed in the previous production stage and the transfer time, select the larger value as the start time of the first sub-batch. Among the end time of the previous sub-batch operation of the selected batch operation, and the sum of the end time of the sub-batch operation in the previous production stage and the transfer time, select the larger value as the start time of the subsequent sub-batches. The end time of each sub-batch operation is the sum of the start time and the processing time of the sub-batch operation. The processing time of the sub-batch operation is the unit processing time multiplied by the size of the sub-batch operation. After calculating the times of all sub-batch operations, judge whether inserting the selected batch operation causes delays to the subsequent sub-batch operations on the selected machine. If there is no delay, insert the selected batch operation into the current position and stop the insertion strategy based on position. Otherwise, skip the current position and continue to try the next insertion point until all positions are traversed or a feasible insertion position is found.

7. The intelligent scheduling method driven by the hybrid of knowledge and data with random arrival of batches according to claim 6, characterized in that The strategy based on the critical sub-batch operation includes the following operation process. Calculate the reduced size of the key sub-batch operations in each production stage except the first production stage. Traverse all production stages except the first production stage and calculate the production stage in the batch operation of the sub-batch operation of the reduced size , The calculation method is as follows: Among them, the symbol represents rounding down; represents a sub-batch operation of the batch operation at the start time in the production stage; in the production stage on the start time; represents a sub-batch operation of the batch operation at the end time in the production stage; in the production stage on the end time; Select the minimum value from the reduced sizes of all production stages calculated, remove the selected minimum reduced size from the critical sub-batch operation, and merge it into the previous sub-batch operation of the critical sub-batch operation. Complete the adjustment of the sizes of the two sub-batch operations, and update the scheduling processes of the two sub-batch operations in each production stage.

8. A kind of intelligent scheduling method driven by the mixture of knowledge and data with batch random arrival, according to claim 7, characterized in that The method for training the policy network using the double deep Q-network based on Steps 2 - 8 is as follows. The double deep Q-network adopts a double-network structure, including a policy network and a target network, which are used for action selection and Q-value evaluation respectively. The double deep Q-network combines the experience replay mechanism for stable training. The specific process is as follows. Construct a training dataset that contains several scheduling instances with a fixed number of production stages, set the maximum number of training rounds, and initialize the current training round to the first round. Initialize two neural networks, namely the policy network and the target network, and establish an experience replay buffer. At the beginning of each round of training, randomly select an instance from the training dataset for training. During the training process, determine whether the maximum number of training rounds has been reached. If not, continue to execute this round of training; if so, output the trained policy network. In each round of training, the policy network obtains the current shop floor state and checks whether the current scheduling task is completed. If it is completed, increment the current training round by one and start the next round; otherwise, continue to execute the subsequent operations. The policy network uses the ε-greedy strategy to select actions, randomly selects an action with probability ε, and selects the current optimal action with the remaining probability 1 - ε. After executing the action, the policy network obtains the new state and the corresponding reward value. The current state, the selected action, the reward value, and the next state are stored as an experience data in the experience replay buffer. The policy network randomly samples a mini-batch of experience data from the experience replay buffer and calculates the target Q value y for each experience data. The calculation method is as follows, where r represents the reward value obtained after the current state s executes the action a; γ is the discount factor, which is used to measure the importance of future rewards; s′ is the next state reached after executing the action; represents the action value estimated by the target network for the state s′ and all actions a′; represents the action value with the largest estimated value among all actions in the state s′; Calculate the loss between the Q value obtained by the policy network and the target Q value, and calculate the gradients of the parameters in the policy network through the backpropagation algorithm to update the parameters in the policy network. After reaching the preset number of training rounds, synchronize the parameters of the target network to the parameters of the policy network. The training process is repeated continuously until the set maximum number of training rounds is reached, and the trained policy network is output.

9. The intelligent scheduling method driven by a mixture of knowledge and data with random batch arrivals according to claim 8, characterized in that, The current state is input into the trained policy network to obtain the Q values corresponding to all actions. An action refers to the rule combination of the batch operation selection strategy and the machine allocation strategy. Select the rule combination with the maximum Q value as the current action, that is, determine the rule combination of the batch operation selection strategy and the machine allocation strategy in the entire scheduling process, execute the scheduling process to obtain the scheduling plan, and complete the scheduling.

Citation Information

Patent Citations

  • Method for solving single-machine batch scheduling problem under condition of random arrival of different workpieces

    CN111260144A

  • Distributed flow shop scheduling method and system with batch delivery constraint

    CN112286152A

  • Hybrid flow shop scheduling method based on time sequence difference

    CN112734172A

  • Unrelated parallel machine dynamic hybrid flow shop scheduling method based on a deep Q network

    CN113406939A

  • Flexible job shop scheduling method and device based on sub-batch division

    CN116880410A