A method for scheduling multi-agent batch processing machines and variable sub-batch hybrid flow shop

By employing a multi-agent batch processing machine and variable sub-batch hybrid flow shop scheduling method, and utilizing a dual-agent collaborative strategy to optimize the scheduling of batch processing machines and variable sub-batch flow shops, the scheduling complexity problem in complex production lines is solved, achieving efficient and stable production requirements.

CN121258124BActive Publication Date: 2026-03-13LIAOCHENG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-02
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

The scheduling of batch processing machines and variable sub-batch mixed flow workshops presents challenges due to their complexity and high scheduling complexity, especially the production line volatility and equipment idleness and schedule delays caused by differences in annealing time for different steel grades and differences in transportation time between processes.

Method used

A multi-agent batch processing machine and variable sub-batch hybrid pipeline scheduling method is adopted. A high-quality initial solution is constructed through an initialization strategy. Combined with a dual-agent cooperative strategy, global destruction-reconstruction parameters and adaptive local search operators are dynamically selected to optimize the total latency, balance the capacity of batch processing machines and the efficiency of sub-batch flow, and coordinate the task allocation between batch processing and discrete processes.

Benefits of technology

It effectively reduced equipment idle time and project delays, lowered the total delay time, shortened the production cycle, reduced energy waste and cost losses, and improved the stability and efficiency of the production line.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121258124B_ABST
    Figure CN121258124B_ABST
Patent Text Reader

Abstract

This invention relates to a method for scheduling multi-agent batch processing machines and variable sub-batch hybrid pipeline workshops. The scheduling method includes the following steps: Step 1: Establish a problem model; Step 2: Set algorithm running parameters; Step 3: Generate individuals using an initialization strategy; Step 4: Decode the current scheduling state; Step 5: A first agent performs a destructive reconstruction operation; Step 6: A second agent performs a local search operation; Step 7: Apply simulated annealing acceptance criteria; update and decode the scheduling state; Step 8: Determine if the termination condition is met. If it is, the iteration ends, and the current optimal scheduling sequence is output; otherwise, return to Step 5 to continue execution. This invention achieves a dynamic balance between global exploration and local development, thereby reducing total latency and improving workshop scheduling efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of workshop scheduling technology, and in particular to a method for scheduling multi-agent batch processing machines and variable sub-batch mixed flow workshops. Background Technology

[0002] The core characteristics of the mixed batch processing machine and variable sub-batch flow shop problem lie in two aspects: firstly, determining the batch segmentation method, managing the batch formation process under batch processing capacity constraints, and optimizing the sequencing among heterogeneous machines; secondly, the volatility of setup time (e.g., fluctuations caused by tool changes between different steel grades) and the differences in transportation time between processes. These problems are highly complex and challenging, and represent an important research direction in modern production management.

[0003] Taking seamless steel pipe manufacturing as an example, this production line includes processes such as heating, drilling, annealing, and hydrostatic testing. The annealing process utilizes a batch processing machine, capable of simultaneously processing multiple steel pipes of the same specifications. However, the annealing time varies for different steel grades, necessitating strict time control to avoid quality issues such as overheating and insufficient hardness. This hybrid production mode, involving both batch processing and discrete operations, leads to frequent batch splitting and merging, causing fluctuations in batch size and significantly increasing scheduling complexity.

[0004] Therefore, how to rationally formulate scheduling strategies for batch processing machines and variable sub-batch mixed flow workshops to minimize total delay time is a key issue with both practical significance and research value. Summary of the Invention

[0005] To more effectively address the hybrid pipelined workshop scheduling problem involving batch processors and variable subbatches, this study proposes a hybrid pipelined workshop scheduling method with multiple agents for batch processors and variable subbatches. This method fully considers the capacity constraints of the batch processors and the dynamic adjustment characteristics of the variable subbatches. It constructs a high-quality initial solution through an initialization strategy and integrates a dual-agent cooperative strategy (the first agent dynamically selects global destruction-reconstruction parameters, and the second agent adaptively selects local search operators) with the goal of minimizing the total latency (TT). Compared with traditional hybrid flow shop scheduling methods, the proposed method has advantages such as clear implementation logic, precise optimization of core parameters through Taguchi methods, and superior solution quality and stability. By exploring and utilizing multi-agent dynamic balancing, it avoids getting trapped in local optima. Combined with the iterative optimization capabilities of Integral Group (IG), a more accurate scheduling scheme can be obtained through multiple iterations, reducing equipment idleness and schedule delays caused by unreasonable batch sorting and unbalanced sub-batch splitting. Simultaneously, this method effectively coordinates the task allocation between batch processing and discrete processes, balancing batch processing machine capacity limitations and sub-batch cross-process flow efficiency, balancing the load of each machine and process, alleviating production bottlenecks and congestion problems, and improving production line stability. The optimized scheduling scheme not only reduces total delay time and shortens the production cycle but also indirectly reduces energy waste and cost losses caused by ineffective equipment operation, meeting the high-efficiency and stable production needs in complex manufacturing scenarios, and possessing significant theoretical value and practical application significance.

[0006] This invention provides a method for scheduling multi-agent batch processing machines and variable sub-batch hybrid production line workshops, comprising the following steps:

[0007] Step 1: With minimizing the total delay time as the optimization objective, establish a model for the hybrid flow shop scheduling problem of batch processing machines and variable sub-batches;

[0008] Step 2: Set the algorithm running parameters for the problem model;

[0009] Step 3: After determining the algorithm parameters, start executing the algorithm and use the initialization strategy to generate initial individuals;

[0010] Step 4: Decode the current scheduling state based on the results of the initialization of the individual;

[0011] Step 5: The first agent performs a destructive reconstruction operation based on the current scheduling state, generates a new agent, and generates an intermediate scheduling state based on the new agent;

[0012] Step 6: The second agent performs a local search operation based on the intermediate scheduling state to generate a new individual;

[0013] Step 7: Apply the simulated annealing criterion to compare the new individual from Step 6 with the individual from Step 5, and accept the poor solution with a certain probability; update the decoding scheduling state based on the new individual from Step 6.

[0014] Step 8: Determine if the termination condition is met. If it is, the iteration ends and the current best scheduling sequence is output. Otherwise, return to step 5 to continue execution.

[0015] Furthermore, the objective function of the problem model in step 1 is:

[0016] ;

[0017] in, To maximize the completion time, For delivery date.

[0018] Furthermore, the parameters in step 2 include:

[0019] Learning rate Discount rate Decaying greed factor End time: ;

[0020] in, is a coefficient, with values ​​of 100, 200, and 300; For the number of stages, ; For batch number, .

[0021] Furthermore, the initialization of individuals in step 3 employs a two-layer encoding comprising batch sequence and batch segmentation, including:

[0022] Batch sequence is represented as: This is used to indicate the order in which batches are arranged;

[0023] Batch splitting is represented as: This is used to represent the sub-batch division of each batch during the processing stage, where And each ;

[0024] By using the earliest available machine rule, batches are allocated to various machines, thereby generating a complete scheduling scheme;

[0025] in, Batch sequence index, For the first One batch of processing, This is the last batch processed. For the index of stages, Indicates the batch is in the stage The batching situation, For the index of the sub-batch, Indicates batch In the stage The number of batches.

[0026] Furthermore, the batch sequence generation method includes: the batch sequence generation method includes:

[0027] (1) According to The delivery dates of each batch are sorted in descending order to obtain the batch sequence;

[0028] (2) According to The batches are sorted in descending order of their completion time in the first stage to obtain the batch sequence;

[0029] (3) The batches are randomly sorted to obtain the batch sequence.

[0030] Furthermore, the batch segmentation generation method includes:

[0031] (1) At all stages, the batch is partitioned according to the capacity of the batch processor. The formula for partitioning the batch according to the capacity of the batch processor is as follows:

[0032] ;

[0033] (2) During the batch processing stage, the data is divided according to the capacity of the batch processing machine, and during the non-batch processing stage, it is divided uniformly. The formula for the uniform batch division method is as follows:

[0034] ;

[0035] (3) During the batch processing stage, the data is divided according to the capacity of the batch processing machine; during the non-batch processing stage, the data is divided randomly. The formula for the random batch division method is as follows:

[0036] ;

[0037] above, Indicates batch Maximum capacity limit during batch processing. This indicates the total number of batches that have been divided. Indicates batch The number of batches in stage k, where k represents stage and l represents batch. Indicates batch The number of items, This indicates the maximum number of sub-batches that can be split. This indicates the total number of batches that have been split. .

[0038] Furthermore, step 5 includes:

[0039] First, the parameters of the disturbance batch and sub-batch are determined during the destruction phase.

[0040] The first agent selects actions based on a deep neural network model trained using the input scheduling state vector and historical scheduling data, choosing appropriate perturbation parameters within its action space. and ,in Indicates the number of batch sequences of the perturbation. Indicates the number of perturbation sub-batches;

[0041] Then, the reconstruction phase adjusts the batch sequence rearrangement and batch splitting;

[0042] The parameters selected by the first agent reorder the batch sequence and remove the... Each batch is removed from the current solution sequence and then sequentially inserted into the position that minimizes the total delay time of the current solution sequence, based on the parameters. Adjust the sub-batch, in the selected Within each batch, two sub-batches are selected for quantity perturbation. The quantity of the sub-batchens containing a larger quantity is reduced by a certain amount, while the quantity of the sub-batchens containing a smaller quantity is increased accordingly.

[0043] The increment or decrement for disturbance adjustments is limited to 1 to... Within the range, This indicates that the selected sub-batch contains a large number of items;

[0044] Finally, after the destruction and reconstruction operations, new candidate scheduling solutions are generated for subsequent local search and acceptance criterion judgment.

[0045] Furthermore, the action space of the first intelligent agent is specifically composed as follows:

[0046] The action space of the first agent includes two adjustable parameters: the number of batch sequence perturbations. The number of sub-batch quantities adjusted , The value can be 2 or 4. The value can be 2, 3, 4, or 5. and Combining the two, eight actions can be obtained. The first agent selects one of the appropriate combination parameters based on the input state to adjust the batch sequence and the number of sub-batches.

[0047] Furthermore, step 6, the local search operation, performed by the second agent, includes:

[0048] The second agent selects actions based on a deep neural network model trained using the input scheduling state vector and historical scheduling data, and chooses an appropriate local search strategy in its action space.

[0049] The local search strategy comprises a combined strategy of local optimization of batch sequence and local optimization of batch segmentation.

[0050] The batch sequence local optimization strategies include the following four categories:

[0051] (1) In each iteration, about 30% of the batches are randomly selected. Each selected batch is removed from the current scheduling solution and then inserted into the position that minimizes the total delay time of the current scheduling solution.

[0052] (2) Based on the first 30% of the batches of the current best scheduling solution sequence, remove the selected 30% of the batches from the current scheduling solution in turn, and then insert them into the position that minimizes the total delay time of the current scheduling solution. This process continues until the first 30% of the batches of the current best scheduling solution sequence are completed.

[0053] (3) In each iteration, select two adjacent batches from the current scheduling solution sequence and swap the positions of these two batches in the current scheduling solution sequence;

[0054] (4) In each iteration, two batches are randomly selected from the current scheduling solution sequence, and their positions in the current scheduling solution sequence are swapped;

[0055] The batch segmentation local optimization strategy includes the following two categories:

[0056] (1) First, identify batches in the non-batch processing stage that contain at least two non-zero sub-batches, record all interchangeable sub-batch pairs, then randomly select 30% of these sub-batch pairs, swap these batches, record the total delay time after the swap, and finally, if the adjusted scheme improves the solution quality, retain the scheme. The iteration continues until no further improvement can be achieved;

[0057] (2) First, identify batches with at least two sub-batches greater than 3 in the non-batch processing stage, record all adjustable sub-batch pairs, and then for each sub-batch pair, transfer 1 to 5 items from one sub-batch to another. Finally, if the adjusted solution improves the solution quality, retain the solution and repeat this process until no further improvement can be achieved.

[0058] Furthermore, the action space of the second intelligent agent is specifically composed as follows:

[0059] The action space of the second agent consists of local search strategies, namely the batch sequence local optimization strategy and the batch segmentation local optimization strategy. These two strategies are combined to obtain eight local search strategies. The second agent selects one of these strategies to execute based on the input state in order to increase the probability of obtaining a high-quality solution.

[0060] Furthermore, the first or second intelligent agent selects actions based on a deep neural network model trained using historical scheduling data, the training process including:

[0061] First, the state vector corresponding to the current scheduling solution is used as input. The state vector contains eighteen feature values, including batch number, stage number, maximum sub-batch number, number of machines, average machine utilization, standard deviation of machine utilization, mean and standard deviation of sub-batch, mean and standard deviation of total duration, mean and standard deviation of preparation time, mean and standard deviation of transportation time, mean and standard deviation of completion time, and mean and standard deviation of sub-batch number.

[0062] Then, the state vector is input into the Q-network, passes through the hidden layer of the Q-network to the output layer, and the first agent outputs a combination of parameters for destructive reconstruction, including the number of batch sequence perturbations. The number of sub-batch quantities adjusted The output of the second agent is the selection of the local search operator;

[0063] Finally, an experience pool is constructed using historical running data and candidate solutions generated during the algorithm process. Training begins when the number of states in the experience pool reaches a preset threshold, and the current state is uniformly sampled from the experience pool. ,action ,award and the next state Quadruples are used to update network parameters;

[0064] Will Input into the Q network and select the action with the highest Q value. The target Q-value is calculated using the target network according to the following formula:

[0065] ;

[0066] Calculate the mean squared error loss function between the predicted Q-value and the target Q-value using the Q-network. The formula is as follows:

[0067] ;

[0068] in, represents the experience pool, which updates network parameters using gradient descent to continuously improve the accuracy of action selection; t represents the current time. This represents the reward received at the current time. Expressing expectations, Represents the experience pool. Indicates the discount rate. This represents the parameters of the main network. express, These represent the parameters of the target network.

[0069] With minimizing the total latency as the optimization objective, a dual-network structure is adopted, including a Q-network and a target network. The parameters of the Q-network are periodically copied into the target network, thereby improving the stability of the training process and avoiding excessive oscillations.

[0070] Through the above training process, the first and second agents can learn the optimal action selection strategy under different scheduling states, thereby achieving dynamic and adaptive scheduling optimization.

[0071] Furthermore, the first or second intelligent agent selects actions based on a deep neural network model trained using historical scheduling data. The training process includes a reward, specifically:

[0072] The first agent, whose reward function is used to maintain the diversity of the search by destroying the current solution and reconstructing a new solution, aims to balance the improvement of solution quality, exploration diversity and performance stability. This design considers not only short-term effects but also long-term contributions. Slight short-term deterioration is not regarded as an absolute negative situation but is allowed to preserve exploration potential.

[0073] The second agent has a reward function used to balance the effects of local and global optimization during the local search process to ensure the overall quality of the final scheduling solution.

[0074] Furthermore, the first and second intelligent agents interact through the state of the solution, specifically as follows:

[0075] After completing the destruction and reconstruction, the first agent passes the generated candidate scheduling solution as a new input state to the second agent. The second agent then performs a local search operation based on this and feeds back the optimized scheduling solution to the overall scheduling process, so as to achieve a synergistic improvement in search diversity and solution quality.

[0076] Furthermore, the simulated annealing acceptance criteria in step 7 include the following steps:

[0077] If the total delay time of a candidate solution is better than that of the current solution, the candidate solution is accepted; if the total delay time of a candidate solution is worse than that of the current solution, the candidate solution is accepted according to a probability function related to the performance difference and temperature parameters, in order to avoid premature convergence.

[0078] Furthermore, step 8 includes the following steps:

[0079] Determine whether the termination time has been reached based on the running time. If the desired result is not achieved, return to step 5 and continue execution; otherwise, the iteration ends and the current optimal scheduling sequence is output.

[0080] The present invention has the following technical effects:

[0081] Methodological aspects:

[0082] 1. An initialization strategy was designed to construct a high-quality solution, so as to achieve a dynamic balance between global search capability and local optimization capability; at the same time, a multi-agent strategy was introduced, in which the first agent is responsible for dynamically selecting global destruction and reconstruction parameters, and the second agent is responsible for adaptively selecting local search operators.

[0083] 2. Eight local search strategies are used to locally perturb the batch sequence and batch segmentation, thereby optimizing the batch sequence and batch segmentation, further enhancing the adaptability of the solution, and improving the local search capability.

[0084] 3. The multidimensional state reflects the characteristics of the assembly line workshop, realizes the comprehensive quantification of scheduling state and solution quality, provides a basis for agent decision-making and local search, and indirectly supports the evaluation of convergence and solution distribution.

[0085] 4. The reward function for multiple agents: the reward for the first agent considers not only short-term effects but also long-term contributions. Slight short-term deterioration is not considered an absolute negative situation but is allowed to preserve exploration potential. The reward for the second agent is to balance the effects of local optimization and global optimization in the local search process to ensure the overall quality of the final scheduling solution.

[0086] Application level:

[0087] 1. Compared with traditional flow workshop scheduling methods, the method provided by this invention has the advantages of simple implementation, easy parameter adjustment, and excellent performance in solution set convergence;

[0088] 2. It can perform multiple iterations of optimization on production scheduling to obtain a more accurate scheduling plan and avoid errors in the scheduling process;

[0089] 3. It can dynamically schedule and optimize tasks across multiple processes and stages based on production needs, helping companies to ensure production efficiency while also taking into account energy conservation and emission reduction.

[0090] 4. It can balance the load between different stages and machines, avoiding bottlenecks and blockages in the production process, and improving the stability and reliability of the production line. The optimized scheduling scheme can reduce equipment idling time and energy consumption, reduce resource waste, and thus effectively reduce production costs.

[0091] Therefore, the scheduling method for a multi-agent batch processing machine and a variable sub-batch hybrid production line provided by this invention can effectively solve the time and energy consumption conflict problem in complex production line scheduling, provide an efficient and low-energy scheduling scheme for batch processing machines and variable sub-batch hybrid production lines, improve scheduling efficiency, and shorten completion time. Attached Figure Description

[0092] Figure 1 A schematic diagram illustrating the implementation process of this invention;

[0093] Figure 2 The algorithm of this invention is in ARDI index comparison box plot;

[0094] Figure 3 The algorithm of this invention is in ARDI index comparison box plot;

[0095] Figure 4 The algorithm of this invention is in The ARDI index is compared with the box plot. Detailed Implementation

[0096] Embodiments of the invention will now be described more fully below with reference to the accompanying drawings, which illustrate examples of the invention. However, the invention may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the invention to those skilled in the art.

[0097] The specific implementation process of this invention is as follows:

[0098] A method for scheduling a multi-agent batch processing machine and a variable sub-batch hybrid flow shop includes the following steps:

[0099] Step 1: With minimizing the total delay time as the optimization objective, establish a model for the hybrid flow shop scheduling problem of batch processing machines and variable sub-batches;

[0100] Step 2: Set the algorithm running parameters for the problem model;

[0101] Step 3: After determining the algorithm parameters, start executing the algorithm and use the initialization strategy to generate initial individuals;

[0102] Step 4: Decode the current scheduling state based on the results of the initialization of the individual;

[0103] Step 5: The first agent performs a destructive reconstruction operation based on the current scheduling state, generates a new agent, and generates an intermediate scheduling state based on the new agent;

[0104] Step 6: The second agent performs a local search operation based on the intermediate scheduling state to generate a new individual;

[0105] Step 7: Apply the simulated annealing criterion to compare the new individual from Step 6 with the individual from Step 5, and accept the poor solution with a certain probability; update the decoding scheduling state based on the new individual from Step 6.

[0106] Step 8: Determine if the termination condition is met. If it is, the iteration ends and the current best scheduling sequence is output. Otherwise, return to step 5 to continue execution.

[0107] Preferably, the objective function of the problem model in step 1 is:

[0108] ;

[0109] in, To maximize the completion time, For delivery date.

[0110] Preferably, the parameters in step 2 include:

[0111] Learning rate Discount rate Decaying greed factor End time: ;

[0112] in, is a coefficient, with values ​​of 100, 200, and 300; For the number of stages, ; For batch number, .

[0113] Preferably, the initialization of individuals in step 3 employs a two-layer encoding including batch sequence and batch segmentation, comprising:

[0114] Batch sequence is represented as: This is used to indicate the order in which batches are arranged;

[0115] Batch splitting is represented as: This is used to represent the sub-batch division of each batch during the processing stage, where And each ;

[0116] By using the earliest available machine rule, batches are allocated to various machines, thereby generating a complete scheduling scheme;

[0117] in, Batch sequence index, For the first One batch of processing, This is the last batch processed. For the index of stages, Indicates the batch is in the stage The batching situation, For the index of the sub-batch, Indicates batch In the stage The number of batches.

[0118] Preferably, the batch sequence generation method includes:

[0119] (1) According to The delivery dates of each batch are sorted in descending order to obtain the batch sequence;

[0120] (2) According to The batches are sorted in descending order of their completion time in the first stage to obtain the batch sequence;

[0121] (3) is correct The batches are randomly sorted to obtain the batch sequence.

[0122] Preferably, the batch segmentation generation method includes:

[0123] (1) At all stages, the batch is partitioned according to the capacity of the batch processor. The formula for partitioning the batch according to the capacity of the batch processor is as follows:

[0124] ;

[0125] in Indicates batch Maximum capacity limit during batch processing. This indicates the total number of batches that have been divided. Indicates batch In the stage Batch quantity;

[0126] (2) During the batch processing stage, the data is divided according to the capacity of the batch processing machine, and during the non-batch processing stage, it is divided uniformly. The formula for the uniform batch division method is as follows:

[0127] ;

[0128] in Indicates batch The number of items, Indicates the maximum number of sub-batches;

[0129] (3) During the batch processing stage, the data is divided according to the capacity of the batch processing machine; during the non-batch processing stage, the data is divided randomly. The formula for the random batch division method is as follows:

[0130] ;

[0131] in This indicates the total number of batches that have been divided.

[0132] Preferably, the decoding process in step 4 includes the following steps:

[0133] First, determine the processing order of each batch based on the batch sequence;

[0134] Secondly, the sub-batch division of each batch is determined according to the batch segmentation scheme;

[0135] Then, the processing order of the sub-batch is arranged according to the earliest arrival priority principle, and each batch is assigned to a specific machine according to the earliest available machine principle;

[0136] Finally, the start time, completion time and total delay time of each batch are calculated to obtain the scheduling status of the current solution.

[0137] The scheduling state of the current solution consists of the following eighteen states:

[0138] Quantity of batches ;

[0139] Number of stages ;

[0140] Maximum number of sub-batches ;

[0141] Number of machines ;

[0142] The average machine utilization rate is calculated by first determining the utilization rate of each machine by dividing its total processing time by its occupied time (including idle time), then summing the utilization rates of all machines and dividing by the total number of machines to obtain the average value. The formula for average machine utilization rate is as follows:

[0143] ;

[0144] in, Indicates in Stage Machine Loading time Indicates in Stage Machine The completion date;

[0145] Formula for the standard deviation of machine utilization:

[0146] ;

[0147] Formula for the average number of sub-batches:

[0148] ;

[0149] in, Indicates batch exist The number of non-zero sub-batches at each stage;

[0150] Formula for the standard deviation of sub-batch quantity:

[0151] ;

[0152] Formula for the average total delay time:

[0153] ;

[0154] Formula for the standard deviation of total delay time:

[0155] ;

[0156] Formula for average preparation time:

[0157] ;

[0158] in, Indicates the current batch of solutions In the stage Preparation time in stages;

[0159] Formula for the standard deviation of preparation time:

[0160] ;

[0161] Formula for average transit time:

[0162] ;

[0163] in, Indicates batch exist The number of non-zero sub-batches at each stage;

[0164] Formula for the standard deviation of transit time:

[0165] ;

[0166] Formula for the average completion time:

[0167] ;

[0168] Formula for the standard deviation of completion time:

[0169] ;

[0170] Formula for sub-batch mean:

[0171] ;

[0172] in, Indicates batch exist The number of non-zero sub-batches at each stage;

[0173] Formula for sub-batch standard deviation:

[0174] .

[0175] Preferably, step 5 includes:

[0176] First, the parameters of the disturbance batch and sub-batch are determined during the destruction phase.

[0177] The first agent selects actions based on a deep neural network model trained using the input scheduling state vector and historical scheduling data, choosing appropriate perturbation parameters within its action space. and ,in Indicates the number of batch sequences of the perturbation. Indicates the number of perturbation sub-batches;

[0178] Then, the reconstruction phase adjusts the batch sequence rearrangement and batch splitting;

[0179] The parameters selected by the first agent reorder the batch sequence and remove the... Each batch is removed from the current solution sequence and then inserted sequentially into the position that minimizes the total delay time of the current solution sequence, based on the parameters. Adjust the sub-batch, in the selected Within each batch, two sub-batches are selected for quantity perturbation. The sub-batches with larger quantities are reduced by a certain amount, while the sub-batches with smaller quantities are increased accordingly.

[0180] The increment or decrement for disturbance adjustments is limited to 1 to... Within the range, This indicates that the selected sub-batch contains a large number of items;

[0181] Finally, after the destruction and reconstruction operations, new candidate scheduling solutions are generated for subsequent local search and acceptance criterion judgment.

[0182] Preferably, step 6, the local search operation, is performed by the second agent and includes:

[0183] The second agent selects actions based on a deep neural network model trained using the input scheduling state vector and historical scheduling data, and chooses an appropriate local search strategy in its action space.

[0184] The local search strategy comprises a combined strategy of local optimization of batch sequence and local optimization of batch segmentation.

[0185] The batch sequence local optimization strategies include the following four categories:

[0186] (1) In each iteration, about 30% of the batches are randomly selected. Each selected batch is removed from the current scheduling solution and then inserted into the position that minimizes the total delay time of the current scheduling solution.

[0187] (2) Based on the first 30% of the best scheduling solution sequence, remove the selected 30% of the batches from the current scheduling solution in turn, and then insert them into the position that minimizes the total delay time of the current scheduling solution. This process continues until the first 30% of the best scheduling solution sequence is completed.

[0188] (3) In each iteration, select two adjacent batches from the current scheduling solution sequence and swap the positions of these two batches in the current scheduling solution sequence.

[0189] (4) In each iteration, two batches are randomly selected from the current scheduling solution sequence and their positions in the current scheduling solution sequence are swapped.

[0190] The batch segmentation local optimization strategy includes the following two categories:

[0191] (1) First, identify batches in the non-batch processing stage that contain at least two non-zero sub-batches, record all interchangeable sub-batch pairs, then randomly select 30% of these sub-batch pairs, swap these batches, record the total delay time after the swap, and finally retain the scheme if the adjusted scheme improves the solution quality. The iteration continues until no further improvement can be achieved.

[0192] (2) First, identify batches with at least two sub-batches greater than 3 in the non-batch processing stage, record all adjustable sub-batch pairs, and then for each sub-batch pair, transfer 1 to 5 items from one sub-batch to another. Finally, if the adjusted solution improves the solution quality, retain the solution and repeat this process until no further improvement can be achieved.

[0193] Preferably, the action space composition of the second intelligent agent includes the following steps:

[0194] The action space of the second agent consists of local search strategies, namely the batch sequence local optimization strategy and the batch segmentation local optimization strategy. These two strategies are combined to obtain eight local search strategies. The second agent selects one of these strategies to execute based on the input state in order to increase the probability of obtaining a high-quality solution.

[0195] Preferably, the first or second intelligent agent selects actions based on a deep neural network model trained using historical scheduling data, and the training process includes the following steps:

[0196] First, the state vector corresponding to the current scheduling solution is used as input. The state vector contains eighteen feature values, including batch number, stage number, maximum sub-batch number, number of machines, average machine utilization, standard deviation of machine utilization, mean and standard deviation of sub-batch, mean and standard deviation of total duration, mean and standard deviation of preparation time, mean and standard deviation of transportation time, mean and standard deviation of completion time, and mean and standard deviation of sub-batch number.

[0197] Then, the state vector is input into the Q-network, passes through the hidden layer of the Q-network to the output layer, and the first agent outputs a combination of parameters for destructive reconstruction, including the number of batch sequence perturbations. The number of sub-batch quantities adjusted The output of the second agent is the selection of the local search operator;

[0198] Finally, an experience pool is constructed using historical running data and candidate solutions generated during the algorithm process. Training begins when the number of states in the experience pool reaches a preset threshold, and the current state is uniformly sampled from the experience pool. ,action ,award and the next state Quadruples are used to update network parameters;

[0199] Will Input into the Q network and select the action with the highest Q value. The target Q-value is calculated using the target network according to the following formula:

[0200]

[0201] Calculate the mean squared error loss function between the predicted Q-value and the target Q-value using the Q-network. The formula is as follows:

[0202]

[0203] in, This represents the experience pool, and the network parameters are updated using the gradient descent method to continuously improve the accuracy of action selection;

[0204] With minimizing the total latency as the optimization objective, a dual-network structure is adopted, including a Q-network and a target network. The parameters of the Q-network are periodically copied into the target network, thereby improving the stability of the training process and avoiding excessive oscillations.

[0205] Through the above training process, the first and second agents can learn the optimal action selection strategy under different scheduling states, thereby achieving dynamic and adaptive scheduling optimization.

[0206] Preferably, the deep neural network structure includes the following steps:

[0207] The deep neural network includes a Q-network and a target network, both with the same structure, comprising a total of six layers: an input layer, an output layer, and hidden layers. The input layer receives eighteen scheduling state features. The output layer of the first agent is a combination of parameters for destructive reconstruction. The number of batch sequence perturbations... The number of sub-batch quantities adjusted The output layer of the second agent is the selection of local search operators, and the hidden layers in the middle have 128, 256, 128, 64, and 32 neurons. Each layer uses the LeakyReLU activation function to improve the model's ability to capture nonlinear patterns. The network parameters are optimized using the Adam optimizer.

[0208] Preferably, the reward includes the following steps:

[0209] The first agent, whose reward function is used to maintain the diversity of the search by destroying the current solution and reconstructing a new solution, aims to balance the improvement of solution quality, exploration diversity and performance stability. This design considers not only short-term effects but also long-term contributions. Slight short-term deterioration is not regarded as an absolute negative situation but is allowed to preserve exploration potential.

[0210] Final contribution Defined as the relative improvement rate of the local optimal solution relative to the solution before destruction, it reflects the actual contribution of the destruction and reconstruction operation to the quality of the final solution, as shown in the following formula:

[0211] ;

[0212] in, This represents the solution after a local search. This indicates the solution after destruction and reconstruction;

[0213] Introducing penalty items To prevent excessive degradation caused by rebuilding, only when short-term performance degradation exceeds a predetermined threshold... The penalty will only be activated if the change in solution quality remains within an acceptable range. This allows for slight fluctuations to preserve the diversity of exploration, as shown in the following formula:

[0214] ;

[0215] in, This indicates the solution before destruction and reconstruction;

[0216] Based on this, the two parts are combined using a weighted average to obtain the final reward for the first agent. The formula is as follows:

[0217] ;

[0218] in, This indicates that the final contribution is given higher weight, prioritizing improvements in solution quality, while This serves only as a slight constraint on overexploration, ensuring that reasonable exploration is not excessively suppressed;

[0219] The second agent has a reward function used to balance the effects of local and global optimization during the local search process to ensure the overall quality of the final scheduling solution.

[0220] Local improvement items The formula for measuring the relative improvement of the solution after local optimization compared to the solution before local search is as follows:

[0221] ;

[0222] in This represents the solution after a local search. This represents the solution before the local search;

[0223] Global Improvement Items The following formula is used to evaluate the improvement of the current optimal solution compared to the previous optimal solution:

[0224] ;

[0225] in, This represents the previous optimal solution. Indicates the current optimal solution

[0226] Based on this, the two parts are combined using a weighted average to obtain the final reward for the first agent. The formula is as follows:

[0227] ;

[0228] in, and These represent the weights for local optimization and global optimization, respectively. The larger the value, the more the algorithm prioritizes local optimization. The smaller the value, the greater the weight of global optimization, thus prioritizing the overall improvement of the solution.

[0229] Preferably, the first agent and the second agent interact through the state of the solution, including the following steps:

[0230] After completing the destruction and reconstruction, the first agent passes the generated candidate scheduling solution as a new input state to the second agent. The second agent then performs a local search operation based on this and feeds back the optimized scheduling solution to the overall scheduling process, so as to achieve a synergistic improvement in search diversity and solution quality.

[0231] Preferably, the simulated annealing acceptance criteria in step 7 include the following steps:

[0232] If the total delay time of a candidate solution is better than that of the current solution, the candidate solution is accepted; if the total delay time of a candidate solution is worse than that of the current solution, the candidate solution is accepted according to a probability function related to the performance difference and temperature parameters, in order to avoid premature convergence.

[0233] Preferably, step 8 includes the following steps:

[0234] Determine whether the termination time has been reached based on the running time. If the desired result is not achieved, return to step 5 and continue execution; otherwise, the iteration ends and the current optimal scheduling sequence is output.

[0235] This application utilizes reinforcement learning and a dual-agent DDQN framework to adaptively decide on global perturbations and local optimizations. Specifically, the first agent dynamically determines the parameters of the global perturbation by integrating accumulated learning experience with real-time state feedback and selecting from eight combinations of destructive and reconstructive parameters. Simultaneously, the second agent adaptively selects the most suitable local search operator from a policy set containing eight combined local search methods. This collaborative mechanism coordinates global batch sorting and splitting adjustments with targeted local optimization, effectively reducing ineffective random exploration. The invention is further described below with reference to specific embodiments:

[0236] The simulation experiment used 100 standard examples. This paper selected two test case sets, denoted as follows: and ,in Depend on Composed of small-scale instances, here , . Depend on Composed of medium to large-scale instances, here , Each combination contains five different instances of the same size, with the number of machines in each stage... Uniformly generated within the range. The sub-batch processing time in the discrete stage is... Randomly generated within a range, the sub-batch processing time in the batch processing stage is within... Randomly generated within a specified range. Sequence preparation time and transport time are respectively drawn from uniform distributions. and Selected from the options.

[0237] To verify the effectiveness of the proposed multi-agent reinforcement learning algorithm (MADDQN_IG), this paper selects several high-performance optimization algorithms proposed in recent years as comparison objects, including: an improved co-evolutionary algorithm (vCCEA) combining variable neighborhood search strategy (VNS), a novel cooperative iterative greedy algorithm (NCIG), a combination of metaheuristic and Q-learning (QABC), and a genetic algorithm (GA). To reduce the randomness of experimental results and enhance the reliability of statistical conclusions, each example is run independently 5 times, and its statistical results are used as the final performance evaluation basis. Regarding performance metrics, the relative deviation index (RDI) is used as the performance evaluation metric. The calculation formula for the evaluation metric is as follows:

[0238] Relative Deviation Index (RDI);

[0239] The Relative Deviation Index (RDI) is a commonly used performance metric in literature related to total delay (TT). The RDI is calculated as follows:

[0240] ;

[0241] in, Indicates the algorithm index. and These represent the best and worst TT results among all algorithms, respectively. RDI values ​​range from 0 to 1, with smaller values ​​indicating better performance. Specifically, when... Approaching When RDI approaches 0; conversely, when Approaching When RDI approaches 1, if all algorithms obtain the same TT value ( To avoid division by zero and to reflect equivalent performance, RDI is defined as 0.

[0242] Average Total Time Delay (ATT) refers to the average of the results obtained by an algorithm solving a single instance multiple times. ATT is a key characteristic of this algorithm. ,and The calculated ARDI is called the Average Relative Bias Index (ARDI), which is a documented metric used to compare the performance of algorithms in repeated experiments.

[0243] Figures 2-4 Examples of the comparison algorithms are given. Different The optimization results are presented intuitively. Specifically, from... Figures 2-4 The ARDI metrics show that MADDQN_IG consistently has a lower ARDI value, indicating that it performs better in terms of performance.

[0244] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Any modifications or changes made to the present invention by those skilled in the art after reading this application and referring to the above embodiments are within the scope of protection claimed in the pending claims of this application.

Claims

1. A multi-agent batch processing machine and variable sub-batch hybrid flow shop scheduling method, characterized in that, The method comprises the following steps: Step 1: a mixed flow shop scheduling problem model of batch machines and variable sub-batches is established with the optimization objective of minimizing total delay time; Step 2: algorithm running parameters are set for the problem model; Step 3: after the algorithm parameters are determined, the algorithm is executed, and an initialization strategy is used to generate an initialization individual; Step 4: a current scheduling state is decoded according to the result of the initialization individual; Step 5: a first agent executes a destruction and reconstruction operation according to the current scheduling state, generates a new individual, and generates an intermediate scheduling state according to the new individual; Step 6: a second agent executes a local search operation according to the intermediate scheduling state, and generates a new individual again; Step 7: a simulated annealing criterion is applied to compare the new individual in the step 6 and the individual in the step 5, and a poor solution is accepted with a preset probability; the decoding scheduling state is updated according to the new individual in the step 6; Step 8: whether a termination condition is met is judged, if yes, the iteration is ended, and a current best scheduling sequence is output, otherwise, the step 5 is returned to continue execution; In the step 3, the initialization individual adopts double-layer encoding including a batch sequence and batch segmentation, and the batch sequence generation mode comprises: (1) According to the delivery dates of the batches are sorted in descending order to obtain a batch sequence; (2) according to the completion times of the batches in the first phase are sorted in descending order to obtain a batch sequence; (3) The batches are randomly sorted to obtain the batch sequence; The batch segmentation generation mode comprises: (1) in all stages, the batches are segmented according to the capacity of the batch machine, and the formula for segmenting the batches according to the capacity of the batch machine is as follows: ; (2) in the batch processing stage, the batches are segmented according to the capacity of the batch machine, and in the non-batch processing stage, the batches are uniformly segmented, and the formula for uniformly segmenting the batches is as follows: ; (3) in the batch processing stage, the batches are segmented according to the capacity of the batch machine, and in the non-batch processing stage, the batches are randomly segmented, and the formula for randomly segmenting the batches is as follows: ; Above, denotes the batch maximum capacity limit at the batch stage, denotes the total number of batches that have been split, denotes the batch number of batches at stage k, k denotes the stage, and I denotes the batch, denotes the batch number of items, denotes the maximum number of split sub-batches; In the step 5, the following steps are included: First, the destruction stage determines the parameters of the disturbed batches and sub-batches The first intelligent agent selects an action according to a deep neural network model trained by an input scheduling state vector and historical scheduling data, and selects a disturbance parameter in the action space thereof and wherein represents the number of disturbance batch sequences, represents the number of disturbance sub-batches; Then, the reconstruction stage adjusts the batch sequence rearrangement and batch segmentation, The parameter selected by the first intelligent agent, reordering the batch sequence, removing the removed batch from the current solution sequence, and then sequentially entering the removed batch into a position that minimizes the total delay time of the current solution sequence according to the parameter adjusting the sub-batches, selecting two sub-batches from the selected batch to perform quantity perturbation, reducing the quantity of the sub-batch containing a larger quantity, and increasing the quantity of the sub-batch containing a smaller quantity; The number of additions or subtractions of the disturbance adjustment is limited to a range of 1 to wherein represents that the number of items of the selected sub-lot is more; Finally, a new candidate scheduling solution is generated after the destruction and reconstruction operations, which is used for subsequent local search and acceptance criterion judgment.

2. The method of claim 1, wherein, The action space of the first agent is composed of the following: The action space of the first agent includes two adjustable parameters: the number of batch sequence perturbations. The number of sub-batch quantities adjusted , The value can be 2 or 4. The value can be 2, 3, 4, or 5. and Combining the two, eight actions can be obtained. The first agent selects one of the combined parameters based on the input state to adjust the batch sequence and the number of sub-batches.

3. The method of claim 1, wherein, The local search operation in the step 6 is executed by the second agent and comprises: The second agent selects an action in its action space according to a deep neural network model trained based on an input scheduling state vector and historical scheduling data; The local search strategy comprises a batch sequence local optimization strategy and a batch segmentation local optimization strategy; The batch sequence local optimization strategy comprises the following four types: (1) in each iteration process, 30% of the batches are randomly selected, each selected batch is removed from the current scheduling solution in turn, and then inserted into a position that minimizes the total delay time of the current scheduling solution; (2) according to the first 30% of the batch sequence of the current scheduling solution, 30% of the batches are removed from the current scheduling solution in turn, and then inserted into a position that minimizes the total delay time of the current scheduling solution, until the first 30% of the batches of the current scheduling solution sequence are completed; (3) In each iteration, select two adjacent batches from the current scheduling solution sequence, and exchange the positions of the two batches in the current scheduling solution sequence; (4) In each iteration, randomly select two batches from the current scheduling solution sequence, and exchange the positions of the two batches in the current scheduling solution sequence; The batch segmentation local optimization strategy includes the following two types: (1) First, identify batches containing at least two non-zero sub-batches in the non-batch processing stage, record all exchangeable sub-batch pairs, then randomly select 30% of the batches from these sub-batch pairs, then exchange the batches, record the total delay time after the exchange, and finally if the adjusted scheme improves the solution quality, keep the scheme; the iteration continues until no further improvement can be made; (2) First, identify batches containing at least two sub-batches with a quantity greater than 3 in the non-batch processing stage, record all adjustable sub-batch pairs, then for each sub-batch pair, transfer 1 to 5 items from one sub-batch to another, and finally if the adjusted scheme improves the solution quality, keep the scheme, and repeat the process until no further improvement can be made.

4. The method of claim 3, wherein, The action space of the second agent is composed of local search strategies, and the batch sequence local optimization strategy and the batch segmentation local optimization strategy are combined to obtain eight local search strategies, and the second agent selects one of them according to the input state to improve the probability of obtaining a high-quality solution. The first agent or the second agent selects actions based on a deep neural network model trained from historical scheduling data, and the training process includes:

5. The method according to claim 1 or 3, characterized in that, First, use the state vector corresponding to the current scheduling solution as input, and the state vector contains eighteen characteristic values, including batch quantity, stage quantity, maximum sub-batch quantity, machine quantity, machine average utilization rate, machine utilization rate standard deviation, sub-batch mean and standard deviation, total duration time mean and standard deviation, preparation time mean and standard deviation, transportation time mean and standard deviation, completion time mean and standard deviation, and sub-batch quantity mean and standard deviation; Minimizing total delay time as the optimization goal, using a double network structure including a Q network and a target network, periodically copying the parameters of the Q network to the target network to improve the stability of the training process and avoid excessive oscillation; Then, the state vector is input to the Q network, and through the hidden layer to the output layer of the Q network, the first agent outputs a parameter combination for destroying the reconstruction, containing the number of batch sequence disturbances and the number of sub-batch quantity adjustments , and the second agent outputs the selection of the local search operator; Finally, an experience pool is constructed by historical running data and candidate solutions generated in the algorithm process. When the number of the experience pool reaches a preset threshold, training is started, and the current state , action , reward and next state four-tuple are uniformly sampled from the experience pool to update the network parameters; Will Input into the Q network and select the action with the highest Q value. The target Q-value is calculated using the target network according to the following formula: ; a mean squared error loss function between the Q value predicted by the Q network and the target Q value , as follows: ; wherein, represents an experience pool, and the network parameters are updated by a gradient descent method to constantly improve the accuracy of action selection; t represents the current time, represents the reward obtained at the current time, represents the expectation, represents the experience pool, represents the discount rate, represents the parameters of the main network, represents, represents the parameters of the target network; Through the above training process, the first agent and the second agent can learn the optimal action selection strategy under different scheduling states, thereby realizing dynamic and adaptive scheduling optimization. The first agent or the second agent selects actions based on a deep neural network model trained from historical scheduling data, and the training process includes rewards, which are specifically:

6. The method of claim 1, wherein, ​ The reward function of the first agent is used to maintain the diversity of the search by destroying the current solution and reconstructing a new solution, aiming to balance the improvement of the quality of the solution, the diversity of exploration and the stability of performance, and the design not only considers the short-term effect, but also considers the long-term contribution, and slight short-term deterioration is not regarded as an absolute negative situation, but is allowed to retain the exploration potential; The reward function of the second agent is used to balance the local optimization and global optimization effects in the local search process, so as to ensure the overall quality of the final scheduling solution.

7. The method according to any of claims 1 to 4, 6, characterized in that, The first agent and the second agent interact through the state of the solution, specifically: After completing the destruction and reconstruction, the first agent transmits the generated candidate scheduling solution as a new input state to the second agent, and the second agent performs a local search operation on the basis of the new input state, and feeds back the optimized scheduling solution to the overall scheduling process, so as to realize the cooperative improvement of the search diversity and the quality of the solution.

Citation Information

Patent Citations

  • Real-time order receiving and scheduling method and system for flexible job shop

    CN119558432A

  • Dynamic optimization and real-time decision-making method and device for multi-agent collaborative target search

    CN119828460A