A buffer workpiece group batch transfer and order allocation collaborative optimization method and system

By employing a buffer zone workpiece batch transfer and order allocation collaborative optimization method in a multi-stage hybrid flow workshop, and utilizing a deep reinforcement learning model to optimize transfer batches and order allocation, the problem of the separation between processing scheduling and material transfer scheduling in existing technologies is solved, achieving efficient, flexible and energy-saving operation of the production system.

CN122175095APending Publication Date: 2026-06-09WUHAN BUSINESS UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
WUHAN BUSINESS UNIV
Filing Date
2026-04-20
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

Existing multi-stage hybrid flow workshop scheduling methods fail to effectively coordinate and optimize processing scheduling and material transfer scheduling, resulting in high AGV idle rates, frequent start-ups and shutdowns, serious energy waste, and difficulty in coping with dynamic disturbances such as emergency order insertions and order changes, leading to poor robustness of the scheduling scheme.

Method used

A collaborative optimization method for buffer workpiece batch transfer and order allocation is adopted. By splitting orders and processing flows, real-time batch clustering and dual threshold triggering mechanisms, combined with a deep reinforcement learning model, batch-transfer equipment-path combination instructions are generated to optimize the allocation of transfer batches and orders, thereby achieving efficient matching of processing resources and transfer resources.

Benefits of technology

It significantly reduces the idle operation and energy consumption of transfer equipment, responds quickly to production disturbances, improves the robustness and energy efficiency of the production system, and achieves global optimization of production efficiency and energy efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122175095A_ABST
    Figure CN122175095A_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of intelligent manufacturing and production scheduling optimization, and particularly relates to a buffer workpiece group batch transfer and order allocation collaborative optimization method and system. The method firstly performs real-time group batch clustering on the to-be-transferred workpieces in the buffer based on attribute correlation to generate a transfer batch, and issues a transfer instruction under a double-threshold control pulse trigger mechanism. Then, a deep reinforcement learning model is constructed. At each pulse trigger time, a global state vector is constructed based on the transfer batches of all issued transfer instructions and the order types in the system, an order allocation decision is output based on the global state vector, and finally a batch-transfer equipment-path combination instruction is generated based on the order allocation decision. The application realizes global optimization of production efficiency and energy efficiency through collaborative optimization of machining scheduling and material transfer scheduling and deep integration of dynamic response of orders.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent manufacturing and production scheduling optimization technology, specifically involving a method and system for collaborative optimization of workpiece batch transfer and order allocation in a buffer zone. Background Technology

[0002] Multi-stage hybrid flow shop scheduling is a core challenge in discrete manufacturing, and its optimization objectives typically include minimizing efficiency indicators such as completion time and delay time. With the deepening of green manufacturing concepts, the energy consumption of material handling equipment (such as automated guided vehicles, AGVs) has become an operating cost that enterprises cannot ignore.

[0003] Existing research has two main shortcomings: First, most scheduling methods consider processing scheduling and material transfer scheduling separately, ignoring the fact that unscientific workpiece batching strategies can lead to high AGV idle rates and frequent start-ups and shutdowns, resulting in huge energy waste. Second, traditional scheduling schemes are mostly static or reactive, making it difficult to effectively cope with dynamic disturbances such as the frequent insertion of emergency orders, order cancellations, or quantity changes in actual production, resulting in poor robustness and low overall energy efficiency of the scheduling scheme. Summary of the Invention

[0004] The purpose of this invention is to address the aforementioned problems in the existing technology by providing a buffer workpiece batch transfer and order allocation collaborative optimization method and system that achieves collaborative optimization of processing scheduling and material transfer scheduling and can deeply integrate dynamic order response to realize global optimization of production efficiency and energy efficiency.

[0005] To achieve the above objectives, the technical solution of the present invention is as follows:

[0006] In a first aspect, the present invention provides a method for collaborative optimization of buffer workpiece batch transfer and order allocation, the method comprising:

[0007] S1. Split the order to get a set of workpieces, and then split the workpiece processing flow in sequence to get a set of processing stages. A buffer is set in front of each machine in each processing stage.

[0008] S2. Perform real-time batch clustering of the workpieces to be transferred in the buffer based on attribute correlation to generate transfer batches, and issue transfer instructions using a pulse triggering mechanism with dual threshold control to transport the workpieces in the transfer batch from the current buffer to the next processing stage through the transfer equipment.

[0009] S3. Construct a deep reinforcement learning model. At each pulse trigger time, construct a global state vector based on all transfer batches and order types that have issued transfer instructions in the system. Input the global state vector into the deep reinforcement learning model to output order allocation decisions. Finally, generate batch-transfer equipment-path combination instructions based on the order allocation decisions.

[0010] In S3, the deep reinforcement learning model selects one scheduling rule from a variety of preset scheduling rules, allocates orders based on the selected scheduling rule and the order type in the global state vector, obtains the order allocation decision, and then constructs an optimization scheduling problem with the objective of minimizing the total order delay and the total energy consumption of the transfer system. Based on the order allocation decision and the global state vector, the optimization scheduling problem is solved to generate batch-vehicle-path combination instructions.

[0011] In S2, the real-time batch clustering includes:

[0012] The attribute correlation of workpiece pairs is calculated based on material correlation, process correlation, and time-domain correlation. Workpiece pairs with attribute correlation higher than a preset correlation threshold are aggregated into the same transfer batch. The functional expression of the attribute correlation is as follows:

[0013] ;

[0014] ;

[0015] ;

[0016] ;

[0017] In the above formula, For attribute relevance; , , respectively workpiece pair Material-related, process-related, and time-related; , , These are material weights, process weights, and time-domain weights, respectively. , These are material weights and size weights, respectively. Indicates workpiece With workpiece Material compatibility, , respectively workpiece With workpiece Material code, Indicates workpiece With workpiece Size similarity; , respectively workpiece , workpiece Dimensional values; , Representing the workpieces respectively , workpiece A set of processes; Indicates workpiece With workpiece The heat treatment similarity; , respectively workpiece , workpiece Delivery time; Indicates the time scaling factor; Indicates priority similarity.

[0018] In S2, the pulse triggering mechanism of dual threshold control means that at each pulse triggering moment, if the number of workpieces in the transfer batch reaches the preset lower limit of the number of workpieces that the transfer equipment can carry, or if the waiting time of the transfer batch in the buffer reaches the preset maximum waiting time, then a transfer instruction is issued for the transfer batch.

[0019] In S3, the global state vector includes: workpiece completion rate, average machine load rate, transfer equipment load rate, batch merging quality, and order type code; wherein,

[0020] The expression for the workpiece completion rate is:

[0021] ;

[0022] In the above formula, Indicates the processing stage The completion rate of the workpiece; For processing stage Workpiece is being processed. The machine number; For processing stage The number of machines; Indicates the machine number index; For workpiece During the processing stage China Machinery Processing time; Indicates workpiece Whether to allocate to the processing stage China Machinery Decision variables on the workpiece Assigned to the processing stage China Machinery The value is 1 when processing is performed on the workpiece, and 0 otherwise; It refers to orders The first in One workpiece;

[0023] The expression for the average machine load rate is:

[0024] ;

[0025] ;

[0026] In the above formula, Indicates the processing stage Average machine load rate; for Time processing stage The machine in Load rate; for Time processing stage The machine in The number of workpieces that have been processed; For processing stage The time when the machine starts processing Duration of the moment;

[0027] The expression for the load rate of the transfer equipment is:

[0028] ;

[0029] In the above formula, for The load rate of the transfer equipment at any given time; For transfer equipment exist The number of workpieces being carried at any given time; Index for transfer equipment numbers; This represents the total number of transfer devices; This refers to the maximum number of workpieces that the transfer equipment can carry.

[0030] The expression for the batch merging quality is:

[0031] ;

[0032] In the above formula, For batch merging quality; This represents the total number of shipments. For transshipment batches Average correlation of attributes;

[0033] The order type coding refers to the use of one-hot coding to represent order types. Order types include ordinary orders, rush orders, and orders with varying scales. Rush orders and orders with varying scales are dynamic orders.

[0034] The scheduling rules include:

[0035] Scheduling rule 1:

[0036] When the order type is an expedited order or an order with a change in size, the order will be assigned to the machine with the lowest average machine load rate.

[0037] When the order type is a regular order, the order will be assigned to the machine with the lowest workpiece completion rate;

[0038] Scheduling rule 2:

[0039] When the order type is an expedited order, calculate the change in the standard deviation of the machine load rate before and after assigning the order to a machine in the current processing stage, and assign the order to the machine with the smallest change in standard deviation.

[0040] When the order type is a regular order or a variable size order, the order will be assigned to the machine with the earliest expected completion time.

[0041] Scheduling rule 3:

[0042] When the order type is expedited, the order will be assigned to the machine with the lowest load rate on the transfer equipment;

[0043] When the order type is a regular order or a variable size order, the order will be assigned to machines other than those with the highest and lowest load rates on the transfer equipment.

[0044] Scheduling rule 4:

[0045] When the order type is an expedited order or an order with a change in size, the order will be assigned to the machine with the highest batch merging quality: ,in This indicates the machine that belongs to the current processing step. The total number of transit batches waiting to be transited in the next buffer; For transshipment batches Average correlation of attributes;

[0046] When the order type is a regular order, the order will be randomly assigned to any machine other than the one with the highest and lowest batch quality. .

[0047] The objective function of the optimization scheduling problem is expressed as:

[0048] ;

[0049] ;

[0050] In the above formula, , These are the first objective function and the second objective function, respectively. , These represent the total order delay and the total energy consumption of the transit system, respectively. , Orders Delivery time and expected completion time; Indicates the total number of orders; Index the number of transfer equipment in the transfer system; For transfer equipment Total runtime; , Transfer equipment Energy consumption per unit time under load operation and no-load operation; , Transfer equipment The time functions corresponding to load operation and no-load operation; For transfer equipment Startup energy consumption; For transfer equipment Number of times it is started;

[0051] The constraints of the optimization scheduling problem include constraints on the number of transfer batches, transfer time, and processing time of the workpiece at each processing stage.

[0052] In S3, a composite reward function is constructed based on the total delay penalty, total energy consumption penalty, constraint penalty, and machine load reward. The parameters of the deep reinforcement learning model are trained based on this composite reward function. The expression of the composite reward function is as follows:

[0053] ;

[0054] ;

[0055] ;

[0056] ;

[0057] ;

[0058] ;

[0059] ;

[0060] ;

[0061] In the above formula, for Momentary rewards; , , , They are respectively Total delay penalty, total energy consumption penalty, constraint penalty, and machine load reward at any given time; , , , These are the weighting coefficients for total delay penalty, total energy consumption penalty, constraint penalty, and machine load reward, respectively. For machines Weighting coefficients; This is the penalty coefficient for delays; This is the total energy consumption penalty coefficient; for Total energy consumption of the transportation system at any given time; , Transfer equipment Energy consumption per unit time under load operation and no-load operation; , Transfer equipment The time functions corresponding to load operation and no-load operation; For transfer equipment Startup energy consumption; For transfer equipment Number of times it is started; For transfer equipment Total runtime; To constrain the penalty coefficient for violations; The overload unit penalty coefficient; For transfer equipment exist The total number of workpieces being carried at any given time; This refers to the maximum number of workpieces that the transfer equipment can carry. Indicates the processing stage Average machine load rate; for Time processing stage Machine load rate; For processing stage The number of machines in the system; for Time processing stage The machine in The number of workpieces that have been processed; This represents the total number of transfer devices; Total number of orders; Indicates the machine number index; Indicates the index of the processing stage; This indicates the index of the transfer equipment.

[0062] In S3, the batch-vehicle-path combination instruction includes the following triplet information: batch information, path information, and time information, wherein the batch information includes the set of workpieces contained in the transfer batch, and the path information includes the starting position of the transfer batch. Destination of the transshipment batch From the starting point of the transfer batch To the finish line The path, the time information includes the time when the transfer instruction was issued. Expected travel time of transfer equipment and expected completion time , .

[0063] Secondly, the present invention provides a buffer workpiece batch transfer and order allocation collaborative optimization system, the system comprising:

[0064] The order and process splitting module is used to split orders to obtain a set of workpieces, and to split the workpiece processing flow in sequence to obtain a set of processing stages. A buffer is set in front of each machine in each processing stage.

[0065] The transfer batch generation module is used to perform real-time batch clustering of workpieces to be transferred in the buffer based on attribute correlation to generate transfer batches, and to issue transfer instructions using a pulse triggering mechanism with dual threshold control so that the workpieces in the transfer batch can be transported from the current buffer to the next processing stage through the transfer equipment.

[0066] The order allocation module is used to build a deep reinforcement learning model. At each pulse trigger time, a global state vector is constructed based on all transfer batches and order types that have issued transfer instructions in the system. The global state vector is input into the deep reinforcement learning model to output the order allocation decision. Finally, a batch-transfer equipment-path combination instruction is generated based on the order allocation decision.

[0067] The deep reinforcement learning model is used to select a scheduling rule from a variety of preset scheduling rules, allocate orders based on the selected scheduling rule and the order type in the global state vector, obtain the order allocation decision, and then construct an optimization scheduling problem with the objective of minimizing the total order delay and the total energy consumption of the transfer system. Based on the order allocation decision and the global state vector, the optimization scheduling problem is solved to generate batch-vehicle-path combination instructions.

[0068] The batch transfer generation module is used to perform real-time batch clustering according to the following steps:

[0069] The attribute correlation of workpiece pairs is calculated based on material correlation, process correlation, and time-domain correlation. Workpiece pairs with attribute correlation higher than a preset correlation threshold are aggregated into the same transfer batch. The functional expression of the attribute correlation is as follows:

[0070] ;

[0071] ;

[0072] ;

[0073] ;

[0074] In the above formula, For attribute relevance; , , respectively workpiece pair Material-related, process-related, and time-related; , , These are material weights, process weights, and time-domain weights, respectively. , These are material weights and size weights, respectively. Indicates workpiece With workpiece Material compatibility, , respectively workpiece With workpiece Material code, Indicates workpiece With workpiece Size similarity; , respectively workpiece , workpiece Dimensional values; , Representing the workpieces respectively , workpiece A set of processes; Indicates workpiece With workpiece The heat treatment similarity; , respectively workpiece , workpiece Delivery time; Indicates the time scaling factor; Indicates priority similarity.

[0075] The transfer batch generation module is used to issue transfer instructions using the following dual-threshold controlled pulse triggering mechanism: at each pulse triggering moment, if the number of workpieces in the transfer batch reaches the preset lower limit of the number of workpieces that the transfer equipment can carry, or if the waiting time of the transfer batch in the buffer reaches the preset maximum waiting time, then a transfer instruction is issued for the transfer batch.

[0076] The global state vector includes: workpiece completion rate, average machine load rate, transfer equipment load rate, batch merging quality, and order type code; among which...

[0077] The expression for the workpiece completion rate is:

[0078] ;

[0079] In the above formula, Indicates the processing stage The completion rate of the workpiece; For processing stage Workpiece is being processed. The machine number; For processing stage The number of machines; Indicates the machine number index; For workpiece During the processing stage China Machinery Processing time; Indicates workpiece Whether to allocate to the processing stage China Machinery Decision variables on the workpiece Assigned to the processing stage China Machinery The value is 1 when processing is performed on the workpiece, and 0 otherwise; It refers to orders The first in One workpiece;

[0080] The expression for the average machine load rate is:

[0081] ;

[0082] ;

[0083] In the above formula, Indicates the processing stage Average machine load rate; for Time processing stage The machine in Load rate; for Time processing stage The machine in The number of workpieces that have been processed; For processing stage The time when the machine starts processing Duration of the moment;

[0084] The expression for the load rate of the transfer equipment is:

[0085] ;

[0086] In the above formula, for The load rate of the transfer equipment at any given time; For transfer equipment exist The number of workpieces being carried at any given time; Index for transfer equipment numbers; This represents the total number of transfer devices; This refers to the maximum number of workpieces that the transfer equipment can carry.

[0087] The expression for the batch merging quality is:

[0088] ;

[0089] In the above formula, For batch merging quality; This represents the total number of shipments. For transshipment batches Average correlation of attributes;

[0090] The order type coding refers to the use of one-hot coding to represent order types. Order types include ordinary orders, rush orders, and orders with varying scales. Rush orders and orders with varying scales are dynamic orders.

[0091] The scheduling rules include:

[0092] Scheduling rule 1:

[0093] When the order type is an expedited order or an order with a change in size, the order will be assigned to the machine with the lowest average machine load rate.

[0094] When the order type is a regular order, the order will be assigned to the machine with the lowest workpiece completion rate;

[0095] Scheduling rule 2:

[0096] When the order type is an expedited order, calculate the change in the standard deviation of the machine load rate before and after assigning the order to a machine in the current processing stage, and assign the order to the machine with the smallest change in standard deviation.

[0097] When the order type is a regular order or a variable size order, the order will be assigned to the machine with the earliest expected completion time.

[0098] Scheduling rule 3:

[0099] When the order type is expedited, the order will be assigned to the machine with the lowest load rate on the transfer equipment;

[0100] When the order type is a regular order or a variable size order, the order will be assigned to machines other than those with the highest and lowest load rates on the transfer equipment.

[0101] Scheduling rule 4:

[0102] When the order type is an expedited order or an order with a change in size, the order will be assigned to the machine with the highest batch merging quality: ,in This indicates the machine that belongs to the current processing step. The total number of transit batches waiting to be transited in the next buffer; For transshipment batches Average correlation of attributes;

[0103] When the order type is a regular order, the order will be randomly assigned to any machine other than the one with the highest and lowest batch quality. .

[0104] The objective function of the optimization scheduling problem is expressed as:

[0105] ;

[0106] ;

[0107] In the above formula, , These are the first objective function and the second objective function, respectively. , These represent the total order delay and the total energy consumption of the transit system, respectively. , Orders Delivery time and expected completion time; Indicates the total number of orders; Index the number of transfer equipment in the transfer system; For transfer equipment Total runtime; , Transfer equipment Energy consumption per unit time under load operation and no-load operation; , Transfer equipment The time functions corresponding to load operation and no-load operation; For transfer equipment Startup energy consumption; For transfer equipment Number of times it is started;

[0108] The constraints of the optimization scheduling problem include constraints on the number of transfer batches, transfer time, and processing time of the workpiece at each processing stage.

[0109] The order allocation module is used to construct a composite reward function based on total delay penalty, total energy consumption penalty, constraint penalty, and machine load reward. This composite reward function is then used to train the parameters of the deep reinforcement learning model. The expression for the composite reward function is:

[0110] ;

[0111] ;

[0112] ;

[0113] ;

[0114] ;

[0115] ;

[0116] ;

[0117] ;

[0118] In the above formula, for Momentary rewards; , , , They are respectively Total delay penalty, total energy consumption penalty, constraint penalty, and machine load reward at any given time; , , , These are the weighting coefficients for total delay penalty, total energy consumption penalty, constraint penalty, and machine load reward, respectively. For machines Weighting coefficients; This is the penalty coefficient for delays; This is the total energy consumption penalty coefficient; for Total energy consumption of the transportation system at any given time; , Transfer equipment Energy consumption per unit time under load operation and no-load operation; , Transfer equipment The time functions corresponding to load operation and no-load operation; For transfer equipment Startup energy consumption; For transfer equipment Number of times it is started; For transfer equipment Total runtime; To constrain the penalty coefficient for violations; The overload unit penalty coefficient; For transfer equipment exist The total number of workpieces being carried at any given time; This refers to the maximum number of workpieces that the transfer equipment can carry. Indicates the processing stage Average machine load rate; for Time processing stage Machine load rate; For processing stage The number of machines in the system; for Time processing stage The machine in The number of workpieces that have been processed; This represents the total number of transfer devices; Total number of orders; Indicates the machine number index; Indicates the index of the processing stage; This indicates the index of the transfer equipment.

[0119] The batch-vehicle-path combination instruction includes the following triplet information: batch information, path information, and time information, wherein the batch information includes the set of workpieces contained in the transfer batch, and the path information includes the starting position of the transfer batch. Destination of the transshipment batch From the starting point of the transfer batch To the finish line The path, the time information includes the time when the transfer instruction was issued. Expected travel time of transfer equipment and expected completion time , .

[0120] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0121] 1. The method described in this invention firstly avoids production conflicts and workpiece backlogs in material flow between different processing stages by structurally decomposing orders and processing flows and rationally setting buffers. Secondly, in the batch transfer stage, real-time clustering based on attribute correlation is combined with a dual-threshold pulse triggering mechanism to accurately aggregate highly correlated workpieces into transfer batches, minimizing idle travel, invalid waiting, and frequent start-stop of transfer equipment, thereby reducing energy consumption in material transfer from the source. Finally, by using a deep reinforcement learning model to comprehensively perceive the dynamics of transfer batches and order type characteristics through global state vectors, it can quickly respond to various production disturbances such as emergency order insertions and order increases or decreases, making the final output order allocation decision and batch-transfer equipment-path combination instructions have real-time and dynamic adaptability, greatly improving the robustness of the production system to dynamic changes, achieving efficient matching of processing resources and transfer resources, shortening the total order delay and total energy consumption of the transfer system, and ultimately achieving global optimization of production efficiency and energy efficiency.

[0122] 2. The method described in this invention dynamically selects scheduling rules through a deep reinforcement learning model, achieving intelligent rule selection and situational adaptability. Compared to a single fixed rule or manual rule switching, it significantly improves the flexibility and response speed of scheduling decisions in complex dynamic environments. It decouples order allocation decisions from subsequent optimization scheduling problems in a layered manner, utilizing the rapid decision-making capability of deep reinforcement learning to meet real-time requirements while ensuring global optimality through optimization scheduling, thus achieving synergistic optimization of both delivery time and energy consumption objectives. Specifically, in urgent order scenarios, the model tends to select rule 1 and combine it with small-scale batching to effectively reduce the risk of order delays. In low-urgency scenarios, rule 2 is selected and large-scale batching is executed, significantly improving AGV load rate and energy utilization efficiency. The introduction of rules 3 and 4 ensures the priority aggregation of highly correlated workpieces and the dynamic balance of system resources. Finally, through the generation of batch-transfer equipment-path combination instructions, it achieves end-to-end optimization from rule selection to path planning, improving the overall operating efficiency and sustainability of the manufacturing system.

[0123] 3. The global state vector in the method described in this invention is a structured representation formed by workpiece completion rate, average machine load rate, transfer equipment load rate, batch merging quality, and order type encoding. The state space of a deep reinforcement learning model is constructed based on the global state vector. The model outputs the optimal order allocation decision adapted to the current state through the mapping relationship learned during training. Finally, based on the optimal order allocation decision, combined with the optimization scheduling problem and constraints, a batch-transfer equipment-path combination instruction containing batch division, transfer equipment assignment, and driving path is generated, realizing end-to-end decision-making from state perception to action execution, and improving the intelligence level of order allocation and transfer scheduling.

[0124] 4. The method described in this invention constructs a composite reward function based on total delay penalty, total energy consumption penalty, constraint penalty, and machine load reward. This multi-dimensional reward mechanism enables the deep reinforcement learning model to fully perceive the comprehensive impact of scheduling decisions. Compared with a single reward mechanism, the order allocation decision output by the trained model is more in line with the complex needs of actual production scenarios. Attached Figure Description

[0125] Figure 1 This is a flowchart of the method described in this invention.

[0126] Figure 2 This is a schematic diagram of the multi-stage mixed flow workshop addressed in this invention.

[0127] Figure 3 The reward value is obtained using the method described in this invention during performance verification.

[0128] Figure 4 The total delay is obtained using the method described in this invention during performance verification.

[0129] Figure 5 The total energy consumption obtained using the method described in this invention during performance verification.

[0130] Figure 6 This is the solution front obtained using the method described in this invention during performance verification.

[0131] Figure 7 This is a structural block diagram of the system described in this invention. Detailed Implementation

[0132] The present invention will now be described in further detail with reference to specific embodiments and accompanying drawings.

[0133] Example 1:

[0134] See Figure 1 A method for collaborative optimization of buffer workpiece batch transfer and order allocation is performed in the following steps:

[0135] S1. Split the order to get a set of workpieces, and split the workpiece processing flow according to the processing order to get a set of processing stages. A buffer is set in front of each machine in each processing stage.

[0136] like Figure 2 As shown, a set of customer orders that the factory needs to process is denoted as... ,in For the total number of orders, each order It can be divided into several standardized workpieces with independent processing procedures, denoted as... ,in For orders The total number of workpieces included. The processing flow includes... A series of sequential processing stages, in each processing stage There is a group of machines with identical functions. Available for selection, among which For processing stage The total number of machines included. Each workpiece can only be processed by one machine in each processing stage, and the processing time and energy consumption required on that machine are known parameters. The movement of workpieces between different processing stages is accomplished through a transfer system, which includes multiple transfer devices (e.g., AGVs). A buffer zone is set up in front of each machine in each processing stage, where workpieces meeting the transfer conditions are dynamically aggregated into transfer batches within the buffer zone. ,in This is the maximum number of transfer batches, determined by the number of transfer devices in the transfer system. Each batch It contains a set of workpieces with a clearly defined destination, each batch The total number of workpieces in the transfer equipment shall not exceed the maximum number of workpieces it can carry. Once all processing steps are completed, workpieces belonging to the same order will be transferred to the designated assembly workshop. Final assembly is carried out, in which This refers to the number of assembly workshops within the factory.

[0137] S2. Based on attribute correlation, perform real-time batch clustering of workpieces to be transferred in the buffer to generate transfer batches, and issue transfer instructions using a pulse triggering mechanism with dual threshold control to transport the workpieces in the transfer batch from the current buffer to the next processing stage through the transfer equipment.

[0138] Specifically, the real-time batch clustering includes:

[0139] Structured data such as workpiece material codes and dimensions are obtained from the enterprise bill of materials (BOM) system; process data such as workpiece process routes and heat treatment requirements are obtained from the manufacturing execution system (MES); and time-domain data such as order delivery time and priority are obtained from the enterprise resource planning (ERP) system. The above data is then normalized to map it to the [0, 1] interval.

[0140] The attribute correlation of workpiece pairs is calculated based on material correlation, process correlation, and time-domain correlation. Workpiece pairs with attribute correlation higher than a preset correlation threshold are aggregated into the same transfer batch. ,in Indicates the first Each batch awaiting transfer is the basic unit for batch scheduling. This represents the transfer relation operator, indicating that two workpieces will be combined into the same transfer batch for task execution. and Both represent workpiece-machine pairs, where , All of these are workpiece identifiers. and respectively workpiece Workpiece The current machine number. Workpieces with attribute correlation exceeding a preset correlation threshold are grouped into the same transfer batch, ensuring high material homogeneity, process compatibility, and processing sequence synchronization within the same transfer batch. This significantly reduces the frequency of loading / unloading and path switching on transfer equipment. Specifically, the introduction of material correlation reduces tool and mold change time caused by material differences, the introduction of process correlation reduces machine waiting time caused by process route conflicts, and the introduction of temporal correlation ensures a relatively balanced delivery time for batch workpieces, preventing urgent orders from being dragged down by non-urgent orders. Ultimately, this achieves a dual optimization of transfer efficiency and production coordination.

[0141] The functional expression for the attribute correlation is as follows:

[0142] ;

[0143] ;

[0144] ;

[0145] ;

[0146] In the above formula, For attribute relevance; , , respectively workpiece pair Material-related, process-related, and time-related; , , These are material weights, process weights, and time-domain weights, respectively. ; , These are material weights and size weights, respectively. ; Indicates workpiece With workpiece Material compatibility, , respectively workpiece With workpiece Material code, if workpiece With workpiece If the material codes are the same, then The value is 1 if it is 1, otherwise it is 0. Indicates workpiece With workpiece Size similarity, ; , respectively workpiece , workpiece Dimensional values; , Representing the workpieces respectively , workpiece The set of processes, Indicates workpiece With workpiece The number of common processes, Indicates workpiece With workpiece The number of different processes; Indicates workpiece With workpiece The heat treatment similarity is calculated as follows: if the heat treatment process requirements are the same, the heat treatment similarity value is 1; otherwise, it is 0. , respectively workpiece , workpiece Delivery time; This represents the time scaling factor, which is set to 0.01 in this embodiment. Indicates workpiece With workpiece The priority similarity is 1 if the priorities are the same, otherwise it is 0; the priority of the workpiece can be divided into different priority levels according to the preset customer importance level corresponding to its order.

[0147] Weight vector The generation and updating of weights occur during the subsequent training of the deep reinforcement learning model. They are modeled as trainable parameters of the policy network in the deep reinforcement learning model. During training, the agent interacts with the state space, obtains rewards, and automatically adjusts all trainable parameters of the policy network through the backpropagation algorithm to maximize the cumulative reward. After the model converges, the weight vectors... This is determined and embedded in the model for online scheduling. The agent adjusts the weight vector in real time based on the current state. This enables the generated transfer batches to autonomously adapt to complex and dynamic production environments.

[0148] Specifically, the pulse-triggered mechanism of the dual-threshold control refers to: setting a basic decision cycle. Every The system will automatically trigger a scheduling pulse at the specified time. If the number of workpieces in the transfer batch reaches the maximum capacity of the transfer equipment at the time of the scheduling pulse trigger, the system will automatically initiate a transfer. Or the transit batch has reached the preset maximum waiting time in the buffer zone. If the system issues a transfer instruction for the transfer batch, the dispatching system will assign transfer equipment to perform the transfer task for that batch.

[0149] By presetting the maximum number of workpieces that can be carried. This avoids resource waste caused by transporting equipment before it is fully loaded, ensures economies of scale in batch transport, reduces transport frequency and equipment start-up and shutdown times, thereby reducing energy consumption; and it also reduces energy consumption by setting a maximum waiting time. This system effectively prevents workpieces from accumulating in the buffer zone for extended periods, avoiding downtime in subsequent processing stages or order delays due to transfer delays, thus ensuring the continuity of the production process. The pulse-triggered mechanism, through fixed-cycle detection and dual-threshold judgment, enables the transfer system to handle both efficient batch transfers under stable loads and rapid response to dynamic disturbances, enhancing the flexibility and robustness of the transfer system. The pulse-triggered mechanism also includes an emergency scheduling pulse. Reaching a critical threshold in the buffer zone is set as a trigger event. When a trigger event occurs, an emergency scheduling pulse is immediately activated, and the scheduling system assigns transfer equipment to perform the transfer task for the corresponding batch in the buffer zone. This quickly responds to abnormal backlog situations, preventing congestion from spreading to preceding processing stages and further optimizing the buffer zone's flow efficiency. Trigger events include: the number of workpieces accumulated in any buffer zone exceeding a preset safe quantity limit.

[0150] S3. At each pulse trigger moment, freeze and collect all transfer batches with transfer instructions issued in the current buffer, forming a transfer batch pool to be decided; construct a deep reinforcement learning model, construct a global state vector based on the transfer batch pool and order type, input the global state vector into the deep reinforcement learning model to output order allocation decision, and finally generate batch-transfer equipment-path combination instructions based on the order allocation decision.

[0151] Specifically, the global state vector includes: workpiece completion rate, average machine load rate, transfer equipment load rate, batch merging quality, and order type code; wherein,

[0152] The expression for the workpiece completion rate is:

[0153] ;

[0154] In the above formula, Indicates the processing stage The completion rate of the workpiece; For processing stage Workpiece is being processed. The machine number; For processing stage The number of machines; Indicates the machine number index; For workpiece During the processing stage China Machinery Processing time; Indicates workpiece Whether to allocate to the processing stage China Machinery Decision variables on the workpiece Assigned to the processing stage China Machinery The value is 1 when processing is performed on the workpiece, and 0 otherwise; It refers to orders The first in One workpiece;

[0155] The expression for the average machine load rate is:

[0156] ;

[0157] ;

[0158] In the above formula, Indicates the processing stage Average machine load rate; for Time processing stage The machine in Load rate; for Time processing stage The machine in The number of workpieces that have been processed; For processing stage The time when the machine starts processing Duration of the moment;

[0159] The expression for the load rate of the transfer equipment is:

[0160] ;

[0161] In the above formula, for The load rate of the transfer equipment at any given time; For transfer equipment exist The number of workpieces being carried at any given time; Index for transfer equipment numbers; This represents the total number of transfer devices; This refers to the maximum number of workpieces that the transfer equipment can carry.

[0162] The expression for the batch merging quality is:

[0163] ;

[0164] In the above formula, For batch merging quality; This represents the total number of shipments. For transshipment batches Average correlation of attributes;

[0165] The order type coding refers to using one-hot encoding to represent order types. Order types include regular orders, rush orders, and orders with varying quantities. Rush orders and orders with varying quantities are dynamic orders, with orders involving increases or decreases in the number of workpieces. Different types of dynamic orders can be distinguished through order type coding.

[0166] In constructing the global state vector, the overall progress of production is reflected by the workpiece completion rate (the proportion of workpieces with completed processes to the total number of workpieces); the average machine load rate (the average ratio of the current load to the rated capacity of each processing machine) characterizes the balanced utilization of manufacturing resources; the transfer equipment load rate (the current task occupancy rate and mileage utilization rate of transfer equipment) reflects the busy state of logistics resources; the batch merging quality (the average comprehensive correlation of workpieces within a transfer batch based on material correlation, process correlation, and time-domain correlation) assesses the rationality of batching decisions; and order type encoding reflects the dynamic changes of orders by one-hot encoding of order types. These five types of data are normalized and concatenated into a fixed-dimensional global state vector, which is then input into a deep reinforcement learning model. The model, through the mapping relationships learned during training, outputs the optimal order allocation decision adapted to the current state to determine the target machine assignment for each order. Finally, based on this optimal order allocation decision and combined with the optimization scheduling problem and constraints, a batch-transfer equipment-path combination instruction, including batch division, transfer equipment assignment, and travel path, is generated, achieving end-to-end decision-making from state perception to action execution.

[0167] Specifically, the state space of a deep reinforcement learning model is constructed based on a global state vector, while its action space consists of several predefined scheduling rules. Different scheduling rules correspond to different machine selection strategies based on order types in the global state vector. During training, the deep reinforcement learning model selects one scheduling rule from multiple options and then... and current status Orders are assigned based on their order types (i.e., the current global state vector), resulting in an order assignment decision. This decision determines which machine a workpiece should be assigned to. Then, an optimization scheduling problem is constructed with the objective of minimizing the total order delay and the total energy consumption of the transfer system. Based on the order assignment decision and the global state vector, this optimization scheduling problem is solved to generate batch-vehicle-path combination instructions. Compared to traditional scheduling models that rely on manual experience or single rules, this approach not only improves the intelligence level of order assignment and transfer scheduling but also achieves synergistic optimization of resource utilization efficiency and order processing timeliness, significantly enhancing the system's robustness to dynamic order changes.

[0168] Specifically, the scheduling rules are used not only to allocate dynamic orders in real time when they occur, but also to regenerate globally consistent batch-vehicle-path combination instructions for all workpieces that have not completed their entire processing step at each scheduling pulse trigger time. In this embodiment, there are four types of scheduling rules, namely:

[0169] Scheduling rule 1:

[0170] When the order type is an expedited order or an order with a change in size, the order will be assigned to the machine with the lowest average machine load rate.

[0171] When the order type is a regular order, the order will be assigned to the machine with the lowest workpiece completion rate;

[0172] Scheduling rule 2:

[0173] When the order type is an expedited order, calculate the change in the standard deviation of the machine load rate before and after assigning the order to a machine in the current processing stage, and assign the order to the machine with the smallest change in standard deviation.

[0174] When the order type is a regular order or a variable size order, the order will be assigned to the machine with the earliest expected completion time.

[0175] Scheduling rule 3:

[0176] When the order type is expedited, the order will be assigned to the machine with the lowest load rate on the transfer equipment;

[0177] When the order type is a regular order or a variable size order, the order will be assigned to machines other than those with the highest and lowest load rates on the transfer equipment.

[0178] In a multi-stage mixed assembly line workshop, different buffer zones are set up in front of different machines within each processing stage. The transfer tasks in these buffer zones are handled by different AGVs. Therefore, the choice of which machine to process the workpiece directly determines which buffer the completed workpiece will enter, thus affecting the subsequent workload of the AGV responsible for transferring tasks in that buffer. When the order type is an urgent order, it is assigned to the machine of the AGV with the lowest current load rate in its corresponding downstream buffer for processing. This ensures that urgent workpieces can be transferred with priority and quickly, avoiding waiting due to busy AGVs.

[0179] Scheduling rule 4:

[0180] When the order type is an expedited order or an order with a change in size, the order will be assigned to the machine with the highest batch merging quality: ,in This indicates the machine that belongs to the current processing step. The total number of transit batches waiting to be transited in the next buffer; For transshipment batches Average correlation of attributes;

[0181] When the order type is a regular order, the order will be randomly assigned to any machine other than the one with the highest and lowest batch quality. ;

[0182] Since high batch merging quality means that the workpiece currently carried by the AGV is highly matched with the expedited / scale change orders in terms of material attributes, process routes, and time requirements, it can reduce the mold change and adjustment time in the downstream processing links. Therefore, when the order type is an expedited order or a scale change order, it is assigned to the machine of the AGV with the highest current batch merging quality in its corresponding downstream buffer for processing. It can be immediately integrated into the existing high batch merging quality batches, realizing plug-and-play.

[0183] Specifically, the objective function of the optimization scheduling problem is expressed as:

[0184] ;

[0185] ;

[0186] In the above formula, , These are the first objective function and the second objective function, respectively. , These represent the total order delay and the total energy consumption of the transit system, respectively. , Orders Delivery time and expected completion time; Indicates the total number of orders; Index the number of transfer equipment in the transfer system; For transfer equipment Total runtime; , Transfer equipment Energy consumption per unit time under load operation and no-load operation; , Transfer equipment The time functions corresponding to load operation and no-load operation; For transfer equipment Startup energy consumption; For transfer equipment The number of times it is started.

[0187] Specifically, the constraints of the optimization scheduling problem include constraints on the number of transfer batches, transfer time, and processing time constraints for the workpiece at each processing stage. The constraints on the number of transfer batches include:

[0188] ;

[0189] In the above formula, for Number of transfers at any given time; This is the maximum quantity allowed for each shipment.

[0190] The transit time constraints include:

[0191] ;

[0192] In the above formula, Indicates workpiece Expected completion time; Indicates from the processing stage To the processing stage The expected travel time of the transfer equipment; Indicates workpiece The start time of the transfer;

[0193] The processing time constraints for the workpiece at each processing stage include:

[0194] ;

[0195] In the above formula, For workpiece During the processing stage China Machinery Start processing time; For workpiece During the processing stage China Machinery Processing time; Indicates workpiece Whether to allocate to the processing stage China Machinery Decision variables on the workpiece Assigned to the processing stage China Machinery The value is 1 when processing is performed, otherwise it is 0; For workpiece During the processing stage China Machinery The completion time.

[0196] Specifically, the batch-vehicle-path combination instruction is an instantaneous execution segment of the global scheduling scheme at the moment of a single pulse trigger. The global scheduling scheme is composed of a series of time-sequential batch-vehicle-path combination instructions. A complete batch-vehicle-path combination instruction includes the following triplet information: batch information, path information, and time information. The batch information includes the set of workpieces contained in the transfer batch, and the path information includes the starting position of the transfer batch. (i.e., the output buffer location), the destination location of the transfer batch. (i.e., the location of the input buffer or assembly workshop), from the starting point of the transfer batch. To the finish line The path (which can be simplified to the starting point of the transfer batch) To the finish line Distance between The time information includes the issuance time of the transfer instruction. Expected travel time of transfer equipment and expected completion time , .

[0197] The number of times a transfer device is started can be calculated based on the batch-transfer device-path combination instructions. and expected completion time of the order : , ,in, Indicates transfer equipment Startup count, , respectively workpiece The start time and required processing time for the final processing stage. For workpiece The cumulative transit time is calculated using the following formula: , This indicates the arrival at the processing stage from the buffer zone. The transfer distance, For transfer equipment The speed of transportation , These refer to the loading and unloading times, respectively.

[0198] Specifically, the deep reinforcement learning decision-making model is a dual deep Q-network (DDQN) model. DDQN has a policy network and a target network. It selects the optimal action by maximizing the cumulative reward. The parameters of the policy network are periodically copied from the target network. A memory pool is also set up to store the agent's historical experience (state, action, reward, new state) of its interaction with the environment. This memory pool is used for random sampling during training to break data correlation. The training steps of the deep reinforcement learning model include:

[0199] First, define the action space and actions. The action space is obtained by encoding the global state vector into a vector representation that can be processed by the neural network through a state encoder. The order type encoding in the global state vector is key input data to ensure that the agent can perceive and respond to dynamic order disturbances. The action space contains four scheduling rules. Initialize the execution network, the target network, and the experience replay pool.

[0200] Then, observe the global state vector at the current moment, according to The strategy selects one of four scheduling rules as the current action output, generates a batch-transfer device-path combination instruction based on the current action, obtains the reward R(t), and observes the global state vector at the next time step; the global state vector at the current time step, the current action, the reward, and the global state vector at the next time step are formed into a quadruple and stored in the experience replay pool.

[0201] Subsequently, a small batch of samples is randomly sampled from the experience replay pool. The target Q-value and the minimization loss function are calculated using the mean squared error as the loss function. Based on the loss function, the execution network parameters are updated using the standard DDQN update method. The policy network parameters are copied to the target network every fixed training step c steps through the Adam optimizer. Specifically, the target Q-value is calculated using the standard DDQN Q-value calculation method, and the mean squared error is used as the loss function.

[0202] Finally, repeat the above steps until the loss function converges.

[0203] Specifically, to guide the agent to achieve multi-objective optimization, a composite reward function is constructed based on total delay penalty, total energy consumption penalty, constraint penalty, and machine load reward. The parameters of the deep reinforcement learning model are trained based on this composite reward function. The expression of the composite reward function is as follows:

[0204] ;

[0205] ;

[0206] ;

[0207] ;

[0208] ;

[0209] ;

[0210] ;

[0211] ;

[0212] In the above formula, for Momentary rewards; , , , They are respectively Total delay penalty, total energy consumption penalty, constraint penalty, and machine load reward at any given time; , , , These are the weighting coefficients for total delay penalty, total energy consumption penalty, constraint penalty, and machine load reward, respectively. For machines Weighting coefficients; This is the penalty coefficient for delays; This is the total energy consumption penalty coefficient; for The total energy consumption increment of the transfer system at each moment is used to represent the time from the previous decision moment to the current decision moment (i.e., Between each decision point, the total energy consumption increment of the transfer system is caused by the execution of the batch-vehicle-route combination instruction of the previous decision point. By minimizing the total energy consumption increment of the transfer system at each decision point, the final total energy consumption of the transfer system can be minimized. , Transfer equipment Energy consumption per unit time under load operation and no-load operation; , Transfer equipment The time functions corresponding to load operation and no-load operation; For transfer equipment Startup energy consumption; For transfer equipment Number of times it is started; For transfer equipment Total runtime; To constrain the penalty coefficient for violations, a positive value is used; This is the penalty coefficient for overloaded units, which is generally taken as a positive value. If there is no overload, the value is 0. For transfer equipment exist The total number of workpieces being carried at any given time; This refers to the maximum number of workpieces that the transfer equipment can carry. Indicates the processing stage Average machine load rate; for Time processing stage Machine load rate; For processing stage The number of machines in the system; for Time processing stage The machine in The number of workpieces that have been processed; This represents the total number of transfer devices; Total number of orders; Indicates the machine number index; Indicates the index of the processing stage; This indicates the index of the transfer equipment.

[0213] This composite reward function, through the introduction of total delay penalties and total energy consumption penalties, guides the model to simultaneously reduce order delays and energy consumption during training, avoiding the unintended consequences of optimizing a single objective. The introduction of constraint penalties effectively prevents illegal scheduling behavior, ensuring the feasibility and compliance of the scheduling scheme. The introduction of machine load rewards promotes balanced utilization of production resources, reducing equipment idleness or overload. This multi-dimensional reward mechanism enables the deep reinforcement learning model to fully perceive the comprehensive impact of scheduling decisions, making the order allocation decisions output by the trained model more closely aligned with the complex needs of actual production scenarios. Specifically, when facing urgent orders, the model can reduce delay penalties while simultaneously considering energy consumption and load balancing, avoiding energy waste or equipment overload caused by prioritizing timeliness. Compared to a single reward mechanism, this composite reward function trains a more robust model with more comprehensive optimization effects, providing efficient, energy-saving, and stable scheduling support for the system.

[0214] Performance verification:

[0215] To verify the effectiveness of the method described in this invention, a multi-stage hybrid assembly line factory of an automobile wheel hub company was used as an example. The method described in this invention was used to generate a buffer workpiece transfer strategy and to dynamically allocate orders. The method described in this invention was compared and analyzed with four other algorithms, including: Non-dominated sorting genetic algorithm (NSGA-II), multi-objective artificial bee colony algorithm (MOABC), standard DQN model, and proximal strategy optimization algorithm (PPO).

[0216] There are five wheel hub models, L1-L5. Specific dimensional parameters and production batches for each workpiece are shown in Table 1. Wheel hub production process information is shown in Table 2. The machining process for automotive wheels begins with the uncoiling of hot-rolled steel sheets and ends with the semi-finished product being stored in the warehouse, comprising a total of 10 processes. Seven of these processes are completed by mechanical equipment, while the other three (grinding, cleaning, and packaging) are completed manually. To simplify the model, this example only sets the first seven machined processes as the machining stage, and the remaining three manual processes as the assembly stage, ignoring the time required for manual loading / unloading and mold changes. The processing time and transfer power for each process are shown in Table 2.

[0217] Algorithm parameter settings: The population size of NSGA-II and MOABC is 200, the crossover probability is 1.0, the mutation probability is 0.3, and the elite set size is 200. In the method described in this invention, the neural network input layer dimension of DDQN, standard DQN, and PPO is 6, the hidden layer is 128, the output layer is 4, the learning rate is 0.0005, the discount rate is 0.95, the memory recycling pool size is 1000, and the optimizer is Adam. Each algorithm is run independently 20 times, and the average value of the obtained multi-objective solution set is taken as the final solution.

[0218] Table 1. Specific dimensional parameters and production batches for wheel models L1-L5:

[0219] ;

[0220] Table 2 Wheel hub production process information:

[0221] ;

[0222] Test results:

[0223] 1. In the method described in this invention, DDQN training stops after 10,000 steps, and its reward value is as follows: Figure 3 As shown, the reward value increases rapidly in the first 2000 steps, then gradually stabilizes. The total delay period is as follows: Figure 4 As shown, the descent rate is stable during training. Its total energy consumption is as follows: Figure 5 As shown, the number of steps gradually stabilized after rapidly decreasing to 6000 during training.

[0224] 2. To evaluate the performance of different algorithms, the final solutions obtained by all algorithms were normalized, and the inverse generational distance (IGD) was selected as the evaluation metric. The calculated IGD values ​​of the five algorithms are shown in Table 3.

[0225] Table 3. IGD value calculation results for different algorithms:

[0226] ;

[0227] Table 3 shows that the IGD value of the method proposed in this invention is smaller than that of the other four algorithms. This indicates that the better the convergence of the method proposed in this invention, the closer the final solution obtained is to the true and optimal Pareto front.

[0228] The objective function values ​​corresponding to the final solutions obtained by the five algorithms are shown in Table 4:

[0229] Table 4. Objective function values ​​corresponding to the final solutions obtained by different algorithms:

[0230] ;

[0231] Table 4 shows the total order delay of the method proposed in this invention ( ) and total energy consumption of the transfer system ( The results are all smaller than the other four algorithms, which shows that the method proposed in this invention can further optimize the total order delay and reduce the total energy consumption of the transfer system.

[0232] The final solutions obtained by the 5 algorithms are as follows Figure 6 As shown, the final solution front obtained by the method proposed in this invention is closer to the origin of the coordinate system, indicating that the method proposed in this invention is closer to the true solution front than the other four algorithms.

[0233] Example 2:

[0234] See Figure 7 A buffer-based workpiece batch transfer and order allocation collaborative optimization system includes an order and process splitting module, a transfer batch generation module, and an order allocation module. The order and process splitting module is used to split orders to obtain workpiece sets, and to sequentially split the workpiece processing flow to obtain processing stage sets. A buffer is set in front of each machine in each processing stage. Specifically, the order and process splitting module is used to execute S1 in Embodiment 1, which will not be elaborated here. The transfer batch generation module is used to perform real-time batch clustering of the workpieces to be transferred in the buffer based on attribute correlation to generate transfer batches, and to issue transfer orders using a pulse triggering mechanism with dual threshold control. The instruction is to transport the workpieces in the transfer batch from the current buffer to the next processing stage via the transfer equipment; specifically, the transfer batch generation module is used to execute S2 in Embodiment 1, which will not be repeated here; the order allocation module is used to construct a deep reinforcement learning model. At each pulse triggering time, a global state vector is constructed based on all transfer batches and order types that have issued transfer instructions in the system. The global state vector is input into the deep reinforcement learning model to output the order allocation decision. Finally, a batch-transfer equipment-path combination instruction is generated based on the order allocation decision. Specifically, the order allocation module is used to execute S3 in Embodiment 1, which will not be repeated here.

[0235] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program goods. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program goods embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0236] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0237] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0238] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0239] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A method for collaborative optimization of workpiece batch transfer and order allocation in a buffer zone, characterized in that: The method includes: S1. Split the order to get a set of workpieces, and then split the workpiece processing flow in sequence to get a set of processing stages. A buffer is set in front of each machine in each processing stage. S2. Perform real-time batch clustering of the workpieces to be transferred in the buffer based on attribute correlation to generate transfer batches, and issue transfer instructions using a pulse triggering mechanism with dual threshold control to transport the workpieces in the transfer batch from the current buffer to the next processing stage through the transfer equipment. S3. Construct a deep reinforcement learning model. At each pulse trigger time, construct a global state vector based on all transfer batches and order types that have issued transfer instructions in the system. Input the global state vector into the deep reinforcement learning model to output order allocation decisions. Finally, generate batch-transfer equipment-path combination instructions based on the order allocation decisions.

2. The method for collaborative optimization of buffer workpiece batch transfer and order allocation according to claim 1, characterized in that: In S3, the deep reinforcement learning model selects one scheduling rule from a variety of preset scheduling rules, allocates orders based on the selected scheduling rule and the order type in the global state vector, obtains the order allocation decision, and then constructs an optimization scheduling problem with the objective of minimizing the total order delay and the total energy consumption of the transfer system. Based on the order allocation decision and the global state vector, the optimization scheduling problem is solved to generate batch-vehicle-path combination instructions.

3. The method for collaborative optimization of buffer workpiece batch transfer and order allocation according to claim 1 or 2, characterized in that: In S2, the real-time batch clustering includes: The attribute correlation of workpiece pairs is calculated based on material correlation, process correlation, and time-domain correlation. Workpiece pairs with attribute correlation higher than a preset correlation threshold are aggregated into the same transfer batch. The functional expression of the attribute correlation is as follows: ; ; ; ; In the above formula, For attribute relevance; , , respectively workpiece pair Material-related, process-related, and time-related; , , These are material weights, process weights, and time-domain weights, respectively. , These are material weights and size weights, respectively. Indicates workpiece With workpiece Material compatibility, , respectively workpiece With workpiece Material code, Indicates workpiece With workpiece Size similarity; , respectively workpiece , workpiece Dimensional values; , Representing the workpieces respectively , workpiece A set of processes; Indicates workpiece With workpiece The heat treatment similarity; , respectively workpiece , workpiece Delivery time; Indicates the time scaling factor; Indicates priority similarity.

4. A method for collaborative optimization of buffer workpiece batch transfer and order allocation according to claim 1 or 2, characterized in that: In S2, the pulse triggering mechanism of dual threshold control means that at each pulse triggering moment, if the number of workpieces in the transfer batch reaches the preset lower limit of the number of workpieces that the transfer equipment can carry, or if the waiting time of the transfer batch in the buffer reaches the preset maximum waiting time, then a transfer instruction is issued for the transfer batch.

5. A method for collaborative optimization of buffer workpiece batch transfer and order allocation according to claim 1 or 2, characterized in that: In S3, the global state vector includes: workpiece completion rate, average machine load rate, transfer equipment load rate, batch merging quality, and order type code; The expression for the workpiece completion rate is: ; In the above formula, Indicates the processing stage The completion rate of the workpiece; For processing stage Workpiece is being processed. The machine number; For processing stage The number of machines; Indicates the machine number index; For workpiece During the processing stage China Machinery Processing time; Indicates workpiece Whether to allocate to the processing stage China Machinery Decision variables on the workpiece Assigned to the processing stage China Machinery The value is 1 when processing is performed on the workpiece, and 0 otherwise; It refers to orders The first in One workpiece; The expression for the average machine load rate is: ; ; In the above formula, Indicates the processing stage Average machine load rate; for Time processing stage The machine in Load rate; for Time processing stage The machine in The number of workpieces that have been processed; For processing stage The time when the machine starts processing Duration of the moment; The expression for the load rate of the transfer equipment is: ; In the above formula, for The load rate of the transfer equipment at any given time; For transfer equipment exist The number of workpieces being carried at any given time; Index for transfer equipment numbers; This represents the total number of transfer devices; This refers to the maximum number of workpieces that the transfer equipment can carry. The expression for the batch merging quality is: ; In the above formula, For batch merging quality; This represents the total number of shipments. For transshipment batches Average correlation of attributes; The order type coding refers to the use of one-hot coding to represent order types. Order types include ordinary orders, rush orders, and orders with varying scales. Rush orders and orders with varying scales are dynamic orders.

6. The method for collaborative optimization of buffer workpiece batch transfer and order allocation according to claim 2, characterized in that: The scheduling rules include: Scheduling rule 1: When the order type is an expedited order or an order with a change in size, the order will be assigned to the machine with the lowest average machine load rate. When the order type is a regular order, the order will be assigned to the machine with the lowest workpiece completion rate; Scheduling rule 2: When the order type is an expedited order, calculate the change in the standard deviation of the machine load rate before and after assigning the order to a machine in the current processing stage, and assign the order to the machine with the smallest change in standard deviation. When the order type is a regular order or a variable size order, the order will be assigned to the machine with the earliest expected completion time. Scheduling rule 3: When the order type is expedited, the order will be assigned to the machine with the lowest load rate on the transfer equipment; When the order type is a regular order or a variable size order, the order will be assigned to machines other than those with the highest and lowest load rates on the transfer equipment. Scheduling rule 4: When the order type is an expedited order or an order with a change in size, the order will be assigned to the machine with the highest batch merging quality: ,in This indicates the machine that belongs to the current processing step. The total number of transit batches waiting to be transited in the next buffer; For transshipment batches Average correlation of attributes; When the order type is a regular order, the order will be randomly assigned to any machine other than the one with the highest and lowest batch quality. .

7. The method for collaborative optimization of buffer workpiece batch transfer and order allocation according to claim 2, characterized in that: The objective function of the optimization scheduling problem is expressed as: ; ; In the above formula, , These are the first objective function and the second objective function, respectively. , These represent the total order delay and the total energy consumption of the transit system, respectively. , Orders Delivery time and expected completion time; Indicates the total number of orders; Index the number of transfer equipment in the transfer system; For transfer equipment Total runtime; , Transfer equipment Energy consumption per unit time under load operation and no-load operation; , Transfer equipment The time functions corresponding to load operation and no-load operation; For transfer equipment Startup energy consumption; For transfer equipment Number of times it is started; The constraints of the optimization scheduling problem include constraints on the number of transfer batches, transfer time, and processing time of the workpiece at each processing stage.

8. A method for collaborative optimization of buffer workpiece batch transfer and order allocation according to claim 1 or 2, characterized in that: In S3, a composite reward function is constructed based on the total delay penalty, total energy consumption penalty, constraint penalty, and machine load reward. The parameters of the deep reinforcement learning model are trained based on this composite reward function. The expression of the composite reward function is as follows: ; ; ; ; ; ; ; ; In the above formula, for Momentary rewards; , , , They are respectively Total delay penalty, total energy consumption penalty, constraint penalty, and machine load reward at any given time; , , , These are the weighting coefficients for total delay penalty, total energy consumption penalty, constraint penalty, and machine load reward, respectively. For machines Weighting coefficients; This is the penalty coefficient for delays; This is the total energy consumption penalty coefficient; for Total energy consumption of the transportation system at any given time; , Transfer equipment Energy consumption per unit time under load operation and no-load operation; , Transfer equipment The time functions corresponding to load operation and no-load operation; For transfer equipment Startup energy consumption; For transfer equipment Number of times it is started; For transfer equipment Total runtime; To constrain the penalty coefficient for violations; The overload unit penalty coefficient; For transfer equipment exist The total number of workpieces being carried at any given time; This refers to the maximum number of workpieces that the transfer equipment can carry. Indicates the processing stage Average machine load rate; for Time processing stage Machine load rate; For processing stage The number of machines in the system; for Time processing stage The machine in The number of workpieces that have been processed; This represents the total number of transfer devices; Total number of orders; Indicates the machine number index; Indicates the index of the processing stage; This indicates the index of the transfer equipment.

9. A method for collaborative optimization of buffer workpiece batch transfer and order allocation according to claim 1 or 2, characterized in that: In S3, the batch-vehicle-path combination instruction includes the following triplet information: batch information, path information, and time information, wherein the batch information includes the set of workpieces contained in the transfer batch, and the path information includes the starting position of the transfer batch. Destination of the transshipment batch From the starting point of the transfer batch To the finish line The path, the time information includes the time when the transfer instruction was issued. Expected travel time of transfer equipment and expected completion time , .

10. A buffer zone workpiece batch transfer and order allocation collaborative optimization system, characterized in that: The system includes: The order and process splitting module is used to split orders to obtain a set of workpieces, and to split the workpiece processing flow in sequence to obtain a set of processing stages. A buffer is set in front of each machine in each processing stage. The transfer batch generation module is used to perform real-time batch clustering of workpieces to be transferred in the buffer based on attribute correlation to generate transfer batches, and to issue transfer instructions using a pulse triggering mechanism with dual threshold control so that the workpieces in the transfer batch can be transported from the current buffer to the next processing stage through the transfer equipment. The order allocation module is used to build a deep reinforcement learning model. At each pulse trigger time, a global state vector is constructed based on all transfer batches and order types that have issued transfer instructions in the system. The global state vector is input into the deep reinforcement learning model to output the order allocation decision. Finally, a batch-transfer equipment-path combination instruction is generated based on the order allocation decision.