Collaborative Optimization Method for Hybrid Flow Shop Scheduling Problem with Limited Transportation Resources
Through structure-aware heterogeneous graph neural network and composite scheduling action selection network, combined with a proximal strategy optimization algorithm, the problem of difficult to effectively coordinate the optimization of mixed flow workshop scheduling with limited transportation resources in the existing technology is solved, and efficient and adaptive scheduling optimization is achieved.
Patent Information
- Application Number
- CN202410956753.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-17
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-07-17
AI Technical Summary
It is difficult for the prior art to effectively coordinate the optimization of mixed flow workshop scheduling problems with limited transportation resources, especially when dealing with small-scale and super-large-scale scheduling problems, it is difficult to ensure the optimization and real-timeness of scheduling results.
Structurally perceived heterogeneous graph neural network and composite scheduling action selection network are adopted. By constructing heterogeneous graph models and graph neural networks, combined with proximal strategy optimization algorithms, end-to-end scheduling state feature characterization and scheduling action selection are realized, avoiding the limitations of manual intervention and traditional methods.
It realizes real-time interaction and optimization of scheduling actions without a complete mathematical model, improves adaptability and scheduling efficiency, and can effectively solve small-scale and super-large-scale scheduling problems and obtains better scheduling results.
Smart Images

Figure CN118966625B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of workshop scheduling, and particularly to a collaborative optimization method for the hybrid flow shop scheduling problem with limited transportation resources. Background Art
[0002] The hybrid flow shop scheduling problem (HFSP), as a complex combinatorial optimization problem with NP-Hard characteristics, widely exists in manufacturing fields such as steelmaking, logistics, and assembly. In previous studies on HFSP, only the selection of the most suitable parallel machine for each workpiece at each stage and the sequencing of the processing order of the workpieces on each parallel machine were considered, without considering the transportation time between adjacent processes of each workpiece, or regarding the transportation time of a workpiece between adjacent stages as part of the processing time of the workpiece. In an actual hybrid flow shop, in order to save production costs and improve production efficiency, the handling of workpieces is usually completed by an automated guided vehicle (AGV). That is, the actual HFSP needs to consider not only the processing of workpieces on parallel machines but also the transportation of workpieces by AGV. In addition, limited by factors such as the spatial layout of the actual hybrid flow shop, the workpiece transportation resources mainly based on AGV are often limited, that is, it is impossible to ensure that each workpiece can obtain the timely transportation response of AGV in adjacent processes. Therefore, proposing a collaborative optimization method for the hybrid flow shop scheduling problem with finite transportation resources (HFSP-FTR), namely the HFSP-FTR collaborative optimization method, has become a key problem that urgently needs to be solved.
[0003] In recent years, most of the research on solving HFSP-FTR has mainly focused on mathematical programming methods, heuristic methods, intelligent optimization methods, and a small number of machine learning methods. Mathematical programming methods can find exact solutions, but they require a high degree of accuracy for the model, have a long solution time, and are difficult to apply to large-scale scheduling problems. Heuristic methods are mainly based on scheduling rules and can quickly find a solution, but the quality of this solution is low and the optimality cannot be guaranteed. Intelligent optimization methods, represented by genetic algorithms, have the advantage of being able to find sub-optimal solutions close to the optimal solution, but the cost is that they require a large number of iterative calculation steps, and they cannot adaptively adjust in the face of disturbances during the scheduling process and need to re-execute the solution process of the algorithm. Machine learning methods, represented by reinforcement learning algorithms, can interact with the environment in real time to learn scheduling strategies for taking the best scheduling actions at a certain scheduling moment to adapt to scheduling problems of various complexities. However, the performance of reinforcement learning algorithms is limited by the capacity of the state space, prone to the problem of "curse of dimensionality", and the selection process of scheduling actions is relatively dependent on manual experience and cannot achieve the "end-to-end" property without any manual intervention. Summary of the Invention
[0004] In view of the above deficiencies of the prior art, the purpose of the present invention is to provide a collaborative optimization method for the hybrid flow shop scheduling problem with limited transportation resources, which can represent the scheduling state characteristics at the current moment in an end-to-end manner without human intervention. In addition, it can solve small-scale and ultra-large-scale scheduling problems and obtain better scheduling results.
[0005] The technical solution of the present invention is: a collaborative optimization method for the hybrid flow shop scheduling problem with limited transportation resources, the method comprising the following steps:
[0006] Step 1: Obtain the HFSP-FTR instance and its production and processing information;
[0007] Step 2: According to the production and processing information of the HFSP-FTR instance, model the HFSP-FTR instance as a heterogeneous graph model simultaneously including process nodes, parallel machine nodes, and AGV nodes, called the HFSP-FTR instance heterogeneous graph model;
[0008] Step 3: According to the existing attention mechanism theory and graph neural network theory, construct a structure-aware heterogeneous graph neural network, and use the structure-aware heterogeneous graph neural network to obtain multi-dimensional feature vectors of the HFSP-FTR instance heterogeneous graph model;
[0009] Step 4: Construct a composite scheduling action selection network ASN t , use the multi-dimensional feature vectors of the instance heterogeneous graph model as the input of ASN t , and calculate each composite scheduling action through ASN t The action selection probability P (iku,h) ;
[0010] Step 5: Input the action selection probability obtained in Step 4 into the proximal policy optimization algorithm to obtain the optimal scheduling action selection strategy π t .
[0011] Furthermore, the heterogeneous graph model of the HFSP-FTR instance is represented as where:
[0012] (1) is the set of process nodes, which contains all the process nodes in the heterogeneous graph model of the HFSP-FTR instance;
[0013] (2) is the set of parallel machine nodes, which contains all the parallel machine nodes in the heterogeneous graph model of the HFSP-FTR instance;
[0014] (3) is the set of AGV nodes, which contains all the AGV nodes in the heterogeneous graph model of the HFSP-FTR instance;
[0015] (4) is the set of all undirected edges connecting process node O (i,h) and parallel machine node M (k,h) in the heterogeneous graph model of the HFSP-FTR instance;
[0016] (5) is the set of all undirected edges connecting process node O (i,j) and AGV node A (u,h) in the heterogeneous graph model of the HFSP-FTR instance;
[0017] (6) is the set of directed edges, which contains the directed edges used to represent the workpiece processing sequence in the heterogeneous graph model of the HFSP-FTR instance;
[0018] where O (i,h) represents the h-th process node in workpiece J i ; M (k,h) represents the k-th parallel machine node in stage s h ; A (u,h) represents the u-th AGV in stage s h .
[0019] Furthermore, the establishment process of the heterogeneous graph model of the HFSP-FTR instance should satisfy the following process constraints:
[0020] (1) Each workpiece can only be processed by one parallel machine in each stage;
[0021] (2) Each workpiece must start processing in the current stage after the previous stage is completed;
[0022] (3) For two workpieces processed on the same parallel machine, only after one workpiece is completed, can the other workpiece start processing;
[0023] (4) All AGVs are the same and travel at a constant speed;
[0024] (5) Once an AGV starts working, it is not allowed to be interrupted midway;
[0025] (6) At the same time, an AGV can only transport one workpiece.
[0026] Furthermore, the structure-aware heterogeneous graph neural network includes 1 global graph embedding network and 3 parallel embedding sub-networks; the 3 parallel embedding sub-networks are the operation node embedding sub-network SEN O , the parallel machine node embedding sub-network SEN M and the AGV node embedding sub-network SEN A ; the global graph embedding network is used to calculate the graph embedding of the entire HFSP-FTR instance heterogeneous graph model; the operation node embedding sub-network SEN O is used to calculate the node embedding of the operation node O (i,h) ; the parallel machine node embedding sub-network SEN M is used to calculate the node embedding of the parallel machine node M (k,h) ; the AGV node embedding sub-network SEN A is used to calculate the node embedding of the AGV node A (u,k) .
[0027] Compared with the prior art, the present invention has the following beneficial effects:
[0028] (1) The present invention adopts the actor-critic algorithm in the proximal policy optimization algorithm as the training algorithm, which can interact with the scheduling environment in real time without a complete mathematical model, and fully explore the optimal scheduling actions in each scheduling environment, improving the adaptability of the present invention in solving more scheduling problems;
[0029] (2) The present invention uses a heterogeneous graph neural network to model each HFSP-FTR instance as a heterogeneous graph model, and uses the workshop production and processing information as scheduling features, avoiding the inaccuracy of manually designing or selecting scheduling state features in traditional reinforcement learning, and thus realizing the "end-to-end" feature; in addition, this heterogeneous graph model can be compatible with both small-scale and large-scale HFSP-FTR scheduling problems, having good technical architecture flexibility;
[0030] (3) Considering that the training process of the proximal policy optimization algorithm is offline, the present invention can trade off between the quality of the scheduling solution and the solving time, and at the same time solves the drawbacks of traditional heuristic methods and intelligent optimization methods. Description of the Drawings
[0031] Figure 1 is a flowchart of the collaborative optimization method for the hybrid flow shop scheduling problem with limited transportation resources in this embodiment;
[0032] Figure 2 is a schematic diagram of a heterogeneous graph model of an HFSP-FTR instance in this embodiment
[0033] Figure 3 is the architecture schematic diagram of the heterogeneous graph neural network in this embodiment;
[0034] Figure 4 is the convergence curve graph of the loss function in the training process for solving HFSP-FTR instances provided in this embodiment;
[0035] Figure 5 is the convergence curve graph of the return function value for solving HFSP-FTR instances provided in this embodiment. Detailed Embodiment
[0036] To facilitate the understanding of the present application, the present application will be described more comprehensively below with reference to the relevant drawings.
[0037] As Figure 1 shown, the collaborative optimization method for the hybrid flow shop scheduling problem with limited transportation resources in this embodiment is abbreviated as the HFSP-FTR collaborative optimization method, and includes the following steps:
[0038] Step 1: Obtain HFSP-FTR instances and obtain the production and processing information of each HFSP-FTR instance, construct an HFSP-FTR instance dataset, and randomly divide the HFSP-FTR instance dataset into a training set and a test set;
[0039] In this embodiment, the number of HFSP-FTR instances in the training set is 100.
[0040] The production and processing information of each HFSP-FTR (Hybrid Flow Shop Scheduling Problem with Finite Transportation Resourses) instance obtained in this embodiment is shown in Table 1.
[0041] Table 1 Production and processing information of HFSP-FTR instances
[0042]
[0043]
[0044] Among them, for |M h |, |O i |, M (k) , A (u) , and TT i(h+1)ku The supplementary explanations are as follows:
[0045] (1) The elements included in |M h | are {M (1,h) , M (2 , h),..., M (g,h) ,..., M (k,j) ,..., M (K,h)} ∈ |M h |. In the heterogeneous graph model and heterogeneous graph neural network of the HFSP-FTR instance established in subsequent steps, M (k,h) represents the k-th parallel machine node in stage s h ;
[0046] (2) The elements included in |O i | are {O (i,1) , O (i,2) ,..., O (i,h) ,..., O (i,H)} ∈ |O i |. In the heterogeneous graph model and heterogeneous graph neural network of the HFSP-FTR instance established subsequently, O (i,h) represents the h-th process node in workpiece J i ;
[0047] (3) M (k) refers to any parallel machine with index k in an HFSP-FTR instance. In the heterogeneous graph model and heterogeneous graph neural network of the HFSP-FTR instance established in subsequent steps, M (k)Only used to represent any parallel machine node with index k among all parallel machine nodes in an HFSP-FTR instance (regardless of the stage it is in), which is convenient for describing the relevant model;
[0048] (4)A (u) Refers to any parallel machine with index u in an HFSP-FTR instance. In the heterogeneous graph model and heterogeneous graph neural network of the HFSP-FTR instance established in the subsequent steps, A (u) Only used to represent any AGV node with index u among all AGV nodes in an HFSP-FTR instance (regardless of the stage it is in), which is convenient for describing the relevant model;
[0049] (5)TT i(h+1)ku Consists of the sum of the following two parts:
[0050] 1)A (u,h) Obtain the time when O that has been processed is completed; ih ;
[0051] 2)A (u,h) Transport O (i,h) to a certain parallel machine in the next stage s h+1 The time on.
[0052] Table 2 shows the data format of an HFSP-FTR instance of "3 jobs - 5 stages".
[0053] Table 2 Schematic of the data format of the HFSP-FTR instance of "3 jobs - 5 stages"
[0054]
[0055] As shown in Table 2, there are a total of 3 rows of data, and each row represents the processing and transportation information of a job, that is, the processing and transportation information of J1, J2, and J3. The operation sets of each job are respectively:
[0056] (1){O 11 ,O 12 ,O 13 ,O 14 ,O 15}∈O1;
[0057] (2){O 21 ,O 22 ,O 23 ,O 24 ,O 25}∈O2;
[0058] (3){O 31 ,O 32 ,O33 , O 34 , O 35}, ∈ O3.
[0059] Each workpiece needs to be processed through 5 stages in sequence according to the process order, and each workpiece needs to be transported by a certain AGV between every two adjacent stages. Taking the data in the first row of the HFSP-FTR example shown in Table 2 as an example, this embodiment details the processing and transportation information of workpiece J1. By analogy, those skilled in the art can understand the processing and transportation information of J2 and J3, which will not be elaborated here.
[0060] In addition, the encoding of all parallel machines in Table 2 adopts the sequential encoding form from beginning to end, without deliberately emphasizing the stage where each parallel machine is located. This encoding method is only used to represent the format of the example for the operation of computer programs, and those skilled in the art should distinguish it from M (k,h) for differentiation.
[0061] The processing and transportation information of J1 is as follows:
[0062] (1) The parallel machines capable of processing operation O (1,1) are M (1) , M (2) and M (3) , and the required unit processing times are 19, 19, 19 respectively; the AGVs capable of transporting operation O (1,1) are A (1) , A (2) and A (3) , and the required unit transportation times are 4, 5, 6 respectively;
[0063] (2) The parallel machines capable of processing operation O (1,2) are M (4) , M (5) and M (6) , and the required unit processing times are 15, 15, 15 respectively; the AGVs capable of transporting operation O (1,2) are A (4) , A (5) and A (6) , and the required unit transportation times are 6, 6, 7 respectively;
[0064] (3) The parallel machine capable of processing operation O (1,3) is M (7) , and the required unit processing time is 14; the AGV capable of transporting operation O 13 is A (7) , and the required unit transportation time is 5;
[0065] (4) The parallel machines capable of processing operation O (1,4) are M (8) , M(9) and M (10) The required processing time units are 5, 5, and 5 respectively; The AGVs that can transport process O (1,4) are A (8) 、A (9) and A (10) The required unit transportation times are 5, 4, and 7 respectively;
[0066] (5)The parallel machines that can process process O (1,5) are M (11) 、M (12) and M (13) The required processing time units are 18, 18, and 18 respectively; The AGVs that can transport process O (1,5) are A (11) 、A (12) and A (13) The required unit transportation times are 3, 3, and 5 respectively.
[0067] Step 2: According to the production and processing information of each HFSP-FTR instance, model each HFSP-FTR instance in the training set as a heterogeneous graph model that simultaneously includes process nodes, parallel machine nodes, and AGV nodes, called the HFSP-FTR instance heterogeneous graph model;
[0068] This embodiment represents the HFSP-FTR instance heterogeneous graph model as Specifically described as follows:
[0069] (1) is the set of process nodes, which contains all the process nodes in the HFSP-FTR instance heterogeneous graph model;
[0070] (2) is the set of parallel machine nodes, which contains all the parallel machine nodes in the HFSP-FTR instance heterogeneous graph model;
[0071] (3) is the set of AGV nodes, which contains all the AGV nodes in the HFSP-FTR instance heterogeneous graph model;
[0072] (4) is the set of all undirected edges connecting process node O (i,h) and parallel machine node M (k,h) in the HFSP-FTR instance heterogeneous graph model;
[0073] (5) is the set of all undirected edges connecting process node O (i,h) and AGV node A (u,h) in the HFSP-FTR instance heterogeneous graph model;
[0074] (6) is a set of directed edges, which contains the directed edges used to represent the processing sequence of workpieces in the heterogeneous graph model of the HFSP-FTR instance.
[0075] Figure 2 The specific structure of the heterogeneous graph model of the HFSP-FTR instance established in this embodiment is shown, and the detailed description of each component is as follows:
[0076] (1) "Start" and "End" are artificially set virtual nodes, which are used to represent the complete heterogeneous graph model. They do not participate in any operations in the heterogeneous graph neural network mentioned later;
[0077] (2) The scale of this HFSP-FTR instance heterogeneous graph model consists of 2 workpieces J1 and J2, and 3 stages s1, s2, and s3. The number of parallel machines in each stage where each process of each workpiece is located is 2, 2, 2 respectively, and the number of AGVs is 2, 2, 2. Taking the information of workpiece J1 in the HFSP-FTR instance heterogeneous graph model as an example: Workpiece J1 has 3 processes, which respectively correspond to 3 process nodes: O (1,1) 、O (1,2) and O (1,3) . In addition,
[0078] 1) The parallel machine nodes that can process O (1,1) are M (1,1) and M (2,1) , and they are connected to O (1,1) by using 2 undirected edges in ; The AGV nodes that can transport O (1,1) are A (1,1) , and they are connected to O (1,1) by using 2 undirected edges in ;
[0079] 2) The parallel machine nodes that can process O (1,2) are M (1,2) and M (2,2) , and they are connected to O (1,2) by using 2 undirected edges in ; The AGV nodes that can transport O (1,2) are A (1,2) and A (2,2) , and they are connected to O (1,2) by using 2 undirected edges in ;
[0080] 3) The parallel machine nodes that can process O (1,3) are M (1,3) and M (2,3) , and they are connected to O(1,3) are connected by using two undirected edges in; can transport O (1,3) The AGV node of is A (1,3) and A (2,3) and they are connected with O (1,3) are connected by using two undirected edges in;
[0081] 4) "Start", process node O (1,1) , process node O (1,2) , process node O (1,3) and "End" are connected by directed edges in, forming the process sequence of workpiece J1.
[0082] By analogy, those skilled in the art can understand the information of J2 in the heterogeneous graph model of the HFSP - FTR instance, which will not be elaborated here.
[0083] In summary, "Start", "End", all process nodes, all parallel machine nodes, all AGV nodes, all undirected edges, and all directed edges together constitute the heterogeneous graph model of the HFSP - FTR instance.
[0084] It should be noted that in the hybrid flow shop scheduling problem with limited transportation resources, the establishment process of the HFSP - FTR instance heterogeneous graph model proposed in this embodiment should satisfy the following main process constraints:
[0085] (1) Each workpiece can only be processed by one parallel machine in that stage in each stage;
[0086] (2) Each workpiece must start the processing of the current stage after the processing of the previous stage is completed;
[0087] (3) For two workpieces processed on the same parallel machine, only after one workpiece is completed, the other workpiece starts to be processed;
[0088] (4) All AGVs are the same and travel at a constant speed;
[0089] (5) Once an AGV starts working, it is not allowed to be interrupted midway;
[0090] (6) At the same time, an AGV can only transport one workpiece.
[0091] Step 3: According to the existing attention mechanism theory and graph neural network theory, construct a Structure-aware Heterogeneous Graph Neural Network (SAHGNN) for extracting the scheduling knowledge contained in the heterogeneous graph model of each HFSP-FTR instance in the training set;
[0092] In this embodiment, as Figure 3 shown, the structure-aware heterogeneous graph neural network consists of 3 parallel embedding subnetworks decomposed based on the SAHGNN model and 1 full-graph embedding network. Among them, the 3 parallel embedding subnetworks decomposed based on the SAHGNN model are the operation node embedding subnetwork SEN O , the parallel machine node embedding subnetwork SEN M , and the AGV node embedding subnetwork SEN A , which are respectively used for parallel computing the node embeddings of operation nodes O (i,n) , parallel machine nodes M (k,h) , and AGV nodes A (u,h) ; the full-graph embedding network SEN G is then used to calculate the graph embedding of the heterogeneous graph model of the entire HFSP-FTR instance;
[0093] The specific implementation process of the SAHGNN described in this embodiment is as Figure 3 shown. Among them, node embeddings and graph embeddings are essentially multi-dimensional feature vectors, which both contain feature information in multiple dimensions. The prerequisite for using the 3 parallel embedding subnetworks decomposed based on the SAHGNN model in Step 3 is to assign original feature vectors to each type of node and edge to characterize the processing attributes of each type of node. The embedding subnetworks of each type of node perform layer-by-layer calculations on the original feature vectors of each type of node to obtain the node embeddings of each type of node, and further use them to obtain the optimal scheduling strategy. To further clearly and conveniently describe the details of each network, Table 3 shows various parameters in the SAHGNN described in this embodiment. For the convenience of description, all the explanations of the set symbols in Table 3 in this embodiment are based on the condition of "at scheduling time t", without additional declarations. In addition, to conveniently illustrate the embedding process of nodes in the SAHGNN, this embodiment only uses symbols to represent each type of node, without writing the full name of the node.
[0094] Table 3 Various symbols in SAHGNN
[0095]
[0096] Step 3.1: Define the operation node O (i,h) , the parallel machine node M (k,h)and the neighbor node set of AGV node A (u,h) ;
[0097] From the neighbor node sets of various nodes shown in Table 3, it can be seen that the SAHGNN proposed in this embodiment uses the information of the neighbor node sets of 3 nodes, namely O (i,h) , M (k,h) , and A (u,h) . It should be noted that all descriptions of the neighbor node set of a certain node in this step are based on the condition of "at scheduling time t". Unless necessary, this embodiment will not repeat the description.
[0098] Step 3.1.1: Define the neighbor nodes of O (i,h) to form the neighbor node set N (i,h) (O t ); (i,h) )
[0099] The elements in N t (O (i,h) ) are composed of the following nodes:
[0100] (1) The previous process node of O (i,h) : O (i,h-1) ;
[0101] (2) The next process node of O (i,h) : O (i,h+1) ;
[0102] (3) The own node of O (i,h) : O (i,h) ;
[0103] (4) The parallel machine node connected to O (i,h) : M (k,h) ;
[0104] (5) The AGV node connected to O (i,h) : A (u,h) .
[0105] In summary, the elements in N t (O (i,h) ) are shown in formula (1):
[0106] N t (O (i,h) ) = {O (i,h-1) , O (i,h+1) , O (i,h) , M (k,h) , A (u,h)} (1)
[0107] It should be noted that during the scheduling process, both O (i,h)The parallel machine scheduling process for processing on a parallel machine also needs to consider O (i,h) The AGV scheduling process transported by the AGV. Therefore, N t (O (i,h) ) can be further divided into the following two parts:
[0108] 1) N t (O (i,h) , M (k,h) ) : The set of parallel machine nodes in the set of neighbor nodes of O (i,h) ;
[0109] 2) N t (O (i,h) , A (u,h) ) : The set of AGV nodes in the set of neighbor nodes of O (i,h) ;
[0110] Therefore, this embodiment mainly uses N t (O (i,h) , M) and N t (O (i,h) , A) to accurately describe SEN M and SEN A in SAHGNN.
[0111] Step 3.1.2: Define the neighbor nodes of M (k,h) to form the set of neighbor nodes N (k,h) (M t ) of M (k,h) ;
[0112] The elements in N t (M (k,h) ) are composed of the following nodes:
[0113] (1) The own node of M (k,h) : M (k,j) ;
[0114] (2) All the operation nodes that M (k,h) can process in the current stage, that is, all the elements in the set N t (M (k,h) , O).
[0115] In summary, the elements in N t (M (k,h) ) are shown in formula (2):
[0116] N t (M (k,h) ) = {M (k,h) , N t (M (k,h) , O)} (2)
[0117] It should be noted that in the subsequent SEN M proposed, N t (M (k,h) , O) is used to replace N t (M (k,h) ) to accurately describe the neighbor node information aggregation process of M (k,h) in SEN M .
[0118] Step 3.1.3: Define the neighbor nodes of A (u,h) to form the neighbor node set N (u,h) (A t ) of A (u,h) ;
[0119] The elements in N t (A (u,h) ) are composed of the following nodes:
[0120] (1) The node itself of A (u,h) : A (u,h) ;
[0121] (2) All process nodes that A (u,h) can transport in the current stage, that is, all elements in the set Ny(A (u,h) , O).
[0122] In summary, the elements in N t (A (u,h) ) are shown in formula (3):
[0123] N t (A (u,h) ) = {A (u,h) , N t (A (u,h) , O)} (3)
[0124] It should be noted that in the subsequent SEN A proposed, N t (A (u,h) , O) is used to replace N t (M (k,h) ) to accurately describe the neighbor node information aggregation process of A (u,h) in SEN A .
[0125] Step 3.2: Assign a multi-dimensional original feature vector to and respectively;
[0126] Step 3.2.1: Assign for O (i,h)Allocate the original feature vector
[0127] Specifically, it includes the processing attributes of the following 7 dimensions:
[0128] (1) The number of neighbor parallel machine nodes of O: |N (i,h) (O t , M); (i,h)
[0129] (2) The number of neighbor AGV nodes of O: |N (i,h) (O t , A); (i,h)
[0130] (3) The number of unscheduled operations of O: NUSO (i,h) ; ih
[0131] (4) The completion time of O: C (i,h) ; (ihk)
[0132] (5) The estimated start processing time or actual start processing time of O: S (i,h) : (ihk)
[0133] (6) The sum of the processing time P (i,h) of O on M (k,h) and the transportation time by A (ihk) : P (u,h) + TT (ihk) ; i(h+1)ku
[0134] (7) Current status Status: At the scheduling moment t, if O (i,h) has been scheduled, it is 1; otherwise it is 0.
[0135] In summary, the original feature vector of O (i,h) is as shown in formula (4):
[0136]
[0137] Among them, for the scheduling moment t, the meaning of C (ihk) needs to be supplemented additionally, and the specific content is as follows:
[0138] (1) The index of the scheduling time mentioned in this embodiment is {1, 2,..., t, t + 1,..., T};
[0139] (2) For C (ihk) mentioned in this embodiment, it is necessary to consider that at the scheduling moment t, O(i,h) Has been scheduled and O (i,h) There are two cases of not being scheduled, which are described as follows:
[0140] 1) At the scheduling moment t, if O (i,h) has been scheduled, then its completion time C (ihk) is the value of the actual completion time in the scheduling process, and this value is automatically obtained during the scheduling process without human participation;
[0141] 2) If O (i,h) is not scheduled, then calculate its completion time C (ihk) according to formula (5):
[0142] C (ihk) = C (i(h-1)g) + P (ihk) (5)
[0143] Among them, C (i(h-1)h) is the completion time on the parallel machine node in the previous stage s (i,h) of the previous process O (i,h-1) of O h-1 . In summary,
[0144] Step 3.2.2: Assign the original feature vector to M (k,h) Specifically, it includes the processing attributes of the following 4 dimensions:
[0145] Specifically, it includes the processing attributes of the following 4 dimensions:
[0146] (1) Two-dimensional variable: I t (M (k,h) ). At the scheduling moment t, if M (k,j) is assigned a process node, I t (M (k,j) ) = 1, otherwise I t (M (k,h) ) = 0;
[0147] (2) The utilization rate of M (k,h) :
[0148] (3) The number of neighbor process nodes of M (k,h) : |N t (M (k,h) , O)|;
[0149] (4) The time when M (k,h) can start processing O (i,h) :
[0150] In summary, M(k,h) The original feature vector As shown in formula (6):
[0151]
[0152] Step 3.2.3: Assign the original feature vector to A (u,h) Assign the original feature vector
[0153] Specifically, it includes the processing attributes of the following 4 dimensions:
[0154] (1) Two-dimensional variable: I t (A (u,h) ) At scheduling time t, if the AGV node A (u,h) is assigned a process node, I t (A (u,h) ) = 1, otherwise I t (A (u,h) ) = 0;
[0155] (2) The utilization rate of A (u,h) :
[0156] (3) The number of neighbor process nodes of A (u,h) : |N t (A (u,h) , O)|;
[0157] (4) The time when A (u,h) can start transporting O (i,h) :
[0158] In summary, the original feature vector of A (u,h) is as shown in formula (7): As shown in formula (7):
[0159]
[0160] Step 3.2.4: Assign the original feature vector to the undirected edge e (i,h) between O (k,h) and M ihk Assign the original feature vector
[0161] Specifically, it includes the processing attributes of the following 1 dimension:
[0162] The processing time of O (i,h) on M (k,h) : P (ihk) .
[0163] In summary, e ihkThe original feature vector As shown in formula (8):
[0164]
[0165] Step 3.2.5: For O (i,h) and A (u,h) The undirected edge e connected between them ihu Allocate the original feature vector
[0166] Specifically, it includes the processing attributes of the following 1 dimension:
[0167] O (i,h) Is transported by A (u,h) The time required: TT i(h+1)ku .
[0168] In summary, the original feature vector of e ihu As shown in formula (9): As shown in formula (9):
[0169]
[0170] Step 3.3: Combine N≥2 attention layers to construct a parallel machine node embedding sub-network SEN M , and obtain the embedding of M M through SEN (k,h) of
[0171] Regarding this step 3.3 of this embodiment, it should be noted that:
[0172] (1) The parallel machine node embedding sub-network SEN M contains N≥2 attention layers (Attention Layer, AL). In this embodiment, N is determined to be 2 according to experience and experimental results. In the network structure of SEN M , the first AL takes the original feature vector of M (k,h) and the original feature vector μ of the extended process node O as inputs, and the embedding of M (ih,ihk) as the output of this AL. Since SEN (ih,ihk) contains multiple AL layers, the output of the first AL (k,h) is used as the input of the second AL, and the output of the second AL is and so on. The output of the l-1th AL M is used as the input of the lth AL, and the output of the lth AL is the embedding of M As the input of the second AL, the output of the second AL is and so on. The output of the l-1th AL is used as the input of the lth AL, and the output of the lth AL is the embedding of M (k,h) of Output of the Nth AL is M (k,h) Final embedding. In this embodiment, only the process of obtaining M at scheduling time t is described in detail (k,h) in SEN M the embedding in the first AL of The specific calculation process of. Those skilled in the art know that the calculation methods in other ALs of SEN M are the same as those in the first AL. Therefore, unless otherwise stated, this embodiment will not repeatedly emphasize "at scheduling time t" and "in the first AL of SEN M " or "the first AL of SEN M " in the following sub-steps
[0173] (2) The descriptions of some parameters have also been shown in Table 3 in this embodiment. Therefore, this embodiment will not repeat the description additionally in the following sub-steps
[0174] At scheduling time t, obtain M (k,h) in SEN M the embedding in the first AL of The specific calculation process includes the following steps
[0175] Step 3.3.1: Expand the process node O (i,h) and splice the process node O (i,h) with the undirected edge e ihk to obtain a new process node, called the expanded process node O (ih,ihk) ;
[0176] Step 3.3.2: Concatenate the original feature vector of the neighbor process node O (k,h) of M (i,h) with the original feature vector of e and e ihk to obtain the original feature vector μ of the expanded process node O (ih,ihk) , and use O (ih,ihk) to replace O (ih,ihk) as the new neighbor node in the neighbor node set of M (i,h) ; (k,h) ;
[0177] μ (ih,ihk) is shown in formula (10):
[0178]
[0179] where is the original feature vector of O (i,h) ; is the original feature vector of e ihkThe original feature vector; Is a splicing operation.
[0180] Step 3.3.3: Calculate M (k,h) And O (ih,ihk) The correlation coefficient calculated from the first AL of SEN between M
[0181] M (k,h) The correlation between M and the point O (ih,ihk) Depends on the query vector of M (k,h) And the key vector of O (ih,ihk)
[0182] This embodiment needs to briefly introduce the query vector, key vector, and value vector involved in SEN M In order to introduce the calculation process of SEN in more detail M
[0183] (1) Query vector: Abbreviated as q, That is, the query vector, which is used to obtain the correlation between a certain node and other nodes;
[0184] (2) Key vector: Abbreviated as q, That is, the key vector, which is used to calculate the similarity between the query vector and the value vector and measure the degree of association between the query vector and other nodes;
[0185] (3) Value vector: Abbreviated as q, That is, the value vector, which contains the information that needs to be weighted and aggregated according to the query vector.
[0186] Based on the above three types of vectors, the specific process of calculating the embedding of M by the first AL of SEN M Is as follows: (k,h)
[0187] (1) First, according to the original feature vector of M (k,h) SEN M Calculate the query vector (k,h) The key vector And the value vector Of M calculated through the first AL, as shown in formulas (11), (12), and (13) respectively, and calculate the query vector of O In the same way (ih,ihk) The key vector And the value vector This embodiment will not be described additionally.
[0188]
[0189] Among them, and are parameter matrices.
[0190] (2) Secondly, according to formula (14) shown below, SEN M Calculate M (k,h) M can be calculated (k,h) and the correlation coefficient between O (ih,ihk)
[0191]
[0192] Among them, is the key vector of O (ih,ihk) ; d k is the dimension of the query vector and the key vector.
[0193] Step 3.3.4: According to the obtained correlation coefficient between M (k,h) and O (ih,ihk) , calculate the attention coefficient between M (k,h) and O (ih,ihk) Furthermore, obtain the attention coefficients between M (k,h) and all neighbor nodes in its neighbor node set;
[0194] It should be noted that the number of neighbors of M (k,h) - the extended process nodes may be greater than or equal to 1. Therefore, in this embodiment, only the calculation method of the attention coefficient between M (k,h) and O (ih,ihk) is described in detail in this step. Those skilled in the art can analogously understand the calculation methods of the attention coefficients between M (k,h) and other neighbor - extended process nodes, so it will not be elaborated in this step in this embodiment.
[0195] The attention coefficient between M (k,h) and O (ih,ihk) is calculated as shown in formula (15):
[0196]
[0197] Among them, O (rh,rhk) is the extended process node after splicing O (r,h) and e ihk ; r is another workpiece index; is M(k,h) The correlation coefficient with O (rh,rhk)
[0198] Step 3.3.5: Calculate M according to the attention coefficients of M with all neighbor nodes (k,h) (k,j) through the first AL of SEN M embedding
[0199] SEN M The first AL of SEN calculates M using formulas (16) and (17) (k,h) in the embedding of SEN M ”
[0200]
[0201] where μ (ih,ihk) is the original feature vector of O (ih,ihk)
[0202] In summary, formulas (10) to (17) together constitute the network structure of SEN M obtain the embedding of M (k,h) in the first AL and use it as the input of the second AL. And so on, SEN M finally obtains the embedding of M (k,h) in the Nth AL Therefore, this embodiment will use SEN M obtained in the lth AL The general form of this process is summarized as shown in formula (18):
[0203]
[0204] Step 3.4: Construct the AGV node embedding sub-network SEN composed of N≥2 attention layers A SEN A outputs the embedding of A (u,h) at the scheduling moment t
[0205] For this step 3.4, this embodiment needs to explain that:
[0206] (1) SEN A contains N≥2 ALs. In the network structure of SEN A the first AL takes the original feature vector of A (u,h) and the original feature vector v of the extended process node O (ih,ihu) (ih,ihu) As the input, A (u,h) is embedded as the output of this AL. Since SEN A contains multiple AL layers, the output of the first AL serves as the input of the second AL, and the output of the second AL is and so on. The output of the (l - 1)-th AL serves as the input of the l-th AL, and the output of the l-th AL is the embedding of A (u,h) The output of the N-th AL is the embedding of A is A (u,h) The final embedding. This embodiment only details obtaining the embedding of A at scheduling time t in (u,h) SEN A in the first AL The specific calculation process. Those skilled in the art know that the calculation processes in other ALs of SEN A are the same as those in the first AL. Therefore, unless otherwise stated, "at scheduling time t" and "in the first AL of SEN" or "the first AL of SEN" will not be repeatedly emphasized in the following sub-steps of this embodiment A A
[0207] (2) Regarding the description of some parameters, this embodiment has also been shown in Table 3. Therefore, this embodiment will not be repeatedly described additionally in the following sub-steps
[0208] At scheduling time t, obtain the embedding of A (u,h) in SEN A in the first AL The specific calculation process includes the following steps
[0209] Step 3.4.1: Expand the process node O (i,h) by splicing the process node O (i,h) with the undirected edge e ihu to obtain a new process node, called the expanded process node O (ih,ihu) ;
[0210] Step 3.4.2: Concatenate the original feature vector of the neighbor process node O (u,h) of A (i,h) with the original feature vector of e to obtain the original feature vector μ ihu of the expanded process node O ; At the same time, O (ih,ihu) will replace O (ih,ihu) as A (ih,ihu) (i,h) as A(u,h) The new neighbor nodes in the set of neighbor nodes of and are used in the process of calculating
[0211] μ (ih,ihu) As shown in formula (19):
[0212]
[0213] Wherein, is the original feature vector of I (i,h) ; is the original feature vector of e ihu ; is the concatenation operation.
[0214] Step 3.4.3: Calculate the correlation coefficient (u,h) between A (ih,ihu) and O A calculated from the first AL in SEN
[0215] A (u,h) The correlation between and O (ih,ihu) depends on the query vector of A (u,h) and the key vector of O (ih,ihu) .
[0216] This embodiment briefly introduced the query vector, key vector, and value vector involved in SEN A in step 3.3.3, so it will not be repeated here. A (u,h) The specific process of obtaining its embedding through the first AL of SEN A is as follows:
[0217] (1) According to the original feature vector of A (u,h) SEN A The first AL calculates the query vector (u,h) A key vector and value vector in accordance with equations (20), (21), and (22), and calculates the query vector O (ih,ihu) key vector and value vector in the same way. This embodiment will not be further described.
[0218]
[0219] Wherein, and is a parameter matrix.
[0220] (2) SEN A The first AL of calculates A according to Equation (23) (u,h) and O (ih,ihu) The correlation coefficient between
[0221]
[0222] where is the key vector of O (ih,ihu) ; d k is the dimension of the query vector and the key vector.
[0223] Step 3.4.4: According to the obtained correlation coefficients between A (u,h) and all its neighbors - the extended process nodes, calculate the attention coefficients between A (u,h) and all its neighbors - the extended process nodes and then obtain the attention coefficients between A (u,h) and all the neighbor nodes in its neighbor node set;
[0224] It should be noted that the number of neighbors - the extended process nodes of A (u,h) may be greater than or equal to 1. Therefore, in this step of this embodiment, only the attention coefficients between A (u,h) and O (ih,ihk) are described in detail. Those skilled in the art can analogously deduce the calculation method of the attention coefficients between A (u,h) and other neighbors - the extended process nodes, so this embodiment will not be elaborated further.
[0225] A (u,h) and O (ih,ihk) The attention coefficient between is calculated according to formula (24):
[0226]
[0227] where O (rh,rhu) is the extended process node after splicing the process node O (r,h) with the undirected edge e ihk It is the new neighbor node of A (u,h) ; r is the index of another workpiece; is the correlation coefficient between A (u,h) and O (rh,rhu)
[0228] Step 3.4.5: According to the obtained attention coefficients between A (u,h) and all neighbor nodes, SENA The first AL of calculates A (u,h) Embedding of
[0229] SEN A The first AL of uses formulas (25) and (26) to calculate A (u,h) Embedding of ”
[0230]
[0231] Where, μ (ih,ihu) is O (ih,ihu) The original feature vector of
[0232] A (u,h) Embedding in the first AL Will be used as the input of the second AL, and so on, SEN A Finally obtains A in the Nth AL (u,h) Embedding of Therefore, this embodiment will use SEN A Obtains in the lth AL The general form of this process is summarized as shown in formula (27):
[0233]
[0234] Step 3.5: Construct the process node embedding sub-network SEN by combining N≥2 attention layers O , SEN O The output of is O (i,h) Embedding at scheduling time t
[0235] Process node embedding sub-network SEN O The network structure of is the same as that of SEN M , SEN A The network structure of is kept consistent, and this embodiment will not elaborate further in this step.
[0236] This embodiment needs to additionally state the following 3 points:
[0237] (1) The set of neighbor nodes of O (i,h) contains M (k,h) and A (u,h) , which are respectively connected to O ihk and e ihu through e (i,h) . Therefore, in the process of SEN O aggregating neighbor node information of O (i,h) , it is necessary to first transform the original feature vector of M (k,h) Concatenate with the original feature vector of e ihk and concatenate the original feature vector of A (u,h) with the original feature vector of e ihu
[0238] (2) In this embodiment, only the specific calculation process of the first AL of SEN A obtaining the embedding of O (i,h) at the scheduling moment t is described in detail. The calculation processes of the other ALs of SEN obtaining the node embeddings of O A are the same as those of the first AL. This embodiment will not repeat the description additionally in the following sub-steps. (i,h)
[0239] (3) For some other parameters, this embodiment has also been described in Table 3.
[0240] At the scheduling moment t, the specific calculation process of the first AL of SEN A obtaining the embedding of O (i,h) includes the following steps:
[0241] Step 3.5.1: Expand the parallel machine node M (k,h) and concatenate M (k,h) with the undirected edge e ihk to obtain a new parallel machine node, called the extended parallel machine node M (kh,ihk) ; Expand the AGV node A (u,h) and concatenate A (u,h) with the undirected edge e ihu to obtain a new AGV node, called the extended AGV node A (uh,ihk) ;
[0242] Step 3.5.2: Concatenate the original feature vector of M (k,h) ∈ N t (O (i,h) , M) with the original feature vector of e to obtain the original feature vector ihk of e of the extended parallel machine node M (kh,ihk) and concatenate the original feature vector of A and the original feature vector of e (u,h) ∈ N t (O (i,h) , A) with the original feature vector of e to obtain the original feature vector ihu of e of the extended AGV node A(uh,ihk) The original feature vector Meanwhile, M (kh,ihk) and A (uh,ihk) will respectively replace M (k,h) and A (u,h) as the new neighbor nodes of O (i,h) and are used in the process of calculating ;
[0243] It is calculated and obtained by formula (28):
[0244]
[0245] wherein, is the original feature vector of M (k,h) ; is the original feature vector of e ihk ; is the concatenation operation.
[0246] It is calculated and obtained by formula (29):
[0247]
[0248] wherein, is the original feature vector of A (u,h) ; is the original feature vector of e ihu ; is the concatenation operation.
[0249] Step 3.5.3: Calculate the correlation coefficient between O (i,h) and M (kh,ihk) as well as the correlation coefficient between O and A (i,h) (uh,ihu) (i,h) The correlation between O
[0250] and M (kh,ihk) and A (uh,ihk) respectively depends on the query vector of O (i,h) and the key vectors of M (kh,ihk) and A (uh,ihk) O (i,h) O The specific calculation process of the first AL of SENis as follows:
[0251] (1) First, according to formula (11), formula (12) and formula (13), use the original feature vector of O (i,h) to calculate O using the first AL of SEN O respectively(i,h) The query vector of The key vector And the value vector Are shown in formulas (30), (31), and (32) respectively, and M is calculated in the same way (hk,ihk) The query vector of The key vector The value vector And A (uh,ihk) The query vector of The key vector The value vector This embodiment will not be described in detail additionally.
[0252]
[0253] Wherein, And Are parameter matrices; Is O (i,h) The original feature vector of
[0254] (2) Secondly, calculate O respectively according to formulas (33) and (34) shown below (i,h) Respectively with M (kh,ihk) And A (uh,ihk) The correlation coefficients between them: And
[0255]
[0256] Wherein, And Are respectively the key vector of M (kh,ihk) And the key vector of A (uh,ihk) ; d k Is the dimension of the query vector and the key vector.
[0257] Step 3.5.4: Calculate the correlation coefficients between O (i,h) Respectively with O (i,h) , O (i,h-1) And O (i,h+1) Are respectively
[0258] It should be noted in this embodiment that step 3.5.3 has clearly written the calculation method of the correlation coefficients between O (i,h) And its neighbor nodes, such as formulas (33) and (34). Therefore, by analogy, those skilled in the art can understand that And Embed the sub-network SEN at the process node O The calculation method in the first AL of and is the same as that in step 3.5.3, and will not be elaborated in this embodiment.
[0259] Step 3.5.5: Calculate the attention coefficients between O (i,h) and each type of its neighbor nodes, and calculate the attention coefficients between O (i,h) and M (kh,ihk) respectively and the attention coefficients between O (uh,ihu) and A
[0260] It should be noted that in N t (O (i,h) ), the numbers of M (kh,ihk) and A (uh,ihu) may both be greater than or equal to 1. Therefore, for this property, the following operations are performed in this embodiment:
[0261] (1) Sum the correlation coefficients of all extended parallel machine nodes in N t (O (i,h) ) and the correlation coefficients of all extended AGV nodes in N t (O (i,h) ) to obtain the total correlation coefficient between O (i,h) and all extended parallel machine nodes and the total correlation coefficient between O (i,h) and all extended AGV nodes The specific calculation methods are shown in formulas (35) and (36):
[0262]
[0263] where M' k and A' h are the total numbers of machine nodes and AGV nodes in stage s h respectively.
[0264] (2) Further, use and to calculate the attention coefficients between O (i,h) and all extended parallel machine nodes in N t (O (i,h) ) and the attention coefficients between O (i,h) and all extended AGV nodes in N t (O (i,h) ) As shown in Formula (37) and Formula (38):
[0265]
[0266]
[0267] Among them, and are the correlation coefficients between O (i,h) and O (i,h) , O (i,h-1) and O (i,h+1) respectively.
[0268] (3) According to the calculation methods of and , those skilled in the art can similarly deduce the attention coefficients of O (i,h) with respect to O (i,h) , O (i,h-1) and O (i,h+1) to be and respectively, which will not be elaborated additionally in this embodiment.
[0269] Step 3.5.6: Calculate the embedding of O (i,h) passing through the first AL of SEN (i,h) based on the attention coefficients of O O with all neighbor nodes
[0270] O (i,h) 's neighbor nodes may include multiple extended parallel machine nodes. Therefore, first, this embodiment uses the existing element-wise sum method to calculate the total original feature vector of all extended parallel machine nodes as shown in Formula (39):
[0271]
[0272] Among them, N t (O (i,h) ) is the set of neighbor nodes of O (i,h) ; is the original feature vector of M (kh,ihk) .
[0273] Secondly, use to calculate the total Value vector (i,h) of all extended parallel machine nodes of O as shown in Formula (40):
[0274]
[0275] Among them, is the parameter matrix.
[0276] Similarly, those skilled in the art can deduce that (i,h) The total Value vector of all extended parallel machine nodes O (i-1,h) Value vector O (i+1)h Value vector and O (i,h) Value vector This embodiment will not be described again.
[0277] Finally, SEN O The first AL calculation O (i,h) Embed As shown in formula (41):
[0278]
[0279] In summary, formula (28) to formula (41) together constitute the process node embedding sub-network SEN O The network structure of process node O is obtained. (i,h) After SEN O The embedding calculated by the first AL And use it as the input of the second AL. And so on, SEN O At the Nth AL, O is finally obtained. (i,h) Embed Therefore, this embodiment will use SEN O Get at the lth AL The general form of this process is summarized as shown in formula (42):
[0280]
[0281] Step 3.6: Construct the full graph embedding network SENG according to formula (43) to obtain the embedding of each HFSP-FTR instance heterogeneous graph model at the scheduling time t
[0282] The SAHGNN mentioned in "Step 3" of this embodiment uses three "parallel node embedding sub-networks based on SAHGNN model decomposition" to obtain O (i,h) 、M (k,h) and A (u,h) In the embedding of "at scheduling time t" and Furthermore, the above nodes are embedded in the heterogeneous graph model of the HFSP-FTR instance The calculation method of is shown in formula (43):
[0283]
[0284] Among them, M' is the total number of parallel machines in an HFSP-FTR instance; A' is the total number of AGVs in an HFSP-FTR instance; O' is the total number of processes in an HFSP-FTR instance; S' is the total number of stages in an HFSP-FTR instance.
[0285] Step 4: According to O in each HFSP-FTR instance in the training set (i,h) Embedding at scheduling time t M (k,h) Embedding at scheduling time t and A (u,h) Embedding at scheduling time t Construct a composite scheduling action selection network (Action Select Network) ASN t , and through ASN t Calculate the action selection probability P of each composite scheduling action ; (iku,h) ;
[0286] This embodiment represents the composite scheduling action as follows:
[0287]
[0288] is specifically defined as the state State at scheduling time t t in which, O (i,h) is transported by A (u,h) to M (k,h) for processing on this complete scheduling action, which simultaneously includes two scheduling processes: parallel machine selection and AGV selection. Therefore, in order to obtain the probability of being selected in the actual scheduling process, this embodiment designs a composite scheduling action selection network ASN t to obtain at scheduling time t, the probability of being selected and executed:
[0289] First, use the action priority score calculation method shown in formula (45) to calculate the priority score of:
[0290]
[0291] Among them, MLP[] is a multi-layer perceptron; τ is a network parameter; is a concatenation operation.
[0292] Second, calculate The action selection probability in state State t is as shown in formula (46):
[0293]
[0294] where AS t is the action space at scheduling time t, which includes all composite scheduling actions at scheduling time t; at' is another composite scheduling action in AS t .
[0295] When obtaining the action selection probabilities of all composite scheduling actions in all action spaces, they will be used as the input of the improved actor-critic algorithm to obtain the optimal scheduling action selection strategy.
[0296] Step 5: Use the action selection probabilities of the composite scheduling actions corresponding to the HFSP-FTR instances in the training set obtained in Step 4 to train the Proximal Policy Optimization (PPO) algorithm, and then input the action selection probabilities of the composite scheduling actions corresponding to the HFSP-FTR instances in the test set into the trained PPO algorithm to obtain the optimal scheduling action selection strategy π t .
[0297] In this embodiment, the Proximal Policy Optimization (PPO) algorithm adopts an algorithm framework in the form of actor-critic.
[0298] In the Actor-Critic algorithm, the actor network is responsible for updating the scheduling action selection strategy, and the critic network is responsible for updating the action value function is The specific principle is as follows:
[0299] (1) The actor network selects scheduling actions based on the action selection probability;
[0300] (2) The critic network evaluates the scheduling actions selected by the actor network and gives a score;
[0301] (3) The actor network modifies the probability of subsequent selection of scheduling actions according to the score of the critic network.
[0302] After the actor network selects scheduling actions, the remaining steps are similar to the general reinforcement learning framework, which will not be elaborated in this embodiment. It should be noted that θ is the network parameter to be updated in the actor network, and ρ is the network parameter to be updated in the critic network.
[0303] To obtain a better-quality scheduling strategy, this embodiment improves the actor-critic algorithm: expands the actor network from one to at least two, and then trains the improved actor-critic algorithm using the action selection probabilities of the composite scheduling actions corresponding to the HFSP-FTR instances in the training set. To enhance the generalization of the obtained optimal scheduling action selection strategy, this embodiment expands the number of actor networks in the existing actor-critic algorithm from 1 to 20 to obtain the improved actor-critic algorithm. The improved actor-critic algorithm can randomly extract the action selection probabilities of 20 (consistent with the number of actor networks) composite scheduling actions corresponding to HFSP-FTR instances from the training set in each training step for training, and in the next training step, randomly extract the action selection probabilities of 20 HFSP-FTR instances corresponding composite scheduling actions from the training set again for training until the end of the entire training process. In this embodiment, the number of training steps is set to 800.
[0304] When the loss function Loss iteration curve of the improved actor-critic algorithm and its training iteration curve on the training set gradually tend to stable convergence, it means that the improved actor-critic algorithm has obtained the optimal scheduling action selection strategy π in the current training set. t 。
[0305] Before the training process, the SAHGNN established in step 3 of this embodiment can encode each training sample into an HFSP-FTR instance heterogeneous graph model and obtain the multi-dimensional feature vectors of each HFSP-FTR instance heterogeneous graph model. Furthermore, the multi-dimensional feature vectors of each HFSP-FTR instance heterogeneous graph model in the training set are used as the input of the composite scheduling action selection network ASN t to obtain each composite scheduling action of each HFSP-FTR instance action selection probability P (iku,h) 。In an HFSP-FTR instance heterogeneous graph model, generation is specifically divided into 2 steps: the operation node O ih is assigned to an AGV node A (u,h) , that is, an undirected edge e (i,h) appears between the operation node O (u,h) and the AGV node A ihu for connection; the operation node O (i,h) is transported by the AGV node A (u,h) to a parallel machine M (k,h) in the current stage for processing, that is, an undirected edge e (i,h) appears between the operation node O (k,h) and the parallel machine node M ihkMake a connection. Further, in this embodiment, the environmental state of the actor-critic algorithm at scheduling time t is represented as State t . Since the sequence of scheduling times is {1, 2,..., t,..., T}, the environmental state sequence for the actor-critic algorithm to obtain the optimal scheduling action selection strategy π t is {State 1 , State 2 ,..., State t ,..., State T}.
[0306] In summary, in an environmental state State t , the improved actor-critic algorithm uses the composite scheduling action selection network ASNt to select and execute the composite scheduling action with the highest action selection probability P t in the current action space AS (iku,h) , then enters the next environmental state State t+1 , until State T and terminates.
[0307] To verify the effectiveness of the obtained optimal scheduling strategy π t , this embodiment defines the representation structure of the HFSP-FTR instance as follows: J-S-A-A’-Z, where J represents the total number of workpieces to be processed; S represents the total number of stages; A represents different distribution types of parallel machines in each stage; A’ represents different distribution types of AGVs in each stage, and the number of AGVs in each stage is equal to the number of parallel machines in each stage; Z represents the index of the instance in the instance set;
[0308] For example, a HFSP-FTR instance is 15-5-A-A’-01, where:
[0309] (1) "15" means that a total of 15 workpieces need to be processed;
[0310] (2) "5" means there are 5 stages;
[0311] (3) A represents the distribution type of parallel machines in each stage, that is, the number of parallel machines in each stage is "3, 3, 1, 3, 3" respectively;
[0312] (4) A’ represents different distribution types of AGVs in each stage, and the number of AGVs in each stage is equal to the number of corresponding parallel machines in each stage, which is "3, 3, 1, 3, 3";
[0313] (5) "01" means that the number of this instance in the instance set is 01.
[0314] To verify the effectiveness of this embodiment, Figure 4 Figure 1 is a convergence curve graph of the loss function during the training process, and the training steps are 1000 steps. The abscissa in the figure is the number of training iterations, and the ordinate is the value of the loss function. It can be seen from this figure that the training network in this embodiment can converge quickly in a short time, reflecting the stability of the training algorithm; in addition, Figure 5 Figure 2 is a convergence curve graph of the return function value of the training algorithm. It can be seen that the reporting function generally tends to be stable, verifying that the scheduling strategy obtained in this embodiment is effective.
[0315] The HFSP-FTR examples selected in this embodiment are as follows:
[0316] (1) 10-5-A-A’-01; (2) 15-5-A-A’-01.
[0317] To highlight the effectiveness of the method of the present invention compared with the scheduling rule algorithms widely used in the job shop scheduling problem, four composite scheduling rules, namely SPT-LTTLJ, LPT-LTTLJ, MWKR-LTTLJ, and LWKR, are selected as comparison algorithms. The final scheduling results are shown in Table 4. It can be seen that the values of the scheduling performance indicators obtained by the method of the present invention are all smaller than those obtained by the other four comparison algorithms. Therefore, the performance of the method of the present invention is better than the above four composite scheduling rule algorithms. Among them, LTTLJ is a scheduling rule specifically used for AGVs, indicating that the AGV that has taken the least time for the most recent transportation job is preferentially selected.
[0318] Table 4 Comparison of scheduling performance between the method of the present invention and other methods
[0319]
[0320] It should be understood that those skilled in the art can make various improvements or transformations based on the above content under the inspiration of the technical concept of the present invention without departing from the content of the present invention, and this still falls within the protection scope of the present invention.
Claims
1. A collaborative optimization method for hybrid flow shop scheduling problem with limited transportation resources, characterized in that: The method comprises the following steps: Step 1: Obtain HFSP-FTR instance and its production and processing information; Step 2: Based on the production and processing information of the HFSP-FTR instance, the HFSP-FTR instance is modeled as a heterogeneous graph model that includes process nodes, parallel machine nodes, and AGV nodes, which is called the HFSP-FTR instance heterogeneous graph model; Step 3: Based on the existing attention mechanism theory and graph neural network theory, a structure-aware heterogeneous graph neural network is constructed, and the structure-aware heterogeneous graph neural network is used to obtain the multi-dimensional feature vector of the HFSP-FTR instance heterogeneous graph model; The structure-aware heterogeneous graph neural network includes a full-graph embedding network and three parallel embedding sub-networks; the three parallel embedding sub-networks are process node embedding sub-networks SEN O , parallel machine node embedding sub-network SEN M And AGV node embedded sub-network SEN A ; The full graph embedding network is used to calculate the graph embedding of the entire HFSP-FTR instance heterogeneous graph model; the process node embedding subnetwork SEN O Used to calculate process node O (i,h) The parallel machine nodes are embedded in the sub-network SEN M For computing parallel machine nodes M (k,h) The AGV node is embedded in the sub-network SEN A Used to calculate AGV node A (u,h) Node embedding of The SEN O , SEN M and SEN A They are constructed by combining N ≥ 2 attention layers, where SEN O The output is O (i,h) Embedding at scheduling time t SEN M The output is M (k,h) Embedding at scheduling time t SEN A The output is A (u,h) Embedding at scheduling time t According to formula (43), the full-graph embedding network SEN is constructed G , obtain the embedding of each HFSP-FTR instance heterogeneous graph model at scheduling time t Wherein, M' represents the total number of parallel machines in the HFSP-FTR instance; A' represents the total number of AGVs in the HFSP-FTR instance; O' represents the total number of processes in the HFSP-FTR instance; S' represents the total number of stages in the HFSP-FTR instance; Step 4: Build a composite scheduling action selection network ASN t , taking the multi-dimensional feature vector of the instance heterogeneous graph model as ASN t Input, through ASN t Calculate each compound scheduling action The action selection probability P (iku,h) ; Indicates that at the scheduling time t, O (i,h) By A (u,h) Delivery to M (j,h) The complete scheduling action of processing includes both parallel machine selection and AGV selection, so it is called a composite scheduling action. The composite scheduling action is expressed as follows: The composite scheduling action selects the network ASN t The construction method is as follows: First, the composite scheduling action is calculated according to formula (45): Action priority scores: Where MLP[] is a multi-layer perceptron; τ is a network parameter; For splicing operation; Secondly, according to formula (46), The probability of action selection at scheduling time t: Among them, AS t is the action space at scheduling time t, which includes all composite scheduling actions at scheduling time t; a t' AS t One other composite dispatch action in ; Step 5: Input the action selection probability obtained in step 4 into the proximal policy optimization algorithm to obtain the optimal scheduling action selection strategy π t .
2. The collaborative optimization method for hybrid flow shop scheduling problem with limited transportation resources according to claim 1 is characterized in that: The heterogeneous graph model of the HFSP-FTR instance described in step 2 is represented as in: (1) is a set of process nodes, including all process nodes in the HFSP-FTR instance heterogeneous graph model; (2) is a set of parallel machine nodes, including all parallel machine nodes in the HFSP-FTR instance heterogeneous graph model; (3) is the AGV node set, which includes all AGV nodes in the HFSP-FTR instance heterogeneous graph model; (4) is the connection process node O in the HFSP-FTR instance heterogeneous graph model (i,h) and parallel machine node M (k,h) The set of all undirected edges between ; (5) is the connection process node O in the HFSP-FTR instance heterogeneous graph model (i,h) and AGV node A (u,h) The set of all undirected edges between ; (6) is a set of directed edges, including the directed edges used to represent the workpiece processing sequence in the HFSP-FTR instance heterogeneous graph model; Among them, (i,h) Indicates workpiece J i The hth process node in M (k,h) Indicates stage h The kth parallel machine node in A (u,h) Indicates stage h The u-th AGV in .
3. The collaborative optimization method for hybrid flow shop scheduling problem with limited transportation resources according to claim 2 is characterized in that: The process of establishing the HFSP-FTR instance heterogeneous graph model should meet the following process constraints: (1) Each workpiece can only be processed by one parallel machine in each stage; (2) Each workpiece must be processed in the previous stage before the current stage can be started; (3) For two workpieces processed on the same parallel machine, processing of the other workpiece can only start after processing of one workpiece is completed; (4) All AGVs are identical and travel at a constant speed; (5) Once the AGV starts working, it is not allowed to be interrupted; (6) An AGV can only transport one workpiece at a time.
4. The collaborative optimization method for hybrid flow shop scheduling problem with limited transportation resources according to claim 1, characterized in that: The SEN in step 3 O , SEN M and SEN A The construction process includes the following steps: Step 3.1: Define process nodes O separately (i,h) , parallel machine node M (k,h) And AGV node A (u,h) The neighbor node set of: Process Node O (i,h) The set of neighbor nodes N at scheduling time t t (O (i,h) )={O (i,h-1) ,O (i,h+1) ,O (i,h) ,M (k,h) ,A (u,h) }, where O (i,h-1) Indicates O (i,h) The previous process node, O (i,h+1) Indicates O (i,h) The next process node, M (k,h) Indicates that (i,h) The connected parallel machine nodes, A (u,h) Indicates that (i,h) Connected AGV nodes; Parallel machine node M (k,h) The set of neighbor nodes N at scheduling time t t (M (k,h) )={M (k,h) ,N t (M (k,h) ,O)}, where N t (M (k,h) ,O) indicates M (k,h) All process nodes that can be processed in the current stage; AGV Node A (u,h) The set of neighbor nodes N at scheduling time t t (A (u,h) )={A (u,h) ,N t (A (u,h) ,O)}, where N t (A (u,h) ,O) indicates A (u,h) All process nodes that can be transported in the current stage; Step 3.2: and Assign a multidimensional raw feature vector: O (i,h) Assigning the original feature vector Where |N t (O (i,h) ,M)| represents O (i,h) The number of neighbor parallel machine nodes; |N t (O (i,h) ,A)| represents O (i,h) The number of neighboring AGV nodes; NUSO ih Indicates O (i,h) The number of unscheduled processes; C (ihk) Indicates O (i,h) Completion time; S (ihk) Indicates O (i,h) The start processing time; P (ihk) +TT i(h+1)ku Indicates O (i,h) In M (k,h) Processing time P (ihk) With A (u,h) The sum of the transportation time; Status represents the current status. At the scheduling time t, if O (i,h) If it has been scheduled, the Status is 1, otherwise the Status is 0; M (k,h) Assigning the original feature vector Among them I t (M (k,h) ) is a two-dimensional variable. At the scheduling time t, if M (k,h) If a process node is assigned, t (M (k,h) )=1, otherwise I t (M (k,h) )=0; Indicates M (k,h) Utilization rate; |N t (M (k,h) ,O)| indicates M (k,h) The number of neighbor process nodes; Indicates M (k,h) Able to start processing O (i,h) time; A (u,h) Assigning the original feature vector Among them I t (A (u,h) ) is a two-dimensional variable. At the scheduling time t, if AGV node A (u,h) If a process node is assigned, t (A (u,h) )=1, otherwise I t (A (u,h) )=0; Indicates A (u,h) Utilization rate; |N t (A (u,h) ,O)| represents A (u,h) The number of neighbor process nodes; Indicates A (u,h) Able to start transporting O (i,h) time; O (i,h) With M (k,h) The undirected edge e between ihk Assigning the original feature vector Where P (ihk) Indicates O (i,h) In M (k,h) Processing time on O (i,h) With A (u,h) The undirected edge e between ihu Assigning the original feature vector TT i(h+1)ku Indicates O (i,h) A (u,h) The time required for transportation; Step 3.3: Construct a parallel machine node embedding subnetwork SEN by combining N ≥ 2 attention layers M , SEN M The output is M (k,h) Embedding at scheduling time t SEN M The input includes M (k,h) The original feature vector of and extended process node O (ih,ihk) The original eigenvector μ (ih,ihk) ; Among them, the extended process node O (ih,ihk) is to pass the process node O (i,h) With undirected edge e ihk The new process node obtained by splicing; (k,h) Neighboring process node O (i,h) The original feature vector of With e ihk The original feature vector of Splice to get the extended process node O (ih,ihk) The original eigenvector μ (ih,ihk) ; Step 3.4: Construct the AGV node embedding subnetwork SEN by combining N ≥ 2 attention layers A , SEN A The output is A (u,h) Embedding at scheduling time t SEN A The input includes A (u,h) The original feature vector of and extended process node O (ih,ihu) The original eigenvector μ (ih,ihu) ; Among them, the extended process node O (ih,ihu) is to pass the process node O (i,h) With undirected edge e ihu The new process node obtained by splicing; (u,h) Neighboring process node O (i,h) The original feature vector of With e ihu The original feature vector of Splice to get the extended process node O (ih,ihu) The original eigenvector μ (ih,ihu) ; Step 3.5: Construct the process node embedding sub-network SEN by combining N ≥ 2 attention layers O , SEN O The output is O (i,h) Embedding at scheduling time t SEN O The input includes O (i,h) The original feature vector of Expand parallel machine node M (kh,ihk) The original feature vector of and extended AGV node A (uh,ihk) The original feature vector of in, Expand parallel machine node M (kh,ihk) By using M (k,h) With undirected edge e ihk Connect the new parallel machine nodes obtained; (k,h) The original feature vector of With e ihk The original feature vector of Connect and get the extended parallel machine node M (kh,ihk) The original feature vector of Expand AGV Node A (uh,ihk) is by (u,h) With undirected edge e ihu New AGV node obtained by splicing; A (u,h) The original feature vector of With e ihu The original feature vector of Splice to obtain extended AGV node A (uh,ihk) The original feature vector of 5. The collaborative optimization method for hybrid flow shop scheduling problem with limited transportation resources according to claim 1, characterized in that: The HFSP-FTR instance is represented as: JSA-A'-Z, where J represents the total number of workpieces that need to be processed; S represents the total number of stages; A represents the different distribution types of parallel machines in each stage; A' represents the different distribution types of AGVs in each stage, and the number of AGVs in each stage is equal to the number of parallel machines corresponding to A in each stage; Z represents the index of the instance in the instance set.
6. The collaborative optimization method for hybrid flow shop scheduling problem with limited transportation resources according to claim 1, characterized in that: The proximal strategy optimization algorithm adopts the actor-critic algorithm.
7. The collaborative optimization method for hybrid flow shop scheduling problem with limited transportation resources according to claim 6 is characterized in that: Improve the actor-critic algorithm: expand the actor network from one to at least two.
Citation Information
Patent Citations
Flexible job shop scheduling method based on graph reinforcement learning under industrial demand response
CN118095753A
HFSP scheduling optimization method based on heterogeneous graph neural network and deep reinforcement learning
CN118295352A