Industrial manufacturing workflow scheduling method based on graph attention network and multi-agent reinforcement learning
Through the combination of graph attention network and multi-agent reinforcement learning, the problem of complex workflow scheduling in industrial manufacturing is solved, efficient and flexible resource allocation and scheduling is achieved, and production efficiency and system stability are improved.
Patent Information
- Application Number
- CN202510328129.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art is difficult to effectively deal with complex workflow scheduling problems in industrial manufacturing, especially under the influence of factors such as task arrival uncertainty and machine failure, resulting in task execution delays and resource waste.
Using a method based on graph attention network and multi-agent reinforcement learning, a workflow scheduling model is built, and the allocation strategy is trained through the Q-learning algorithm, and the task progress and machine load changes are responded in real time, and the container execution order is dynamically updated.
Improve the flexibility and adaptability of workflow scheduling, optimize resource allocation, reduce machine idle time, improve scheduling success rate and reliability, reduce operational costs, and achieve green production.
Smart Images

Figure CN120256052A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent manufacturing, and particularly to an industrial manufacturing workflow scheduling method based on graph attention network and multi-agent reinforcement learning. Background Art
[0002] Internetware can effectively combine software entities with autonomous characteristics on various nodes in the Internet environment according to predetermined function and performance index requirements. Based on interconnection, intercommunication, and cooperation, these entities execute component services in the form of a workflow. In the processing of manufacturing tasks, this method usually transforms tasks and their dependencies into directed acyclic graphs (DAGs), and determines the optimal execution order of tasks through graph simplification operations and the like to achieve the best production efficiency and maximize resource utilization.
[0003] Although there are also some existing technologies providing workflow scheduling methods, for example, Chinese Patent CN118409848A discloses a cloud computing workflow scheduling method, system, processing device, and medium. The method includes: constructing a workflow model of a to-be-tested application based on an application request; solving a pre-constructed workflow scheduling model based on the workflow model of the to-be-tested application to obtain an optimal workflow scheduling scheme for the to-be-tested application, where the optimal workflow scheduling scheme includes which virtual machine each task is assigned to execute, as well as the total completion time and total cost of the entire workflow, and the workflow scheduling model is trained using a deep reinforcement learning algorithm.
[0004] However, since tasks originate from different devices, and each task is accompanied by complex dependencies, this constitutes an NP-hard problem involving multi-objective optimization, including key factors such as machine rental cost, success rate, and deadline violation rate. Compared with traditional workflow scheduling problems, computational tasks in industrial clouds are more complex and changeable, covering various types of computational tasks such as equipment predictive maintenance, equipment control, and product detection. These tasks not only have complex structures but also involve a large amount of data processing and complex machine learning models, further increasing the difficulty of the workflow scheduling problem. In addition, the arrival of tasks in industrial manufacturing has a high degree of uncertainty. Tasks may appear alternately frequently, and factors such as machine failures or equipment maintenance often cause delays in task execution, and these factors may all affect the final scheduling result. Therefore, how to effectively perform workflow scheduling to ensure the efficient execution of tasks and the stable operation of the system. Summary of the Invention
[0005] The purpose of the present invention is to provide an industrial manufacturing workflow scheduling method based on graph attention network and multi-agent reinforcement learning.
[0006] The purpose of the present invention can be achieved through the following technical solutions:
[0007] An industrial manufacturing workflow scheduling method based on graph attention network and multi-agent reinforcement learning, comprising:
[0008] Step S1: Obtain a workflow and construct a workflow scheduling model based on the workload;
[0009] Step S2: Construct a reinforcement learning model according to the workflow scheduling model and the machine cluster resource model, and use the Q-learning algorithm based on deep learning for training to output the final allocation strategy;
[0010] Step S3: During the execution of the workflow tasks, according to the execution progress of the tasks and the load conditions of the machines, respond to the changes in task progress and machine load in real time, and dynamically update the execution order of the containers.
[0011] The said Step S1 includes:
[0012] Step S1-1: Model the workflow w i as a directed acyclic graph:
[0013] G=(T,E)
[0014] Where: G is the directed acyclic graph, T is the set of nodes of the directed acyclic graph, each node includes a task, E is the set of edges of the directed acyclic graph, and each edge represents the control dependency between two tasks;
[0015] Step S1-2: Configure the workflow in the form of a configuration file, and the configuration content includes: the start time of the workflow, the included tasks, the fixed attributes of the tasks, and the dependency relationships between the tasks. Among them, the fixed attributes of the tasks include: the number of CPU cores required to run the task and the execution time;
[0016] Step S1-3: Match the tasks in the workflow with the corresponding containers. Among them, for any task t i , if the set of containers that can be allocated when it runs is C={c1,c2,…,c m}, where c m represents each independent container, and each task is calculated by an independent container. A workflow w i requires multiple containers to work together to complete;
[0017] Step S1-4: Construct a container execution queue through sub-deadlines to complete the construction of the workflow scheduling model.
[0018] In the said Step S1-1, any task in the set of nodes of the directed acyclic graph is an indivisible task;
[0019] If any task ti There is an edge e with a data transfer volume of data ij , representing the execution order constraint between tasks t ij and t i , that is, the priority of subtask t j is lower than that of the parent task t j . i .
[0020] In the steps S1-4, the calculation method for the container sub-deadline is as follows
[0021]
[0022]
[0023] Where: R m is the upward rank obtained by container c m according to the topological structure, c n is the sub-container of container c m , R n is the upward rank obtained by container c n according to the topological structure, T mn is the transmission time between containers, t m * is the shortest execution time of the container, sd m is the sub-deadline of the container, d i is the deadline of the workflow, R entry is the upward rank of the starting task.
[0024] The state space of the reinforcement learning model includes the execution time of the container, the amount of data required to be transmitted by the container, the estimated start time of the container on the server, the weight value of the server, and the rental cost of the server;
[0025] The action space A of the reinforcement learning model t is:
[0026] A t ={a t |a t =<c m ,m k >},c m ∈C,m k ∈M
[0027] Where: a t is the action, indicating that container c k is run by machine m m , m k is the kth machine;
[0028] The reward function r of the reinforcement learning model is:
[0029] r = -ωcost(m k ) - (1 - ω)ECT(c m , m k )
[0030] Where: ω is the weight value in the overall optimization objective, cost(m k ) is the cost of the machine, ECT(c m , m k ) is the estimated completion time.
[0031] Since the container features and machine features are a low-dimensional input, the Q-network of the reinforcement learning model includes:
[0032] Perform an embedding representation on it through the structure of a multi-layer perceptron MLP, mapping this inseparable data into high-dimensional features:
[0033] h c = MLP(d c ) = σ(W c d c + b c )
[0034] h m = MLP(d m ) = σ(W m d m + b m )
[0035] Where: h c is the output feature vector of the container, h m is the output feature vector of the machine, MLP(d c ) is the multi-layer perceptron (Multilayer Perceptron) network of the container features, with the input being the feature vector d c of the container, MLP(d m ) is the multi-layer perceptron network representing the machine features, with the input being the feature vector d m of the machine, W c is the weight matrix in the multi-layer perceptron network of the container features, used for linear transformation of the input container feature vector d c , W m is the weight matrix in the multi-layer perceptron network of the machine features, used for linear transformation of the input machine feature vector d m , b c is the bias vector in the multi-layer perceptron network of the container features,, b m is the bias vector in the multi-layer perceptron network of the machine features, σ(·) is the activation function;
[0036] Machine node feature extraction is achieved using a graph attention network. Assume there are N machine nodes in the input graph, and the input feature of each machine node i is where D is the feature dimension of the node. For each node i, its attention weight α ij represents the relationship strength between machine node i and its neighbor node j, and this weight is calculated by the following formula:
[0037] e ij = LeakReLU(a T [Wh i ||Wh j )
[0038] where: W is the learned weight matrix for feature transformation of the input features, || represents the feature concatenation operation, and a is the shared attention parameter for calculating the relationship between node pairs
[0039] The coefficients of the entire neighborhood are normalized using the softmax function, and the specific processing method is as follows:
[0040]
[0041] where: α ij is the attention coefficient of node i to its neighbor node j, obtained through normalization by the softmax function, and is used for subsequent weighted summation of the features of neighbor nodes. N(i) represents the set of neighbor nodes of node i;
[0042] Update the feature of node i based on all neighbor nodes:
[0043]
[0044] where: represents the feature vector of the i-th container at the H-th layer, and h′ m is the output feature vector after passing through multiple attention layers.
[0045] Merge the container feature and the machine feature to obtain h1, and output the final feature vector to the Q network through a two-layer perceptron:
[0046] h1 = concat(h t , h′ m )
[0047] h2 = MLP(h1) = σ(W1h1 + b1)
[0048] h3 = MLP(h2) = σ(W2h2 + b2)
[0049] where: h1 represents the combination of the container feature h t and the machine feature h′m Combined, h2 represents the feature vector after being processed by the first-layer multi-layer perceptron (MLP), which is the non-linear transformation result of h1. h3 represents the feature vector after being processed by the second-layer multi-layer perceptron (MLP), which is the further non-linear transformation result of h2, and is finally output to the Q network, h t is the feature vector of the container, W1 is the weight of the first-layer MLP, W2 is the weight of the second-layer MLP, b1 is the bias vector of the first-layer MLP, and b2 is the bias vector of the second-layer MLP.
[0050] In the step S2, during the training process of the Q-learning algorithm based on deep learning, Q values are used to evaluate the quality of each state-action pair. The ultimate goal is to maximize the cumulative reward. The Q network generates a richer state representation through the graph attention layer during this process. During its training and update process, the input features include the machine resource information of the graph structure. The update formula of the network is as follows:
[0051] Q(s t ,a t )←Q(s t ,a t )+α*[r t +γ*max a′ Q(s t+1 ,a′)-Q(s t ,a t )]
[0052] Where: Q(s t ,a t ) represents the action value function of taking action a t in state s t , r t is the reward at the current time t, γ is the discount factor, α is the learning rate, a′ represents the set of possible actions in the next state s t+1 , and the update objective is to minimize the difference between the predicted value of the evaluated Q network and the target Q value through backpropagation.
[0053] The loss function L(θ) of the training process of the deep learning model is:
[0054] L(θ)=E[(r t +γ*max a′ Q(s t+1 ,a′|θ′)-Q(s t ,a t |θ)) 2
[0055] Where: r t is the reward for executing action a t obtained later, E[·] is the expected value of the loss function, and θ is the weight parameter for evaluating the Q network.
[0056] In step S3, when an execution delay occurs in a certain container, the expected start time and expected completion time of related subsequent containers are updated in real time, and the update formula is as follows:
[0057]
[0058] where: Δt m is the delay time of the container, EST(c n ,m k ) is the expected start time, ECT(c n ,m k ) is the expected completion time, and succ(c m ) represents the set of child nodes of the container.
[0059] An industrial manufacturing workflow scheduling device based on graph attention network and multi-agent reinforcement learning includes a memory, a processor, and a program stored in the memory. When the processor executes the program, the above method is implemented.
[0060] Compared with the prior art, the present invention has the following beneficial effects:
[0061] 1. Improve the flexibility and adaptability of workflow scheduling: By integrating graph attention network and multi-agent reinforcement learning, the workflow scheduling system can monitor the changes in the production environment in real time, such as task arrival, machine failure, or resource availability changes, and quickly adjust the scheduling strategy to adapt to these changes.
[0062] 2. Optimize resource allocation and reduce machine idle time: Adopt an intelligent scheduling algorithm to accurately allocate tasks to the most suitable machines, avoid resource waste and machine overload, thereby improving the utilization efficiency of machines. By dynamically adjusting the execution order of tasks, the system can ensure that tasks are efficiently completed within the specified time, reducing machine idle time. This not only improves production efficiency but also helps to reduce energy consumption and operating costs, achieving green production and sustainable development.
[0063] 3. Improve the success rate of scheduling: By combining the reinforcement learning framework and graph attention network, the information aggregation and collaboration between machine nodes are optimized, minimizing the machine rental cost while meeting the scheduling deadline requirements from containers to machines. This optimization strategy significantly improves the success rate of scheduling, ensuring that tasks can be completed within the specified time.
[0064] 4. Improve the reliability of scheduling: Introduce a dynamic adjustment strategy to enhance the fault tolerance and robustness of the system. When the execution of a container is delayed, the system can update the expected start time and completion time of relevant subsequent containers in real time, avoid task execution conflicts, and ensure the maximization of resource utilization efficiency. This dynamic adjustment mechanism enables the system to maintain the reliability of scheduling and reduce the risk of production interruption in the face of sudden situations such as machine failures and task delays.
[0065] 5. Combine the Graph Attention Network (GAT) with a reinforcement learning framework to optimize the information aggregation and collaboration between machine nodes. This fusion method can more effectively handle the dependencies between tasks and make optimal scheduling decisions in real time in a dynamic environment. By enhancing the GAT's ability to model task dependencies and machine states, it shows higher efficiency and adaptability in complex workflow scheduling. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 It is a framework diagram of an industrial manufacturing workflow scheduling based on the Graph Attention Network and multi-agent reinforcement learning;
[0067] Figure 2 It is an experimental result diagram of the industrial manufacturing workflow scheduling method based on the Graph Attention Network and multi-agent reinforcement learning using four scientific workflows under different workflow arrival factors;
[0068] Figure 3 It is an experimental result diagram of the industrial manufacturing workflow scheduling method based on the Graph Attention Network and multi-agent reinforcement learning using the Alibaba cluster workflow under different workflow arrival factors;
[0069] Figure 4 It is an experimental result diagram of the industrial manufacturing workflow scheduling method based on the Graph Attention Network and multi-agent reinforcement learning using four scientific workflows under different machine failure probabilities;
[0070] Figure 5 It is an experimental result diagram of the industrial manufacturing workflow scheduling method based on the Graph Attention Network and multi-agent reinforcement learning using the Alibaba cluster workflow under different machine failure probabilities;
[0071] Figure 6 It is a schematic diagram of the main step process of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0072] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and gives detailed implementation manners and specific operation processes, but the protection scope of the present invention is not limited to the following embodiments.
[0073] Some formal definitions used are as follows:
[0074] (1) Execution time: When the task is assigned to container c m , the actual processing time is related to the computing performance of the machine to which the container is assigned. The execution time t(c m ,m k ) is calculated as follows:
[0075] t(c m ,m k ) = T(c m ) * F k
[0076] In the formula, T(c m ) represents the expected execution time of container c m , and F k represents the computing performance of machines at different levels. For a machine of type k, its weight F k is defined as the weight value of the execution time of this type of machine for this computing task.
[0077] (2) Transmission time: If the execution sub-container c m and its parent container c n are assigned to different machines, the data transmission time between them depends on the size of the data to be transmitted and the bandwidth between the machines. This means that if c m and c n are assigned to the same machine, the corresponding data transmission time is zero.
[0078]
[0079] In the formula, bw is the bandwidth size of the machine, and data is the size of the data to be transmitted.
[0080] (3) Machine failure probability: Machine failure is a common dynamic factor in industrial manufacturing. Prolonged use will cause machine equipment. When a machine failure occurs, the probability ρ k is calculated as follows:
[0081]
[0082] (4) Expected start time and expected completion time: The expected start time of container c n depends on the actual completion time of all its parent containers, the expected data transmission time T mn between it and all its parent containers, and the available time of the machine m k to be assigned. The corresponding recursive relationship is:
[0083]
[0084] At the same time, define container cn The estimated completion time on machine m k is:
[0085] ECT(c n ,m k ) = EST(c n ,m k ) + t(c m ,m k )
[0086] (5) Machine rental cost: The rental cost of machine m k is related to the rental time and rental price of the machine, and can be calculated as:
[0087] cost(m k ) = (LCT(m k ) - LST(m k )) · price(m k )
[0088] The formula for calculating the total rental cost is as follows:
[0089]
[0090] Therefore, according to the above definitions, the optimization objective of this scheduling problem is defined as minimizing the total execution cost TEC(W) under the constraint of meeting the deadline d i of each workflow w i , which can be expressed as:
[0091] min TLC(W)
[0092] S.T. TET(w i ) ≤ d i
[0093]
[0094] where is used to determine whether the container c m is allocated on the machine m k :
[0095]
[0096] An industrial manufacturing workflow scheduling method based on graph attention network and multi-agent reinforcement learning, as shown in Figure 1 and Figure 6 , includes:
[0097] Step S1: Obtain the workflow and construct a workflow scheduling model based on the workload. Step S1 includes:
[0098] Step S1-1: Model the workflow w i as a directed acyclic graph:
[0099] G = (T, E)
[0100] where: G is the directed acyclic graph, T is the set of nodes of the directed acyclic graph, each node includes a task, E is the set of edges of the directed acyclic graph, and each edge represents the control dependency relationship between two tasks;
[0101] Step S1-2: Configure the workflow in the form of a configuration file. The configuration content includes: the start time of the workflow, the tasks included, the fixed attributes of the tasks, and the dependency relationships between the tasks. Among them, the fixed attributes of the tasks include: the number of CPU cores required to run the task and the execution time;
[0102] Step S1-3: Match the tasks in the workflow with the corresponding containers. Among them, for any task t i , if the set of containers that can be allocated when it runs is C = {c1, c2,..., c m}, where c m represents each independent container, and each task is calculated by an independent container. Therefore, a workflow w i requires multiple containers to work together to complete;
[0103] Step S1-4: Build a container execution queue through sub-deadlines to complete the construction of the workflow scheduling model.
[0104] In Step S1-1, any task in the set of nodes of the directed acyclic graph is an indivisible task;
[0105] If any task t i has an edge e ij with a data transfer volume of data ij , it represents the execution order constraint between task t i and t j , that is, the priority of subtask t j is lower than that of the parent task r i .
[0106] In Step S1-4, the calculation method for the sub-deadline of the container is as follows
[0107]
[0108] where: R m is the upward rank obtained by container c m according to the topological structure, c n is the sub-container of container c m , and R n is container cn The upward rank obtained according to the topological structure, T mn is the transmission time between containers, t m * is the shortest execution time of the container, sd m is the sub-deadline of the container, d i is the deadline of the workflow, R entry is the upward rank of the starting task.
[0109] Step S2: Construct a reinforcement learning model according to the workflow scheduling model and the machine cluster resource model, and use the Q-learning algorithm based on deep learning for training to output the final allocation policy;
[0110] The state space of the reinforcement learning model includes the execution time of the container, the amount of data to be transmitted by the container, the expected start time of the container on the server, the weight value of the server, and the rental cost of the server;
[0111] The action space A of the reinforcement learning model t is:
[0112] A t ={a t |a t =<c m ,m k >}, c m ∈C, m k ∈M
[0113] where: a t is the action, indicating that the container c k is run by the machine m m , m k is the kth machine;
[0114] The reward function r of the reinforcement learning model is:
[0115] r = -ωcost(m k ) - (1 - ω)ECT(c m , m k )
[0116] where: ω is the weight value in the overall optimization objective, cost(m k ) is the cost of the machine, and ECT(c m , m k ) is the expected completion time.
[0117] Since the container features and machine features are a low-dimensional input, the Q-network of the reinforcement learning model includes:
[0118] It is embedded and represented through the structure of a multi-layer perceptron (MLP), mapping this inseparable data into high-dimensional features:
[0119] h c = MLP(d c ) = σ(W c d c + b c )
[0120] h m = MLP(d m ) = σ(W m d m + b m )
[0121] Where: h c is the output feature vector of the container, h m is the output feature vector of the machine, MLP(d c ) is the multi-layer perceptron (MLP) network of the container features, with the input being the feature vector d c of the container, MLP(d m ) is the multi-layer perceptron network representing the machine features, with the input being the feature vector d m of the machine, W c is the weight matrix in the multi-layer perceptron network of the container features, used for linearly transforming the input container feature vector d c , W m is the weight matrix in the multi-layer perceptron network of the machine features, used for linearly transforming the input machine feature vector d m , b c is the bias vector in the multi-layer perceptron network of the container features,, b m is the bias vector in the multi-layer perceptron network of the machine features, σ(·) represents the activation function;
[0122] The machine node feature extraction is achieved using the graph attention network. Assume there are Ni machine nodes in the input graph, and the input feature of each machine node i is where D is the feature dimension of the node. For each node i, its attention weight α ij represents the relationship strength between machine node i and its neighbor node j, and this weight is calculated by the following formula:
[0123] e ij = LeakyReLU(a T [Wh i || Wh j )
[0124] Where: W is the learned weight matrix for feature transformation of the input features, || represents the feature concatenation operation, and a is the shared attention parameter for calculating the relationship between node pairs.
[0125] The coefficients of the entire neighborhood are normalized using the softmax function, and the specific processing method is as follows:
[0126]
[0127] Where: α ij is the attention coefficient of node i to its neighboring node j, obtained through normalization by the softmax function, and is used for subsequent weighted summation of the features of neighboring nodes. N(i) represents the set of neighboring nodes of node i;
[0128] Update the feature of node i based on all neighboring nodes:
[0129]
[0130] Where: represents the feature vector of the i-th container at the H-th layer, and h′ m is the output feature vector after passing through multiple attention layers.
[0131] Merge the container feature and the machine feature to obtain h1, and output the final feature vector to the Q network through two-layer perceptrons:
[0132] h1 = concat(h t , h′ m )
[0133] h2 = MLP(h1) = σ(W1h1 + b1)
[0134] h3 = MLP(h2) = σ(W2h2 + b2)
[0135] Where: h1 represents the combination of the container feature h t and the machine feature h′ m h2 represents the feature vector after being processed by the first-layer multi-layer perceptron (MLP), which is a non-linear transformation result of h1. h3 represents the feature vector after being processed by the second-layer multi-layer perceptron (MLP), which is a further non-linear transformation result of h2, and is finally output to the Q network. h t is the feature vector of the container, W1 is the weight of the first-layer MLP, W2 is the weight of the second-layer MLP, b1 is the bias vector of the first-layer MLP, and b2 is the bias vector of the second-layer MLP.
[0136] In step S2, the deep learning-based Q-learning algorithm uses Q-values to evaluate the quality of each state-action pair, and the ultimate goal is to maximize the cumulative reward. In this process, the Q-network in this paper generates a richer state representation through the graph attention layer. Its training and update process is similar to the traditional deep learning-based Q-learning algorithm. However, due to the introduction of graph attention, the input features are more complex and include the machine resource information of the graph structure. The following is the update formula of the network:
[0137] Q(s t ,a t )←Q(s t ,a t )+α*[r t +γ*max a′ Q(s t+1 ,a′)-Q(s t ,a t )]
[0138] Where: Q(s t ,a t ) represents the action value function of taking action a t under state s t , r t is the reward at the current time t, γ is the discount factor, α is the learning rate, a′ represents the set of possible actions in the next state s t+1 . The goal of the update is to minimize the difference between the predicted value of the evaluated Q-network and the target Q-value through backpropagation.
[0139] The loss function L(θ) in the training process of the deep learning model is:
[0140] L(θ)=E[(r t +γ*max a′ Q(s t+1 ,a′|θ′)-Q(s t ,a t |θ)) 2
[0141] Where: r t is obtained after executing action a t at the current time t, E[·] is the expected value of the loss function, and θ is the weight parameter of the evaluated Q-network.
[0142] Step S3: During the execution of the workflow task, according to the execution progress of the task and the load situation of the machine, respond to the changes in the task progress and machine load in real time, and dynamically update the execution order of the containers.
[0143] In step S3, during the execution of the workflow task, the system dynamically adjusts the scheduling policy according to the execution progress of the task and the load condition of the machine, ensuring that the task can be completed on time and optimizing the resource utilization efficiency.
[0144] When the execution of a certain container is delayed, the expected start time and expected completion time of the relevant subsequent containers need to be updated in real time. Such update not only needs to consider the execution delay of the container itself, but also the dependencies between containers. The update formula is as follows:
[0145]
[0146] where: Δt m is the delay time of the container, EST(c n ,m k ) is the expected start time, ECT(c n ,m k ) is the expected completion time, succ(c m ) represents the set of child nodes of the container. By this method, the system can dynamically adjust the priorities of all relevant containers according to the actual completion time of the delayed container, thus avoiding container execution conflicts and maximizing the resource utilization efficiency.
[0147] As Figure 2 - As Figure 5 shown, the experimental results of using four scientific workflows and the Alibaba cluster workflow under different conditions for the method of this embodiment are presented. In the experimental setup, the Q network consists of one GAT layer, three fully connected layers, and one dense linear output layer. The learning rate α is set to 0.009, the discount factor γ is 0.9, and the exploration rate range ε is between 0.99 and 0.1. The evaluation metrics used in the experiment are as follows:
[0148] (1) Average rental cost: To examine the rental cost under each strategy, the average rental cost after k experiments is defined as:
[0149]
[0150] (2) Success rate: The success rate can be used to represent the final completion situation of the workflow scheduling. The specific calculation method is as follows:
[0151]
[0152] In the formula, represents the number of workflows that meet the deadline constraint, and I is the total number of workflows.
[0153] (3) Deadline deviation value: To evaluate the completion situation of the workflow under the deadline constraint, the average deadline deviation value of all workflows is defined as:
[0154]
[0155] where makespan i is the running duration of workflow W i and a i represents the arrival time of the workflow.
[0156] According to the experiments, it can be found that this method performs best in terms of average rental cost and success rate, and its deadline violation rate is also relatively low. It can be concluded that it can perform resource scheduling more stably and reliably in the face of dynamic and uncertain environments.
[0157] The above-mentioned datasets used are described as follows: four scientific workflows, namely LIGO, Cybershake, SIPHT, and MONTAGE. These workflows have different structures and properties and cover diverse scheduling scenarios. The LIGO workflow mainly focuses on the detection and analysis of gravitational waves, Cybershake is used to simulate the impact of earthquakes on buildings, the SIPHT workflow involves sequence analysis in bioinformatics, and MONTAGE is widely used in the processing of astronomical image data. These workflows not only play important roles in scientific research but also have different requirements for computing resources, thus providing rich test scenarios for the research of workflow scheduling algorithms. The Alibaba cluster workflow contains metadata and runtime information of 4000 machines, 71000 online services, and 4 million batch jobs within 8 days, including DAG information of production batch workloads.
[0158] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
Claims
1. An industrial manufacturing workflow scheduling method based on graph attention network and multi-agent reinforcement learning, characterized in that Including: Step S1: Obtain a workflow and build a workflow scheduling model based on the workload. Step S2: Build a reinforcement learning model according to the workflow scheduling model and the machine cluster resource model, and use the Q-learning algorithm based on deep learning for training to output the final allocation policy. Step S3: During the execution of the workflow tasks, according to the execution progress of the tasks and the load conditions of the machines, respond in real time to the changes in task progress and machine load, and dynamically update the execution order of the containers.
2. The industrial manufacturing workflow scheduling method based on graph attention network and multi-agent reinforcement learning according to claim 1, characterized in that, The said step S1 includes: Step S1-1: Model the workflow w i as a directed acyclic graph: G=(T, E) Where: G is a directed acyclic graph, T is the set of nodes of the directed acyclic graph, each node includes a task, E is the set of edges of the directed acyclic graph, and each edge represents the control dependency relationship between two tasks. Step S1-2: Perform workflow configuration in the form of a configuration file. The configuration content includes: the start time of the workflow, the included tasks, the fixed attributes of the tasks, and the dependency relationships between the tasks. Among them, the fixed attributes of the tasks include: the number of CPU cores required to run the task and the execution time. Step S1-3: Match the tasks in the workflow with the corresponding containers. For any task t i , if the set of containers that can be allocated when it runs is C = {c1, c2, …, c m}, where c m represents each independent container, and each task is calculated by an independent container. A workflow w i requires multiple containers to work together to complete; Step S1-4: Build a container execution queue through sub-deadlines to complete the construction of the workflow scheduling model.
3. The industrial manufacturing workflow scheduling method based on graph attention network and multi-agent reinforcement learning according to claim 2, characterized in that, In the said step S1-1, any task in the set of nodes of the directed acyclic graph is an indivisible task. If any task t i has an edge e with a data transfer volume of data ij ij which represents the execution order constraint between task t i and t j That is, the priority of subtask t j is lower than that of the parent task t i . 4. An industrial manufacturing workflow scheduling method based on graph attention network and multi-agent reinforcement learning according to claim 2, characterized in that, In the said step S1-4, the calculation method for the container sub-deadline is as follows Where: R m is the upward rank obtained for container c m according to the topology, and c n is a sub-container of container c m , and R n is the upward rank obtained for container c n according to the topology, T mn is the transfer time between containers, and t m * is the shortest execution time of the container, sd m is the sub-deadline of the container, d i is the deadline of the workflow, and R entry is the upward rank of the starting task.
5. The industrial manufacturing workflow scheduling method based on graph attention network and multi-agent reinforcement learning according to claim 2, characterized in that, The state space of the said reinforcement learning model includes the execution time of the container, the amount of data to be transmitted by the container, the expected start time of the container on this server, the weight value of the server, and the rental cost of the server. The action space A of the reinforcement learning model t is as follows: A t = {a t | a t = <c m , m k >}, c m ∈ C, m k ∈ M where: a t is an action, indicating that machine m k runs container c m , m k is the k-th machine; The reward function r of the said reinforcement learning model is: r = -ωcost(m k ) - (1 - ω)ECT(c m , m k ) Where: ω is the weight value of the table in the overall optimization goal, cost(m k ) is the cost of the machine, and ECT(c m , m k ) is the estimated completion time.
6. The industrial manufacturing workflow scheduling method based on graph attention network and multi-agent reinforcement learning according to claim 5, characterized in that Since the container features and machine features are a low-dimensional input, the Q network of the said reinforcement learning model includes: Perform an embedding representation on it through the structure of a multi-layer perceptron MLP, and map this indivisible data into high-dimensional features: h c = MLP(d c ) = σ(W c d c + b c ) h m = MLP(d m ) = σ(W m d m + b m ) Where: h c is the output feature vector of the container, h m is the output feature vector of the machine, MLP(d c ) is the Multilayer Perceptron network for container features, with the input being the feature vector d c of the container, MLP(d m ) is the Multilayer Perceptron network representing machine features, with the input being the feature vector d m of the machine, W c is the weight matrix in the Multilayer Perceptron network for container features, used to linearly transform the input container feature vector d c , W m is the weight matrix in the Multilayer Perceptron network for machine features, used to linearly transform the input machine feature vector d m , b s is the bias vector in the Multilayer Perceptron network for container features, b m is the bias vector in the Multilayer Perceptron network for machine features, σ(·) represents the activation function; Machine node feature extraction is achieved using a graph attention network. Assume that there are N machine nodes in the input graph, and the input feature of each machine node i is where D is the feature dimension of the node. For each node i, its attention weight α ij represents the relationship strength between machine node i and its neighbor node j, and this weight is calculated by the following formula: e ij = LeakReLU(a T [Wh i ||Wh j ) Where: W is the learned weight matrix for performing feature transformation on the input features, || represents the feature splicing operation, and a is the shared attention parameter for calculating the relationship between node pairs. Use the softmax function to normalize the coefficients of the entire neighborhood. The specific processing method is as follows: where: α ij is the attention coefficient of node i to its neighbor node j, obtained by normalization through the softmax function, and is used for subsequent weighted summation of the features of neighbor nodes. N(i) represents the set of neighbor nodes of node i; Update the features of node i based on all neighbor nodes: Wherein: represents the feature vector of the i-th container at the H-th layer, h′ m is the output feature vector after passing through multiple attention layers. Merge the container features and machine features to obtain h1, and output the final feature vector to the Q network through a two-layer perceptron: h1 = concat(h t , h' m ) h2 = MLP(h1) = σ(W1h1 + b1) h3 = MLP(h2) = σ(W2h2 + b2) Where: h1 represents combining the container feature h t and the machine feature h′ m Combining them, h2 represents the feature vector after being processed by the first - layer multi - layer perceptron (MLP), which is the non - linear transformation result of h1. h3 represents the feature vector after being processed by the second - layer multi - layer perceptron (MLP), which is the further non - linear transformation result of h2. Finally, it is output to the Q network, and h t is the feature vector of the container, W1 is the weight of the first - layer MLP, W2 is the weight of the second - layer MLP, b1 is the bias vector of the first - layer MLP, and b2 is the bias vector of the second - layer MLP.
7. An industrial manufacturing workflow scheduling method based on a graph attention network and multi-agent reinforcement learning according to claim 6, characterized in that, In the said step S2, during the training process of the Q-learning algorithm based on deep learning, Q values are used to evaluate the pros and cons of each state-action pair. The ultimate goal is to maximize the cumulative reward. The Q network generates a richer state representation through the graph attention layer during this process. During its training and update process, the input features include the machine resource information of the graph structure. The update formula of the network is as follows: Q(s t ,a t ) ← Q(s t ,a t ) + α * [r t + γ * max a′ Q(s t+1 ,a′) - Q(s t ,a t )] Where: Q(s t , a t ) represents the action-value function when taking action a t in state s t , r t is the reward at the current time t, γ is the discount factor, α is the learning rate, a′ represents the set of possible actions in the next state s t+1 , and the update objective is to minimize the difference between the predicted value of the evaluation Q network and the target Q value through backpropagation.
8. The industrial manufacturing workflow scheduling method based on graph attention network and multi-agent reinforcement learning according to claim 6, characterized in that, The loss function L(θ) of the training process of the said deep learning model is: L(θ) = E[(r t + γ * max a′ Q(s t+1 , a′|θ′) - Q(s t , a t |θ)) 2 ) where: r t is obtained after performing action a at the current time t, E[·] is the expected value of the loss function, and θ is the weight parameter for evaluating the Q-network. t 9. A workflow scheduling method for industrial manufacturing based on graph attention network and multi-agent reinforcement learning according to claim 8, characterized in that, In step S3, when an execution delay occurs in a certain container, the estimated start time and the estimated completion time of related subsequent containers are both updated in real time, and the update formula is as follows: where: Δt m is the delay time of the container, EST(c n ,m k ) is the estimated start time, ECT(c n ,m k ) is the estimated completion time, succ(c m ) represents the set of child nodes of the container.
10. An industrial manufacturing workflow scheduling device based on a graph attention network and multi-agent reinforcement learning, comprising a memory, a processor, and a program stored in the memory, characterized in that, When the processor executes the program, the method described in any one of claims 1-9 is implemented.
Citation Information
Patent Citations
Cloud computing workflow scheduling method and system, processing equipment and medium
CN118409848A