Multi-target flexible workshop scheduling method based on double-attention network and deep reinforcement learning

By adopting dual attention network and deep reinforcement learning methods in flexible operation workshops, a multi-objective optimization model is built, and the scheduling problem of multi-objective flexible operation workshops is solved, achieving efficient production scheduling and cost reduction.

CN120010417APending Publication Date: 2025-05-16HEFEI UNIV OF TECH

Patent Information

Application Number
CN202510172701.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art is difficult to effectively solve the problem of multi-objective flexible operation workshop scheduling, resulting in waste of resources and increased costs.

Method used

The static flexible workshop scheduling method based on dual attention network and deep reinforcement learning is adopted. By constructing a multi-objective optimization model, the dual attention network is used to extract the characteristic information of the process and machine, and the optimal scheduling scheme is obtained through the actor critic network.

Benefits of technology

It realizes a scheduling solution that can not only minimize time and minimize the total machine load, significantly improving production efficiency and reducing production costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120010417A_ABST
    Figure CN120010417A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-target flexible workshop scheduling method based on a double-attention network and deep reinforcement learning. The method comprises the following steps: 1, acquiring basic information in flexible job shop scheduling; 2, constructing a workshop scheduling mathematical model based on a target function and constraint conditions; 3, formalizing a flexible job shop scheduling problem into a Markov decision process, and defining state parameters, action parameters, state transition parameters and reward parameters in a deep reinforcement learning method; 4, constructing a double-attention network, and extracting feature information of procedures and machines; and 5, obtaining an optimal scheduling scheme by using a deep reinforcement learning algorithm. According to the invention, the scheduling scheme with shortest time consumption and minimum machine load can be obtained, direct learning from original data to decision can be realized, dependence on expert knowledge is reduced, production scheduling time is reduced, workshop generation efficiency is improved, and decision basis is provided for actual production scheduling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of workshop production scheduling, and in particular relates to a multi-objective flexible workshop static scheduling method, which is used for scheduling and management of workshop production plans. Background Art

[0002] With the rapid development of modern manufacturing, the fixed mode of traditional workshop scheduling has indeed become difficult to adapt to the diversified and personalized production needs. In this context, the importance of Flexible Job Shop Scheduling (FJSP) has become increasingly prominent. As a core component of the flexible manufacturing system, the core concept of FJSP is flexibility and adaptability. A good scheduling method can not only quickly adjust the configuration of the production line according to production needs, but also optimize resource allocation, thereby improving production efficiency.

[0003] In current production practice, although some studies have used deep reinforcement learning (DRL) to solve the FJSP problem, the quality of these solutions relative to precise methods (such as OR-Tools) still has room for improvement. In addition, the solutions to the single-objective flexible job shop scheduling problem often fail to fully balance the various influencing factors in the production process, such as production speed, machine load, and energy consumption optimization. This leads to waste of resources and increased costs, making multi-objective optimization an important research direction. Summary of the invention

[0004] In order to overcome the shortcomings of the prior art, the present invention proposes a multi-objective flexible workshop scheduling method based on dual attention network and deep reinforcement learning, in order to obtain a high-quality production scheduling plan, thereby improving the overall production efficiency and reducing the operating cost of the workshop.

[0005] To achieve the above object, the technical solution of the present invention is as follows:

[0006] The static flexible workshop scheduling method based on dual attention network and deep reinforcement learning of the present invention is characterized in that it comprises the following steps:

[0007] Step 1: Obtain the basic information of the flexible job shop scheduling, including: all workpieces, the process corresponding to each workpiece, and the processing time of each process on different machines; where the total number of workpieces is n; the total number of machines is m; i, Respectively represent the two workpiece numbers, i, =1,2,…,n;j, Respectively represent the sequence numbers of the two processes, j, =1,2,…,n i , n irepresents the i-th workpiece J i The number of processes; M k represents the kth machine, where k represents the machine number, k=1,2,…,m; the i-th workpiece J i The jth process is O ij ;

[0008] Step 2: Based on the basic information, taking minimizing the maximum completion time and minimizing the total machine load as the objective function, a flexible job shop scheduling model based on the objective function and constraint conditions is constructed;

[0009] Step 3: Convert the flexible job shop scheduling model into a Markov decision process, and define states, actions, state transitions, and reward values;

[0010] Step 4: Based on the state in the Markov decision process, a dual attention network is constructed to extract feature information of the process and the machine;

[0011] Step 5: Based on the actions in the Markov decision process and the feature information extracted by the dual attention network, an actor-critic network is constructed to obtain the probability distribution of actions in a given state and select actions with higher cumulative reward values, thereby obtaining the optimal scheduling solution.

[0012] The static flexible workshop scheduling method based on dual attention network and deep reinforcement learning described in the present invention is also characterized in that step 2 includes:

[0013] Step 2.1: Use formula (1) to establish the first objective function f1 of minimum and maximum completion time:

[0014] (1)

[0015] In formula (1), C i represents the i-th workpiece J i completion time;

[0016] The second objective function f2 of minimum machine total load is established using formula (2):

[0017] (2)

[0018] In formula (2), x ijk A Boolean value indicating that the kth machine M k Whether to process the i-th workpiece J i The jth process of the kth machine M k Processing the i-th workpiece J i The jth process O ij , then let x ijk =1, otherwise, let xijk =0; represents the kth machine M k Processing the i-th workpiece J i The jth process O ij The time taken;

[0019] Step 2.2: Use equations (3) and (4) to construct the process sequence constraints for each workpiece:

[0020] (3)

[0021] (4)

[0022] In formula (3)-formula (4), s ij represents the i-th workpiece J i The jth process O ij The processing start time; c ij represents the i-th workpiece J i The jth process O ij The processing completion time; represents the i-th workpiece J i The j+1th process O i(j+1) The processing start time;

[0023] Use formula (5) to construct the completion time constraint of the workpiece:

[0024] (5)

[0025] In formula (5), c ij represents the i-th workpiece J i The jth process O ij The processing completion time; C max represents the maximum completion time;

[0026] Using equations (6) and (7), we construct the constraint that the same machine can only process one process at the same time:

[0027] (6)

[0028] (7)

[0029] In formula (6)-formula (7), B is a positive number; represents the i-th workpiece J i The jth process O ij Is it before the Workpieces No. Process On the kth machine M kOn processing, if before, then let =1, otherwise, let =0;

[0030] Formula (8) is used to construct the constraint that the same process can only be processed by one machine at the same time:

[0031] (8)

[0032] In formula (8), m ij represents the i-th workpiece J i The jth process O ij The number of optional processing machines.

[0033] Furthermore, the step 3 comprises:

[0034] Step 3.1: Define the state s at time t t Includes: the i-th workpiece J i The jth process O ij The eigenvector of ; The kth machine M k The eigenvector of ; Compatible process-machine pairs (O ij , M k )’s eigenvector h(O ij , M k );

[0035] Step 3.2: Define action a at time t t is the process-machine pair (O ij , M k ), indicating the use of the kth machine M at time t k Processing the i-th workpiece J i The jth process O ij ;

[0036] Step 3.3, define state transition: when executing action a at time t t After that, update the state s at time t t , and get the state s at time t+1 t+1 ;

[0037] Step 3.4: Use formula (9) to construct the reward function at time t :

[0038] (9)

[0039] In formula (9), w1 and w2 are two weight coefficients used to balance the completion time reward and load reward; Indicates the current state s t and the next state st+1 The maximum completion time difference between them; Indicates the current state s t and the next state s t+1 The difference in the total load of the machines between them;

[0040] Step 3.5: Define strategy π(a t |s t ), indicating a given state s t Next select action a t probability.

[0041] Furthermore, the dual attention network in step 4 includes: L-layer process information attention blocks, L-layer machine information attention blocks, and a global feature fusion module:

[0042] Step 4.1: Define and initialize the current layer ; Define the current The i-th workpiece J output by the layer process information attention block i The jth process O ij The eigenvector of , and initialize ; Define the current The k-th machine M output by the attention block of the layer machine information k The eigenvector of , and initialize ;

[0043] Step 4.2: The process information attention block is used to capture the dependencies between processes and encode them into the feature representation of the process, thereby updating , get the The i-th workpiece J output by the layer process information attention block i The jth process O ij The eigenvector of ;

[0044] Step 4.3: The machine information attention block is used to capture the competitive relationship between machines and encode it into the feature representation of the machine, thereby updating , get the The k-th machine M output by the layer process information attention block k The eigenvector of ;

[0045] Step 4.4: Assign to Then, return to step 4.2 and execute sequentially until the i-th workpiece J output by the L-th layer process information attention block is obtained. i The jth process O ij The eigenvector of And the kth machine M output by the Lth layer machine information attention block k The eigenvector of , so that the global feature fusion module uses formula (15) to obtain the global feature representation :

[0046] (15)

[0047] In formula (15), O u Represents a set of unfinished processes. Indicates the number of processes in the unfinished process set; M u represents the set of idle machines, The number of machines representing the idle machine set.

[0048] Further, the step 4.2 includes:

[0049] Step 4.2.1. Calculate the i-th workpiece J using formula (10): i The jth process O ij The first Layer Attention Coefficient , and use the softmax function to Normalize it and get the normalized Layer Attention Coefficient ;

[0050] (10)

[0051] In formula (10), is the weight vector of the attention mechanism; represents the i-th workpiece J i The pth process O ip No. Layer feature vector, where p takes values ​​of j-1, j, and j+1; represents the linear transformation matrix; || represents the vector connection operation; LeakyReLU is the activation function; T represents transposition;

[0052] Step 4.2.2: Use formula (11) to obtain Layer i workpiece J i The jth process O ij The eigenvector of :

[0053] (11)

[0054] In formula (11), σ represents the ELU nonlinear activation function.

[0055] Further, the step 4.3 includes:

[0056] Step 4.3.1. Calculate the kth machine M using formula (12) k and the qth machine M q The first Layer competition intensity :

[0057] (12)

[0058] In formula (12), represents the distance between the kth machine and the qth machine. Layer process competition set; Indicates the unscheduled first Layer process collection;

[0059] Step 4.3.2: Use formula (13) to calculate the kth machine M k and the qth machine M q The first Layer Attention Coefficient , and use the softmax function to Normalize and get the normalized attention coefficient :

[0060] (13)

[0061] In formula (13), represents the qth machine M q No. layer feature vector; represents the first linear transformation matrix; represents the second linear transformation matrix; Represents the weight vector of the attention mechanism;

[0062] Step 4.3.3: Use formula (14) to get The kth machine M of the layer k The eigenvector of :

[0063] (14)

[0064] In formula (14), N k Represents all the connections with the kth machine M k There is a competing relationship with the qth machine.

[0065] Furthermore, the actor-critic network in step 5 includes: an actor network and a critic network:

[0066] Step 5.1: The actor network uses formula (16) to calculate the state s t Next select action a t Tendency , and use the softmax function to Normalize and get the state s at time t t Next action a t The probability distribution of being selected , and as the t-time strategy of the actor network;

[0067] (16)

[0068] In formula (16), MLP θ represents a multilayer perceptron; θ represents the parameters of the actor network;

[0069] Step 5.2: The critic network is based on the global features , using formula (17) to calculate in state s t The t-time strategy is used at all subsequent times The accumulated depreciation bonus value ;

[0070] (17)

[0071] In formula (17), is the discount factor; is The reward value obtained at the moment; D represents the moment when the accumulated depreciation reward value reaches convergence; represents the strategy at time t The mathematical expectation of represents the parameters of the critic network;

[0072] Step 5.3: Use the proximal strategy optimization algorithm to train the parameters of the actor-critic network and obtain the actor-critic network corresponding to the optimal parameters, thereby outputting the optimal scheduling plan.

[0073] The electronic device of the present invention includes a memory and a processor, wherein the memory is used to store a program that supports the processor to execute the multi-objective flexible workshop scheduling method, and the processor is configured to execute the program stored in the memory.

[0074] The present invention provides a computer-readable storage medium, wherein a computer program is stored on the computer-readable storage medium, and the computer program executes the steps of the multi-objective flexible workshop scheduling method when the computer program is executed by a processor.

[0075] Compared with the prior art, the present invention has the following beneficial effects:

[0076] 1. The present invention fully considers the actual production situation of the flexible operation workshop, introduces multiple constraints, and constructs a multi-objective optimization model. By using a deep reinforcement learning algorithm based on a dual attention network to solve the model, the present invention can derive a scheduling solution that takes the shortest time and minimizes the total machine load. This solution significantly improves production efficiency while reducing production costs.

[0077] 2. This invention proposes a compact state representation method for accurately describing the process and machine information in FJSP. As the scheduling process proceeds, the state space will gradually decrease, and the optimal scheduling plan can be found more quickly during the optimization of the workshop scheduling plan, thereby reducing the scheduling time.

[0078] 3. The present invention designs a dual attention network, which is composed of multiple process and machine message attention blocks, and can perform in-depth feature extraction on processes and machines, thereby improving the efficiency of workshop production. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] Figure 1 It is an overall flow chart of the dual attention network and deep reinforcement learning algorithm of the present invention;

[0080] Figure 2 is a flow chart of the dual attention network of the present invention;

[0081] Figure 3 Flowchart of the actor-critic network of the present invention. DETAILED DESCRIPTION

[0082] In this embodiment, a multi-objective flexible workshop scheduling method based on dual attention network and deep reinforcement learning is to minimize the maximum completion time and minimize the total machine load, and provide a solution for flexible workshop scheduling. It introduces dual attention network and deep reinforcement learning to achieve efficient optimization of flexible workshop scheduling problems. This method can not only find the scheduling solution with the shortest time and the smallest machine load, but also, as the scheduling proceeds, its compact state representation allows the state space to be dynamically reduced, reducing the scheduling time and improving the workshop production efficiency. Specifically, Figure 1 As shown, the method comprises the following steps:

[0083] Step 1: Obtain the basic information of the flexible job shop scheduling, including: all workpieces, the process corresponding to each workpiece, and the processing time of each process on different machines; where the total number of workpieces is n; the total number of machines is m; i, Respectively represent the two workpiece numbers, i, =1,2,…,n;j, Respectively represent the sequence numbers of the two processes, j, =1,2,…,n i , n i represents the i-th workpiece J i The number of processes; M k represents the kth machine, where k represents the machine number, k=1,2,…,m; the i-th workpiece J i The jth process is O ij ;

[0084] Step 2: Construct a flexible job shop scheduling model based on objective function and constraints;

[0085] Step 2.1: Use formula (1) to establish the first objective function f1 of minimum and maximum completion time:

[0086] (1)

[0087] In formula (1), C i represents the i-th workpiece J i completion time;

[0088] The second objective function f2 of minimum machine total load is established using formula (2):

[0089] (2)

[0090] In formula (2), x ijk A Boolean value indicating that the kth machine M k Whether to process the i-th workpiece J i The jth process of the kth machine M k Processing the i-th workpiece J i The jth process O ij , then let x ijk =1, otherwise, let x ijk =0; represents the kth machine M k Processing the i-th workpiece J i The jth process O ij The time taken.

[0091] Step 2.2: Use equations (3) and (4) to construct the process sequence constraints for each workpiece:

[0092] (3)

[0093] (4)

[0094] In formula (3)-formula (4), s ij represents the i-th workpiece J i The jth process O ijThe processing start time; c ij represents the i-th workpiece J i The jth process O ij The processing completion time; represents the i-th workpiece J i The j+1th process O i(j+1) The processing start time.

[0095] Use formula (5) to construct the completion time constraint of the workpiece:

[0096] (5)

[0097] In formula (5), c ij represents the i-th workpiece J i The jth process O ij The processing completion time; C max represents the maximum completion time;

[0098] Using equations (6) and (7), we construct the constraint that the same machine can only process one process at the same time:

[0099] (6)

[0100] (7)

[0101] In formula (6)-formula (7), B is a positive number; represents the i-th workpiece J i The jth process O ij Is it before the Workpieces No. Process On the kth machine M k On processing, if before, then let =1, otherwise, let =0.

[0102] Formula (8) is used to construct the constraint that the same process can only be processed by one machine at the same time:

[0103] (8)

[0104] In formula (8), m ij represents the i-th workpiece J i The jth process O ij The number of optional processing machines.

[0105] Step 3: Convert the scheduling model into a Markov decision process and define states, actions, state transitions, and reward values;

[0106] Step 3.1: Define the state s at time t t Includes: the i-th workpiece J i The jth process O ij The eigenvector of , including 4 static features and 6 dynamic features; the kth machine M k The eigenvector of , including 2 static features and 6 dynamic features; compatible process-machine pairs (O ij , M k )’s eigenvector h(O ij , M k ), including 2 static features and 6 dynamic features. As production scheduling proceeds, the processes that have been scheduled in the state space and will not have an impact will be removed;

[0107] The i-th workpiece J i The jth process O ij The static characteristics of are: on all available machines, the i-th job J i The jth process O ij The minimum processing time, average processing time, the difference between the maximum and minimum processing time, and the proportion of machines that can be processed. The dynamic characteristics are: the i-th workpiece J i The jth process O ij Whether it is processed, the earliest possible completion time, waiting time, and remaining processing time; the i-th workpiece J i The number of unprocessed processes and the remaining workload.

[0108] The kth machine M k The static characteristics of the k-th machine M are: k The shortest processing time and average processing time of the k-th machine M k The number of unprocessed processes, the number of candidate processes, idle time, waiting time, whether idle; remaining processing time.

[0109] Process-Machine Pair (O ij , M k ) have the following static features: the i-th workpiece J i The jth process O ij On the kth machine M k The dynamic characteristics are: the i-th workpiece J i The jth process O ij On the kth machine M kThe ratio of processing time to the maximum processing time of the machine's candidate processes, the maximum processing time of the unprocessed processes, the maximum processing time of the unprocessed processes that the machine can handle, the ratio of processing time to the maximum processing time of the process-machine compatibility pair, the ratio of processing time to the remaining workload of the job, and the total waiting time.

[0110] Step 3.2: Define action a at time t t is the process-machine pair (O ij , M k ), indicating the use of the kth machine M at time t k Processing the i-th workpiece J i The jth process O ij ;

[0111] Step 3.3, define state transition: when executing action a at time t t After that, update the state s at time t t , and get the state s at time t+1 t+1 ;

[0112] Step 3.4: Use formula (9) to construct the reward function at time t :

[0113] (9)

[0114] In formula (9), w1 and w2 are two weight coefficients used to balance the completion time reward and load reward; Indicates the current state s t and the next state s t+1 The maximum completion time difference between them; Indicates the current state s t and the next state s t+1 The difference in the total machine load between the two.

[0115] Step 3.5: Define strategy π(a t |s t ), indicating a given state s t Next select action a t probability.

[0116] Step 4: Based on the state of the Markov decision process, construct an L-layer dual attention network to extract the feature information of the process and machine and fuse the global features, such as Figure 2 As shown, this is conducive to accurately and concisely representing the complex relationship between processes and machines and improving the quality of decision-making;

[0117] Step 4.1: Define and initialize the current layer ; Define the current The i-th workpiece J output by the layer process information attention blocki The jth process O ij The eigenvector of , and initialize ; Define the current The k-th machine M output by the attention block of the layer machine information k The eigenvector of , and initialize ;

[0118] Step 4.2: The process information attention block is used to capture the dependencies between processes and encode them into the feature representation of the process, thereby updating , get the The i-th workpiece J output by the layer process information attention block i The jth process O ij The eigenvector of ;

[0119] Step 4.2.1. Calculate the i-th workpiece J using formula (10): i The jth process O ij The first Layer Attention Coefficient , and use the softmax function to Normalize it and get the normalized Layer Attention Coefficient , in order to distribute attention among different processes;

[0120] (10)

[0121] In formula (10), is the weight vector of the attention mechanism; represents the i-th workpiece J i The pth process O ip No. Layer feature vector, where p takes values ​​of j-1, j, and j+1; Represents the linear transformation matrix, which is used to transform the eigenvector and Make linear changes; || represents vector concatenation operation; LeakyReLU is the activation function; T represents transposition.

[0122] Step 4.2.2: Use formula (11) to obtain Layer i workpiece J i The jth process O ij The eigenvector of :

[0123] (11)

[0124] In formula (11), σ represents the ELU nonlinear activation function;

[0125] Step 4.3: The machine information attention block is used to capture the competitive relationship between machines and encode it into the feature representation of the machine, thereby updating , get the The k-th machine M output by the layer process information attention block k The eigenvector of ;

[0126] Step 4.3.1. Calculate the kth machine M using formula (12) k and the qth machine M q The first Layer competition intensity :

[0127] (12)

[0128] In formula (12), represents the distance between the kth machine and the qth machine. The layer process competition set represents the set of processes that can be processed by both the k-th machine and the q-th machine; Indicates the unscheduled first Layer process collection;

[0129] Step 4.3.2: Use formula (13) to calculate the kth machine M k and the qth machine M q The first Layer Attention Coefficient , and use the softmax function to Normalize and get the normalized attention coefficient :

[0130] (13)

[0131] In formula (13), represents the qth machine M q The eigenvector of represents the first linear transformation matrix for the eigenvector and Do linear transformations; Represents the second linear transformation matrix, which is used to transform the k-th machine M k and the qth machine M q The intensity of competition between Do linear transformations; Represents the weight vector of the attention mechanism;

[0132] Step 4.3.3: Use formula (14) to get The kth machine M of the layer k The eigenvector of :

[0133] (14)

[0134] In formula (14), N k Represents all the connections with the kth machine M k The qth machine in a competitive relationship;

[0135] Step 4.4: Assign to Then, return to step 4.2 and execute sequentially until the i-th workpiece J output by the L-th layer process information attention block is obtained. i The jth process O ij The eigenvector of And the kth machine M output by the Lth layer machine information attention block k The eigenvector of , so that the global feature fusion module uses formula (15) to obtain the global feature representation :

[0136] (15)

[0137] In formula (15), O u Represents a set of unfinished processes. Indicates the number of processes in the unfinished process set; M u represents the set of idle machines, The number of machines representing the idle machine set.

[0138] Step 5: Based on the actions in the Markov decision process and the feature information extracted by the dual attention network, an actor-critic network is constructed, such as Figure 3 As shown, the probability distribution of actions in a given state is obtained, and the action with a higher cumulative reward value is selected after evaluating the reward value of the action to obtain the optimal scheduling plan;

[0139] Step 5.1: The actor network uses formula (16) to calculate the state s t Next select action a t Tendency , and use the softmax function to Normalize and get the state s at time t t Next action a t The probability distribution of being selected , and as the t-time strategy of the actor network;

[0140] (16)

[0141] In formula (16), MLP θ represents a multilayer perceptron; θ represents the parameters of the actor network;

[0142] Step 5.2: The critic network is based on the global features , using formula (17) to calculate in state s t The t-time strategy is used at all subsequent times The accumulated depreciation bonus value ;

[0143] (17)

[0144] In formula (17), is the discount factor; is The reward value obtained at the moment; D represents the moment when the cumulative depreciation reward value reaches convergence; represents the strategy at time t The mathematical expectation of represents the parameters of the critic network.

[0145] Step 5.3: Use the proximal policy optimization algorithm to train the parameters of the actor-critic network and obtain the actor-critic network corresponding to the optimal parameters, thereby outputting the optimal scheduling solution. The proximal policy optimization algorithm is an optimization algorithm based on policy gradients. It improves the stability of the training process by limiting the amplitude of policy updates and improves training efficiency. The actor-critic network corresponding to the optimal parameters is trained and generates the probability distribution of selecting the optimal action under a given state. The combination of the optimal actions constitutes the optimal scheduling solution with the maximum cumulative reward value.

[0146] In this embodiment, an electronic device includes a memory and a processor, wherein the memory is used to store a program that supports the processor to execute the above method, and the processor is configured to execute the program stored in the memory.

[0147] In this embodiment, a computer-readable storage medium stores a computer program on the computer-readable storage medium, and the computer program executes the steps of the above method when executed by a processor.

Claims

1. A multi-objective flexible workshop scheduling method based on dual attention network and deep reinforcement learning, characterized in that: The steps include: Step 1: Obtain the basic information of the flexible job shop scheduling, including: all workpieces, the process corresponding to each workpiece, and the processing time of each process on different machines; where the total number of workpieces is n; the total number of machines is m; i, Respectively represent the two workpiece numbers, i, =1,2,…,n;j, Respectively represent the sequence numbers of the two processes, j, =1,2,…,n i , n i represents the i-th workpiece J i The number of processes; M k represents the kth machine, where k represents the machine number, k=1,2,…,m; the i-th workpiece J i The jth process is O ij ; Step 2: Based on the basic information, taking minimizing the maximum completion time and minimizing the total machine load as the objective function, a flexible job shop scheduling model based on the objective function and constraint conditions is constructed; Step 3: Convert the flexible job shop scheduling model into a Markov decision process, and define states, actions, state transitions, and reward values; Step 4: Based on the state in the Markov decision process, a dual attention network is constructed to extract feature information of the process and the machine; Step 5: Based on the actions in the Markov decision process and the feature information extracted by the dual attention network, an actor-critic network is constructed to obtain the probability distribution of actions in a given state and select actions with higher cumulative reward values, thereby obtaining the optimal scheduling solution.

2. The multi-objective flexible workshop scheduling method based on dual attention network and deep reinforcement learning according to claim 1 is characterized in that: The step 2 comprises: Step 2.1: Use formula (1) to establish the first objective function f1 of minimum and maximum completion time: (1) In formula (1), C i represents the i-th workpiece J i completion time; The second objective function f2 of minimum machine total load is established using formula (2): (2) In formula (2), x ijk A Boolean value indicating that the kth machine M k Whether to process the i-th workpiece J i The jth process of the kth machine M k Processing the i-th workpiece J i The jth process O ij , then let x ijk =1, otherwise, let x ijk =0; represents the kth machine M k Processing the i-th workpiece J i The jth process O ij The time taken; Step 2.2: Use equations (3) and (4) to construct the process sequence constraints for each workpiece: (3) (4) In formula (3)-formula (4), s ij represents the i-th workpiece J i The jth process O ij The processing start time; c ij represents the i-th workpiece J i The jth process O ij The processing completion time; represents the i-th workpiece J i The j+1th process O i(j+1) The processing start time; Use formula (5) to construct the completion time constraint of the workpiece: (5) In formula (5), c ij represents the i-th workpiece J i The jth process O ij The processing completion time; C max represents the maximum completion time; Using equations (6) and (7), we construct the constraint that the same machine can only process one process at the same time: (6) (7) In formula (6)-formula (7), B is a positive number; represents the i-th workpiece J i The jth process O ij Is it before the Workpieces No. Process On the kth machine M k On processing, if before, then let =1, otherwise, let =0; Formula (8) is used to construct the constraint that the same process can only be processed by one machine at the same time: (8) In formula (8), m ij represents the i-th workpiece J i The jth process O ij The number of optional processing machines.

3. The multi-objective flexible workshop scheduling method based on dual attention network and deep reinforcement learning according to claim 2 is characterized in that: The step 3 comprises: Step 3.1: Define the state s at time t t Includes: the i-th workpiece J i The jth process O ij The eigenvector of ; The kth machine M k The eigenvector of ; Compatible process-machine pairs (O ij , M k )’s eigenvector h(O ij , M k ); Step 3.2: Define action a at time t t is the process-machine pair (O ij , M k ), indicating the use of the kth machine M at time t k Processing the i-th workpiece J i The jth process O ij ; Step 3.3, define state transition: when executing action a at time t t After that, update the state s at time t t , and get the state s at time t+1 t+1 ; Step 3.4: Use formula (9) to construct the reward function at time t : (9) In formula (9), w1 and w2 are two weight coefficients used to balance the completion time reward and load reward; Indicates the current state s t and the next state s t+1 The maximum completion time difference between them; Indicates the current state s t and the next state s t+1 The difference in the total load of the machines between them; Step 3.5: Define strategy π(a t |s t ), indicating a given state s t Next select action a t probability.

4. The multi-objective flexible workshop scheduling method based on dual attention network and deep reinforcement learning according to claim 3 is characterized in that: The dual attention network in step 4 includes: L-layer process information attention blocks, L-layer machine information attention blocks, and a global feature fusion module: Step 4.1: Define and initialize the current layer ; Define the current The i-th workpiece J output by the layer process information attention block i The jth process O ij The eigenvector of , and initialize ; Define the current The k-th machine M output by the attention block of the layer machine information k The eigenvector of , and initialize ; Step 4.2: The process information attention block is used to capture the dependencies between processes and encode them into the feature representation of the process, thereby updating , get the The i-th workpiece J output by the layer process information attention block i The jth process O ij The eigenvector of ; Step 4.3: The machine information attention block is used to capture the competitive relationship between machines and encode it into the feature representation of the machine, thereby updating , get the The k-th machine M output by the layer process information attention block k The eigenvector of ; Step 4.4: Assign to Then, return to step 4.2 and execute sequentially until the i-th workpiece J output by the L-th layer process information attention block is obtained. i The jth process O ij The eigenvector of And the kth machine M output by the Lth layer machine information attention block k The eigenvector of , so that the global feature fusion module uses formula (15) to obtain the global feature representation : (15) In formula (15), O u Represents a set of unfinished processes. Indicates the number of processes in the unfinished process set; M u represents the set of idle machines, The number of machines representing the idle machine set.

5. The multi-objective flexible workshop scheduling method based on dual attention network and deep reinforcement learning according to claim 4 is characterized in that: The step 4.2 comprises: Step 4.2.

1. Calculate the i-th workpiece J using formula (10): i The jth process O ij The first Layer Attention Coefficient , and use the softmax function to Normalize it and get the normalized Layer Attention Coefficient ; (10) In formula (10), is the weight vector of the attention mechanism; represents the i-th workpiece J i The pth process O ip No. Layer feature vector, where p takes values ​​of j-1, j, and j+1; represents the linear transformation matrix; || represents the vector connection operation; LeakyReLU is the activation function; T represents transposition; Step 4.2.2: Use formula (11) to obtain Layer i workpiece J i The jth process O ij The eigenvector of : (11) In formula (11), σ represents the ELU nonlinear activation function.

6. The multi-objective flexible workshop scheduling method based on dual attention network and deep reinforcement learning according to claim 5 is characterized in that: The step 4.3 comprises: Step 4.3.

1. Calculate the kth machine M using formula (12) k and the qth machine M q The first Layer competition intensity : (12) In formula (12), represents the distance between the kth machine and the qth machine. Layer process competition set; Indicates the unscheduled first Layer process collection; Step 4.3.2: Use formula (13) to calculate the kth machine M k and the qth machine M q The first Layer Attention Coefficient , and use the softmax function to Normalize and get the normalized attention coefficient : (13) In formula (13), represents the qth machine M q No. layer feature vector; represents the first linear transformation matrix; represents the second linear transformation matrix; Represents the weight vector of the attention mechanism; Step 4.3.3: Use formula (14) to get The kth machine M of the layer k The eigenvector of : (14) In formula (14), N k Represents all the connections with the kth machine M k There is a competing relationship with the qth machine.

7. The multi-objective flexible workshop scheduling method based on dual attention network and deep reinforcement learning according to claim 6 is characterized in that: The actor-critic network in step 5 includes: actor network, critic network: Step 5.1: The actor network uses formula (16) to calculate the state s t Next select action a t Tendency , and use the softmax function to Normalize and get the state s at time t t Next action a t The probability distribution of being selected , and as the t-time strategy of the actor network; (16) In formula (16), MLP θ represents a multilayer perceptron; θ represents the parameters of the actor network; Step 5.2: The critic network is based on the global features , using formula (17) to calculate in state s t The t-time strategy is used at all subsequent times The accumulated depreciation bonus value ; (17) In formula (17), is the discount factor; is The reward value obtained at the moment; D represents the moment when the cumulative depreciation reward value reaches convergence; represents the strategy at time t The mathematical expectation of represents the parameters of the critic network; Step 5.3: Use the proximal strategy optimization algorithm to train the parameters of the actor-critic network and obtain the actor-critic network corresponding to the optimal parameters, thereby outputting the optimal scheduling plan.

8. An electronic device, comprising a memory and a processor, characterized in that: The memory is used to store a program that supports the processor to execute the multi-objective flexible shop scheduling method described in any one of claims 1 to 7, and the processor is configured to execute the program stored in the memory.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the multi-objective flexible shop scheduling method described in any one of claims 1 to 7 are executed.

Citation Information

Patent Citations

  • Super-resolution reconstruction algorithm for medical imaging

    CN110717856A

  • Flexible job shop real-time scheduling method based on Dueling architecture deep reinforcement learning

    CN116880425A

  • Cloud cluster resource scheduling method based on deep reinforcement learning

    CN117555683A

  • Flexible workshop operation dynamic scheduling method based on deep reinforcement learning

    CN117892969A

  • Wind power assembly workshop multi-objective optimization scheduling method based on reinforcement learning

    CN118690897A

Cited By

  • State representation modeling method and system based on multi-domain graph attention network

    CN120672037A

  • Flexible job shop multi-target scheduling method and system based on preference driving

    CN121119638A

  • Task scheduling management method and system for AI finance and taxation robot

    CN122132143A