An Adaptive Scheduling Method for Job Shop Based on Deep Reinforcement Learning
By applying deep reinforcement learning and near-end strategy optimization algorithms in the work workshop, the problem of recalculating the scheduling solution when scale changes is solved, efficient and adaptive adaptive scheduling is achieved, and scheduling efficiency and quality are significantly improved.
Patent Information
- Application Number
- CN202210406935.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-18
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-04-18
AI Technical Summary
When solving the problem of scheduling in the work workshop, it is difficult to balance the time cost and algorithm quality, and the scheduling plan needs to be recalculated when the scale changes, resulting in a shutdown of production resources.
Adaptive scheduling method based on deep reinforcement learning is adopted, and the scheduling function model for job workshop scheduling problems is constructed, and combined with action optimization search strategies and asynchronous update mechanisms, a direct and efficient exploration and asynchronous update near-end strategy optimization algorithm (E2APPO) is proposed to minimize completion time, with the optimization goal.
Efficient adaptive scheduling is achieved, with a 5.6% increase in scheduling score, a minimum completion time reduced by 8.9% compared with traditional deep Q network algorithms, and is adapted to workshop environments of different sizes and complexities.
Smart Images

Figure CN114707881B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of job shop adaptive scheduling, and relates to a job shop adaptive scheduling method based on deep reinforcement learning. Background Art
[0002] With the development of information technology in the manufacturing industry, intelligent manufacturing and reconfigurable manufacturing have emerged as the times require. The job shop scheduling problem has attracted much attention because it can optimally allocate limited resources and improve production efficiency. The JSSP is essentially a combinatorial optimization problem, and traditionally, it is divided into exact algorithms (mathematical methods) and myopic algorithm methods. The exact algorithms for solving the JSSP are mainly based on operations research, such as mathematical programming method, Lagrangian relaxation method, and branch and bound method, etc. These methods can theoretically obtain the optimal solution. However, because this method requires precise modeling and a large amount of calculation, most of them still remain at the theoretical level and cannot be applied to actual production.
[0003] To solve this problem, many scholars have shifted their attention to approximation algorithms, such as priority rules or metaheuristic algorithms. These priority rules, such as First In First Out (FIFO), Longest Processing Time (LPT), Most Operation Remaining (MOPR), Most Work Remaining (MWKR), etc., are faster in calculation speed and can naturally handle the uncertainties in practice, but are prone to short-sightedness and fall into local optima, making it difficult to obtain the global optimal solution. When the scheduling scale expands, it will lead to a decline in the quality of the scheduling solution. Scholars have also proposed many composite rules based on domain knowledge, which have shown good scheduling performance. Designing an effective composite scheduling rule requires a large amount of prior knowledge and a lot of time. In terms of metaheuristic algorithms, there are many swarm intelligence algorithms, such as genetic algorithms, particle swarm algorithms, and ant colony algorithms, etc. These algorithms can obtain relatively optimal solutions through continuous exploration and iteration. However, the same problem faced by metaheuristics and priority rules is that once the scale of the scheduling problem changes, the scheduling scheme is no longer applicable and needs to be recalculated and solved. In large-scale production, it is unimaginable to stop production resources for a long time or even several hours for the scheduling scheme.
[0004] To seek a balance between time cost and algorithm quality, reinforcement learning (RL) has been proposed to train scheduling models and has achieved many successful applications in actual scheduling cases. There are still two issues that need attention. First, due to the existence of manual metrics, the feature extraction of the workshop state will be affected by humans. Second, taking scheduling rules as the action space, since the selection of the job sequence returns to the selection of rules, it will inevitably consume more time.
[0005] Many scholars have applied reinforcement learning (RL) to the research of scheduling strategies, providing new ways and directions for efficient decision-making in job shop scheduling. Reinforcement learning (RL) is an unsupervised learning that does not require pre-prepared labeled data. It has unique advantages in situations where labeled data is difficult to collect and obtain. A job shop can be regarded as a similar scenario, where an agent selects actions based on the current workshop state. The job shop scheduling process can be transformed into a Figure 1 Markov decision process (MDP) as shown, whose key elements are state, action, and reward.
[0006] The applications of RL in scheduling can be mainly divided into the following four categories. First, combine reinforcement learning (RL) with heuristic algorithms to improve algorithm performance by optimizing algorithm parameters; second, combine reinforcement learning (RL) with priority rules and design the rule set as the action space; reinforcement learning (RL) is used to find the optimal rule at each scheduling point to achieve the optimal strategy. Third, directly design the processing procedures of workpieces as the action space. Reinforcement learning (RL) directly selects the processing procedures at each scheduling point, that is, the optimal solution is obtained. Finally, define the machine ID or the transferred material as the action space selected by the agent. The above categories usually correspond to four different types of action spaces in reinforcement learning (RL), namely, optimizing parameters, optimizing rules, processing procedures, and machine equipment.
[0007] The present invention proposes an explicit exploration and asynchronous update proximal policy optimization algorithm (E2APPO) based on the Job shop scheduling problem (JSSP), with the optimization goal of minimizing the makespan. The main work of this paper is as follows: (1) By designing a dynamic optimization exploration strategy and an asynchronous update mechanism, an explicit exploration and asynchronous update proximal policy optimization algorithm (E2APPO) is constructed to obtain the mapping relationship between the production state and the action probability distribution to obtain the optimal operation sequence. (2) An adaptive scheduling scheme is constructed for different production states, especially different case scales. (3) A real-time scheduling system is established to achieve offline training and online execution; this system can allocate a trained model to cope with the uncertain workshop environment to improve the scheduling efficiency. (4) The numerical experimental results prove the effectiveness and generality of the proposed explicit exploration and asynchronous update proximal policy optimization algorithm (E2APPO). Summary of the Invention
[0008] The technical problem to be solved by the present invention is: to provide an adaptive scheduling method for a job shop based on deep reinforcement learning to solve the technical problems existing in the prior art.
[0009] The technical solution adopted by the present invention is: an adaptive scheduling method for a job shop based on deep reinforcement learning, the method comprising the following steps:
[0010] (1) Construct a scheduling function model for the job shop scheduling problem: Suppose there are n jobs and m machines, and each job includes m different operations. In job shop scheduling, n jobs J = {J1, J2..., J n} must be processed on m machines m = {M1, M2..., M m} in a different order known in advance. Let O k,b represent the k-th operation of workpiece b, and each operation O k,b must be executed on a specific machine within a specific time period. Workpiece b is on machine M kThe processing time on it is denoted as t b,k is marked as t b,k which is pre-determined. The actual completion time of workpiece b on machine M k is denoted as C b,k and it is equal to A b,k +t b,k , where A b,k represents the start processing time of workpiece b on machine M k A workpiece is completely finished after the completion of its last process. All scheduling objectives depend on the completion times of all workpieces; the objective function of minimizing the makespan corresponds to the length of the schedule; the scheduling function model of the Job shop scheduling problem (JSSP) is defined as:
[0011] C max = min max{C b,k} (1)
[0012] where b = 1, 2... n; k = 1, 2..., m;
[0013] C bk -t bk + M(1 - y bhk ) ≥ C bh (2)
[0014] where M is a very large value, b = 1, 2... n; h, k = 1, 2..., m; C bk represents the actual completion time of workpiece b on machine M k ; t b,k represents the processing time of workpiece b on machine M k ; C bh represents the actual completion time of workpiece b on machine M h ; y bhk represents the conditional function as in (4). If workpiece b is processed on machine h before machine k, y bhk is equal to 1, otherwise it is equal to 0.
[0015] C ak - C bk + M(1 - x bak ) ≥ t ak (3)
[0016] where M is a very large value, a, b = 1, 2... n; k = 1, 2..., m; C ak represents the actual completion time of workpiece a on machine M k , C bk represents the actual completion time of workpiece b on machine M k ; ta,k Denotes the processing time of workpiece a on machine M k ; x bhk Denotes a conditional function such as (5). If workpiece b is processed on machine k before workpiece a, x bhk is equal to 1; otherwise it is equal to 0;
[0017]
[0018]
[0019] Equation (1) is the overall objective function to minimize the completion time of all workpieces; equations (2)-(3) are the constraints of the scheduling process; equation (2) indicates that workpiece b is processed on machine h before machine k, and equation (3) indicates that workpiece b is processed on machine k before workpiece a.
[0020] (2) After introducing the optimization strategy and asynchronous update mechanism into the proximal policy optimization algorithm, the direct and efficient exploration and asynchronous update proximal policy optimization algorithm is formed;
[0021] (3) Combine the graph neural network with the hierarchical non-linear refinement of the original state information, and based on the direct and efficient exploration and asynchronous update proximal policy optimization algorithm in step (2), a kind of end-to-end deep reinforcement learning method is given;
[0022] (4) Based on the end-to-end deep reinforcement learning method in step (3), an adaptive scheduling decision is made for the job shop in step (1).
[0023] The action policy adopts a new exploration strategy
[0024]
[0025] In step (2.4), the following loss function is adopted
[0026]
[0027] Among them,
[0028] Among them, x i , y i respectively represent the target value and the predicted value. In the region where the error is close to 0, the average value of the square of the difference between the target value and the predicted value is used, and in the region where the error is far from 0, the average value of the absolute value of the difference between the target value and the predicted value is used.
[0029] Both network A and network C adopt the activation function
[0030] f(x) = x.sigmoid(βx) (10)
[0031] Among them, x is the input of the network, f(x) is the output after the non-linear transformation of the network, and β is the trainable parameter.
[0032] The A network and the C network are updated using an asynchronous update mechanism: K = 2 means that the A network is updated once after the C network is updated twice.
[0033] Advantages of the present invention: Compared with the prior art, the effects of the present invention are as follows:
[0034] 1) For the job shop scheduling problem, based on the proximal policy optimization algorithm that combines the action optimization search strategy and the asynchronous update mechanism, the present invention proposes a direct and efficient exploration and asynchronous update proximal policy optimization algorithm; the direct and efficient exploration and asynchronous update proximal policy optimization algorithm of the present invention has high robustness, the scheduling score is increased by 5.6% compared with the proximal policy optimization algorithm, and the minimum completion time is reduced by 8.9% compared with the traditional deep Q-network algorithm. The experimental results prove the effectiveness and generality of the proposed adaptive scheduling strategy;
[0035] 2) The action strategy draws on the ε-greedy strategy in the value deterministic strategy, and selects the action with a high action probability as the best action, as shown in Equation (8). This method reduces meaningless searches and enhances the search direction and small-scale traversal. This strategy can learn the optimal scheduling strategy faster and is more suitable for the dynamic complexity, variability and uncertainty of the workshop;
[0036] 3) The advantage function is evaluated, and a delayed strategy is introduced to form an asynchronous update mechanism between the C network and the A network. The asynchronous update mechanism reduces the incorrect update of the A network because the update speed of the A network is slower than that of the critic network. Such an advantage can avoid unnecessary repeated updates and reduce the cumulative error of repeated updates. K is the update delay coefficient between the A and C networks;
[0037] 4) The smooth loss function is used instead of the mean square error loss function; this loss function is insensitive to outliers and ensures stability. In job shop scheduling, outliers will inevitably appear in the exploration of spatial values. The model generated by the smooth loss function is more suitable for complex manufacturing, has better robustness, and can adapt to different scheduling situations. To maximize the model performance, the activation function adopted by the neural network can be regarded as a smooth function between the linear function and the Relu function, combining the advantages of both. This activation function has better performance than the Relu activation function. Description of the Drawings
[0038] Figure 1 It is a schematic diagram of a Markov chain for production scheduling;
[0039] Figure 2 It is a flowchart of the algorithm based on PPO2;
[0040] Figure 3 It is a diagram of a real-time scheduling system based on E2APPO;
[0041] Figure 4 It is a convergence comparison diagram of the ε-greedy strategy and the softmax strategy;
[0042] Figure 5 It is a convergence comparison diagram of different ε parameters;
[0043] Figure 6 It is a convergence comparison diagram under different k coefficients;
[0044] Figure 7 It is a comparison diagram of E2APPO and the GA algorithm;
[0045] Figure 8 It is a performance score diagram of E2APPO and the GA;
[0046] Figure 9 It is a generalization test diagram of E2APPO for large-scale algorithms;
[0047] Figure 10 It is a scheduling score diagram of E2APPO and basic PPO;
[0048] Figure 11 It is a comparison diagram of E2APPO and MDQN in terms of training stability. Specific implementation manners
[0049] The present invention will be further introduced below in conjunction with specific embodiments.
[0050] Embodiment 1: As Figure 1 - 11 shown, a job shop adaptive scheduling method based on deep reinforcement learning, the method includes the following steps:
[0051] The correct processing sequence and process scheduling are crucial for maximizing the productivity of the workshop. The job shop scheduling problem can be regarded as a sequential decision-making problem. The goal of scheduling is to determine the processing sequence of each process on each machine and the start time of each process to minimize the makespan.
[0052] For the convenience of modeling, several predefined constraints are agreed for this problem. These constraints are the same as the methods in the prior art and are as follows: (1) The sequential relationship and processing time of different processes of the same workpiece are known in advance; (2) Each machine can perform at most one operation at a time; (3) Each operation can only be performed on one machine; (4) Any started processing should be continuous without interruption until completion; (5) There is no sequential constraint between the processes of different workpieces; (6) All workpieces arrive available at time 0.
[0053] (1) Construct a scheduling function model for the job shop scheduling problem: Suppose there are n jobs and m machines. Each job consists of m different processes. In job shop scheduling, the n jobs J = {J1, J2..., J n} must be processed on the m machines m = {M1, M2..., M m} in a pre-known different order. Let O k,b represent the k-th process of workpiece b. Each process O k,b must be executed on a specific machine within a specific time period. The processing time of workpiece b on machine M k is marked with t b,k , and t b,k is pre-determined. The actual completion time of workpiece b on machine M k is represented by C b,k , which is equal to A b,k + t b,k , where A b,k represents the start processing time of workpiece b on machine M k . A workpiece is completely finished after the completion of its last process. All scheduling objectives depend on the completion times of all workpieces; the objective function of minimizing the makespan corresponds to the length of the schedule; the scheduling function model of the job shop scheduling problem (Job shop scheduling problem, JSSP) is defined as:
[0054] C max = min max{C b,k} (1)
[0055] where b = 1, 2... n; k = 1, 2..., m;
[0056] C bk - t bk + M(1 - y bhk ) + C bh (2)
[0057] where M is a very large value, b = 1, 2... n; h, k = 1, 2..., m; C bk represents the actual completion time of workpiece b on machine M k ; t b,k represents the processing time of workpiece b on machine M k ; C bh represents the actual completion time of workpiece b on machine M h ; y bhk represents a conditional function as in (4). If workpiece b is processed on machine h before machine k, y bhk is equal to 1, otherwise it is equal to 0.
[0058] C ak -C bk +M(1 - x bak )≥t ak (3)
[0059] Among them, M is a maximum value, a, b = 1, 2... n; k = 1, 2..., m; C ak represents the actual completion time of workpiece a on machine M k ; C bk represents the actual completion time of workpiece b on machine M k ; t a,k represents the processing time of workpiece a on machine M k ; x bhk represents a conditional function such as (5). If workpiece b is processed on machine k before workpiece a, x bhk is equal to 1, otherwise it is equal to 0.
[0060]
[0061]
[0062] Equation (1) is the total objective function to minimize the completion time of all workpieces; Formulas (2)-(3) are the constraint conditions of the scheduling process; Formula (2) means that workpiece b is processed on machine h before machine k, and formula (3) means that workpiece b is processed on machine k before workpiece a. In this case, the present invention is to find the best strategy to solve the scheduling problem;
[0063] The algorithm adopted is to improve the traditional proximal policy optimization algorithm (PPO) for job shop scheduling. Combining with graph neural network forms an end-to-end reinforcement learning method, which can effectively extract the job shop state features and help the agent learn more accurate strategies.
[0064] The proximal policy optimization algorithm is based on the typical AC network framework, where the A network is used for action selection and the C network is used to evaluate the state value function V(s t ) to evaluate the decisions made by the actor. The proximal policy optimization algorithm restricts the update range of the new and old policies to ensure its stability, making the Policy Gradient (PG) algorithm less sensitive to a larger learning rate. It adopts a clip function to limit the update degree between 1 - ∈ and 1 + ∈, as shown in Equation (6), where ε is a hyperparameter.
[0065]
[0066] A(st , a t ) = ∑ t′>t γ t′-t r t′ -V(s t ) (7)
[0067] The advantage function formula (7) is defined as the integral of the state value function V(s t ) and the discounted reward, representing the additional benefit of taking action a t . The state value function V(s t ) is negative, so the variance is smaller; the network is trained by applying an optimizer (Adam).
[0068] The present invention uses an agent to interact with the production workshop to generate scheduling data, such as processing time, machine allocation, scheduling the current process, etc. These data are collected and stored in a buffer. After one trajectory, the actor network and the critic network use the stored scheduling data to learn experience. The critic network is updated by gradient descent using the temporal-difference (TD) error, and the actor network is updated by gradient ascent using the policy gradient to find the optimal actor network for coping with changes in the production state. The specific process of scheduling is as Figure 2 shown.
[0069] (2) The direct and efficient exploration and asynchronous update proximal policy optimization algorithm proposes the Markov decision process transformation of the job shop scheduling environment, such as using a graph neural network to extract job shop features, the action space composed of optional operations, the reward design in the model training process, etc. Based on the consistent performance of the proximal policy optimization algorithm in the discrete action space, after introducing the greedy policy and the asynchronous update method into the proximal policy optimization algorithm, the direct and efficient exploration and asynchronous update proximal policy optimization algorithm is formed, and this algorithm adaptively schedules the job shop in step (1);
[0070] The steps of the direct and efficient exploration and asynchronous update proximal policy optimization algorithm are as follows:
[0071] (2.1) Input: Network π of A with training parameter θ θ ; Network v of C with training parameter ω ω , clipping coefficient ∈, update frequency multiple K of network C relative to network A, discount factor λ, greedy factor ε;
[0072] (2.2) Model the Markov process of the production environment, design the environmental state (s t ), action set (a t ), reward value (r t );
[0073] (2.3) Perform 1-N rounds of scheduling training; for 1-J steps in this round of training; sense the state s t , select the action a based on the action policy t ; obtain the immediate reward r t and the next state s t+1 ; collect the above parameters {s t , r t , a t} into the experience pool, and determine whether this round of scheduling is completed;
[0074] (2.4) After the scheduling is completed, evaluate the advantage function of this round of training by inputting the experience pool data into the C network
[0075] (2.5) Update the C network by backpropagation
[0076] (2.6) When the number of training times is an integer multiple of K, update the parameters θ of the A network according to the following formula, reflecting the asynchronous update of the AC network
[0077]
[0078] (2.7) Assign the updated parameters to the A network π old ← π θ .
[0079] The role of step 2.2 is to utilize the key elements {s t , r t , a t} used in the reinforcement learning process designed by Markov, which will be introduced in detail in the following section. N is the number of trajectories, and J is the number of training steps for each trajectory. In each trajectory, the content in step 2.3 "Perform 1-N rounds of scheduling training; for 1-J steps in this round of training; sense the state s t , select the action a based on the optimized action policy t ; obtain the immediate reward r t and the next state s t+1 " means that the agent interacts with the production environment and collects data. The action policy draws on the ε-greedy policy in Q-learning (Q-learning), selects the action with a high probability as the best action, as shown in Equation (8). This method reduces meaningless searches and enhances the search direction and small-scale traversal. This policy can learn the optimal scheduling policy faster and is more suitable for the dynamic complexity, variability, and uncertainty of the workshop. ε is the balance between exploration and exploitation, generally adjusted between 0.5 and 0.15. In the simulation experiment of the present invention, 0.1 is adopted.
[0080]
[0081] At the end of a trajectory, in step 2.4, the parameters collected by the agent interacting with the environment in the previous three steps are input into the C network to evaluate the advantage function And a delayed policy is introduced in step 2.6. When the number of updated steps is an integer multiple of K, the A network is updated, forming an asynchronous update mechanism between the A network and the C network. The asynchronous update mechanism reduces incorrect updates because the actor updates more slowly than the critic network. Such an advantage can avoid unnecessary repeated updates and reduce the cumulative error of repeated updates. K is the actor update delay coefficient, and its optimal value is 2 in the training experiment. Different from most algorithms, the present invention uses a smooth loss function as shown in Equation (9) instead of the mean squared error loss function. This loss function is insensitive to outliers and ensures stability. In job shop scheduling, outliers inevitably occur during the exploration of spatial values. The model generated by the smooth loss function is more suitable for complex manufacturing, has better robustness, and can adapt to different scheduling situations; the smooth loss function adopted:
[0082]
[0083] where,
[0084] where x i , y i represent the target value and the predicted value respectively. The average value of the square of the difference between the target value and the predicted value is used in the region where the error is close to 0, and the average value of the absolute value of the difference between the target value and the predicted value is used in the region where the error is far from 0.
[0085] To maximize the model performance, the neural network of the present invention adopts an activation function as shown in Equation (10), which can be regarded as a smooth function between a linear function and a Relu function, combining the advantages of both. This activation function has better performance than the Relu activation function. This activation function is used in the experiment, and the results show that this method has good accuracy. The activation functions of the A network and the C network:
[0086] f(x) = x.sigmoid(βx) (10)
[0087] where, β is a trainable parameter
[0088] where x is the input of the network, f(x) is the output after the nonlinear change of the network, and β is a trainable parameter.
[0089] (3) After introducing the optimization strategy and the asynchronous update mechanism into the proximal policy optimization algorithm, a direct and efficient exploration and asynchronous update proximal policy optimization algorithm is formed;
[0090] (4) Combine the graph neural network with the hierarchical non - linear refinement of the original state information, and directly and efficiently explore and asynchronously update the proximal policy optimization algorithm based on step (2), to give an end - to - end deep reinforcement learning method;
[0091] (5) Make an adaptive scheduling decision for the job shop in step (1) based on the end - to - end deep reinforcement learning method in step (4).
[0092] The Markov process modeling of the job shop is as follows:
[0093] Reinforcement learning applies an agent to continuously interact with the environment. The agent obtains the mapping between states and actions through interaction with the environment, and learns the optimal policy to maximize the cumulative reward. The basic reinforcement learning task is usually transformed into a Markov decision process (MDP). The Markov decision process (MDP) framework describes the environment with a 5 - tuple <S, a, P, r(S, a), γ>. S represents the set of environmental states, A represents the set of actions that the agent can execute, P is the probability of state transition, representing the probability of transitioning from the previous state to the current state. The reward r(s t , a t ) represents the reward for taking action a t ∈ S under state S t ∈ A. The most important property of Markov is that the next state is independent of past states and only related to the current state.
[0094] Job shop scheduling is very suitable for transformation into a Markov decision process. The agent observes the job shop scheduling state, selects actions, immediately obtains rewards after the operation is completed, and then maximizes the cumulative reward to learn the optimal scheduling policy. The Markov model of job shop scheduling has the following key elements.
[0095] (1) Feature extraction of job shop state based on the Graph Neural Networks (GNN) method
[0096] The job shop scheduling state can be represented by a selection graph. The selection graph provides a comprehensive view, including the processing time on each machine and the pre - constraint sequence. The decision - making point in job shop scheduling is represented as a disjunctive graph G=(N, A, E). The nodes N describe the set of all operations of all workpieces, including start and end virtual nodes, N = O ∪ {O s , O e} = {O s , O 1,1 ,..., O 1,v1..., O n,1 ...O n,vn , O e}; The connection arc set A represents the set of all processes of the same workpiece. For each node, A contains the directed edge O(j,k) → O(j,k+1); the disjunctive set E reflects the undirected arcs, and each arc connects a pair of processes that require the same machine for processing. Therefore, finding a solution for the job scheduling instance is the same as determining the direction of each separation point, thus generating a directed acyclic graph (DAG). Minimizing the longest path in the disjunctive graph is exactly the optimal solution for minimizing the makespan.
[0097] The method based on Graph Neural Networks (GNN) is an effective method for extracting the features of the disjunctive graph and updating the disjunctive graph as input. The method based on the spatial domain obtains features by neighborhood sampling, calculating the correlation between the target node and neighborhood nodes, and aggregating the received messages into a single vector to represent the shop floor state. Taking G=(N,A,E) as an example, the graph neural network (Graph Neural Networks, GNN) is used to iterate each node to obtain multi-dimensional embeddings, and the update equation for the k-th iteration is described by formula (11). A single heuristic rule only uses a single attribute as the basis for the scheduling sequence. It only considers local information and may produce different scheduling performances in different situations. In contrast, the features extracted by the graph neural network (Graph Neural Networks, GNN) method are based on the original data, can better represent the current state, and avoid the deficiencies of artificial features.
[0098]
[0099] In the formula, σ is non-linear, W is the weight matrix, h is the node feature, k is the depth, and N is the neighborhood function.
[0100] (2) Modeling the action space of the agent in the job shop
[0101] A represents the set of actions that can be selected at each scheduling point. In the field of job shop scheduling, the action space generally refers to the operations or heuristic rules that can be executed. In addition, there are some different forms, such as equipment settings and parameter selections. In the present invention, the process is designed as the action space. Select O t ∈A t as the action at decision step t. Assuming that each workpiece can only have one processable process at time t, then the size of the action set is equal to the number of workpieces and decreases as the workpieces are completed.
[0102] (3) Modeling the reward for the agent to execute actions in the job shop
[0103] The reward function is essentially designed to guide the agent to obtain the maximum cumulative reward. Our agent's goal is to minimize the makespan C under the optimal scheduling strategy max 。C max is the maximum completion time of all jobs and is the same as the scope of the entire schedule. The reward function is defined as formula (12), where r(a t , s t ) represents the reward value obtained by the agent after executing action a t , and is also the value difference between state s t and state s (t+1) . Maximizing the accumulation of immediate rewards is consistent with the effect of minimizing the completion time. Reward design is the key to the success of production scheduling, and this invention takes the completion time as the most critical factor in production scheduling.
[0104] r(a t , s t ) = T(s t ) - T(s t+1 ) (12)
[0105] where T(s t ) represents the completion time in state s t , and T(s t+1 ) represents the completion time of the next state.
[0106] Instance simulation: A real-time scheduling system was established to verify the performance of the algorithm, and algorithm tests and comparisons were carried out under the system. First, a real-time scheduling system with a deep reinforcement learning algorithm model was established to enhance the immediate scheduling ability of the production workshop. The parameter optimization and setting of the training and testing process will be introduced later. Then, the performance of the proposed direct and efficient exploration and asynchronous update proximal policy optimization algorithm was compared with that of classical heuristic algorithms and other reliable scheduling rules. To further verify the advantages of the proposed adaptive scheduling strategy, the direct and efficient exploration and asynchronous update proximal policy optimization algorithm was also compared with two other methods trained using reinforcement learning. The results of the comparative experiments verified the effectiveness and generality of the proposed adaptive scheduling strategy.
[0107] Job shop real-time scheduling system based on this method: Real-time is a significant difference between the workshop production scheduling system based on deep reinforcement learning and traditional scheduling algorithms. Our goal is not only to develop an advanced solution applicable to small instances, but also to find a solution that can quickly obtain an approximate solution in the best case for large-scale scenarios. The system proposed in this invention is as Figure 3As shown. On the one hand, the system can use historical data or simulation data to describe the state of the job shop, and perform offline training on the model in advance, and then store the trained model for future use. On the other hand, the system can evaluate the current state of the job shop through real-time sensing technology in the shop or Internet of Things technology, and then select a well-trained model for real-time scheduling. At the same time, the trained model has strong generalization ability for scheduling instances of different sizes, avoiding the time consumption of retraining, and has real-time scheduling performance compared with traditional methods.
[0108] Experimental parameters: The training process is carried out under the above scheduling system; the processing time of the operations and the machine task assignment of training instances of various sizes are randomly generated within the range of 1-99. Experiments show that convergence can be achieved after 10,000 training trajectories. The proposed direct efficient exploration and asynchronous update proximal policy optimization algorithm runs on a computer with an Intel Core i7-6700@4.0GHz CPU, a GEFORCE RTX 2080Ti GPU, and 8GB of RAM. Table 1 shows the parameters of the training process. New instances are randomly generated in each round of training, improving the generality of the direct efficient exploration and asynchronous update proximal policy optimization algorithm in the training process, similar to a complex manufacturing environment. After each training stage, the trained direct efficient exploration and asynchronous update proximal policy optimization algorithm is tested on a validation instance to evaluate the effectiveness of the trained model.
[0109] Table 1 Parameter settings of the algorithm in training
[0110] Parameter Name Value Number of Training Times 10000 Memory Pool Capacity 1e6 Clipping Coefficient ∈ 0.2 Innovation Exploration Strategy Parameter ε 0.05-0.15 Learning Rate lr 2e-5 Delay Coefficient K 2 Discount Factor γ 1 GAE Parameter λ 0.98 Optimizer Adam
[0111] The innovative exploration strategy combines the advantages of the random strategy and the deterministic strategy. Compared with the deterministic strategy, the innovative exploration strategy can avoid falling into local optima; on the other hand, compared with the random strategy, the innovative exploration strategy has a more precise exploration direction, preventing meaningless exploration and consumption. Figure 4 Shows the convergence of the innovative exploration strategy and the softmax strategy during the training process. The reward curve of the innovative exploration strategy is basically higher than that of other strategies, indicating that the cumulative reward value of the innovative exploration strategy is greater than that of softmax. The performance of the innovative exploration strategy is better than that of the Softmax strategy in searching for the action space during the process.
[0112] The parameter ε is the balance between space exploration and exploitation, as Figure 5As shown. The ε-greedy parameter ε is the exploration probability, which is optimized within the range of 0.05 - 0.15, and ε = 1 is for random action. The experimental results show that except for ε = 1, within the range of ε = 1, the reward curve has a gradually increasing trend. After about 3000 episodes, the ε = 0.1 curve has reached the top, while in the later chapter, the reward value of ε = 0.15 will decrease. The reason may be that the increase of ε leads to insufficient exploitation. Through the comparison of the training process, the optimal value of ε is obtained as 0.1.
[0113] In the delayed update mechanism, the parameter K represents the delayed update frequency of the actor network relative to the critic network. The best value of K is selected as a multiple from 1 - 3. To better show the convergence under different coefficients K, the number of training times in this experiment is expanded to 16,000 times. As Figure 6 shown, the convergence curves of K = 1 and K = 2 are always higher. K = 1 is at a higher level at the beginning of the training stage, but in the later stage, due to the frequent update of the actor under the uncertainty of the critic, it is below the K = 2 curve. It can be concluded that the asynchronous update strategy with coefficient K = 2 stabilizes the whole training compared with K = 1 and converges to the highest point in the later stage of training.
[0114] Performance metrics and test datasets: For the present invention, the goal is to find a scheduling scheme to minimize the makespan. To comprehensively evaluate various scheduling methods, as shown in Equation (13), the performance score represents the gap between the minimum makespan obtained by different methods and the optimal solution (OR-Tools). The higher the performance score, the more effective the method.
[0115] Performance score = (1 - (T i - T best ) / T best ) * 100% (13)
[0116] where T i is the completion time of different methods, and T bestis the completion time of the OR-Tools solution. Two benchmark datasets used in this invention are the well-known public Job shop scheduling problem (JSSP) datasets and the generated instances; nearly 90 cases are selected from the public benchmarks. Among them, small and medium-sized examples are from FT, LA, and ORB. Large-scale examples are selected from the DMU dataset for comparison with the literature "C.-C. Lin, D.-J. Deng, Y.-L. Chih, and H.-T. Chiu (2019) Smart Manufacturing Scheduling With Edge Computing Using Multiclass Deep Q Network. IEEE Trans. Ind. Informatics 15(7): 4276–4284". The same instances generated in the literature "C. Zhang, W. Song, Z. Cao, J. Zhang, P. S. Tan, and C. Xu (2020) Learning to Dispatch for Job Shop Scheduling via Deep Reinforcement Learning. NeurIPS 1: 1–17" are adopted for comparison with the algorithms therein.
[0117] Results and Discussion:
[0118] Comparison with Heuristic Algorithms: To demonstrate the superiority of the direct and efficient exploration and asynchronous update proximal policy optimization algorithm proposed in this invention over heuristic algorithms, it was compared with the genetic algorithm (GA) in the literature "Y. Zhan and C. Qiu (2008) Genetic algorithm application to the hybrid flow shop scheduling problem. Proc. IEEE Int. Conf. Mechatronics Autom. ICMA2008". Several commonly used high-performance priority rules were selected for comparison in the literature "V. Sels, N. Gheysen, and M. Vanhoucke (2012) A comparison of priority rules for the job shop scheduling problem under different flow time-and tardiness-related objective functions. Int. J. Prod. Res. 50(5): 4255–4270". The genetic algorithm has good performance in solving the JSSP problem; the disadvantage is that it needs to be solved when encountering different job shop scheduling problem (JSSP) instances and spends a lot of time again.
[0119] For the scale of 15*15, 25 examples were selected for comparison with the genetic algorithm (GA). As Figure 7 shown, the method of this invention is superior to the genetic algorithm in 15 cases, equal to the genetic algorithm in 5 cases, and slightly lower than the genetic algorithm in the remaining 5 cases. From the above results and combined with Figure 8 it can be seen that the direct and efficient exploration and asynchronous update proximal policy optimization algorithm does not have an absolute advantage in quality compared with the genetic algorithm (GA). The main advantage of the direct and efficient exploration and asynchronous update proximal policy optimization algorithm is that it can obtain approximately excellent results in different sizes without retraining and has obvious advantages in large-size instances.
[0120] The comparison priority rules are as follows.
[0121] Shortest Processing Time (SPT): Select the next operation with the shortest processing time;
[0122] First In First Out (FIFO) rule: Select the next operation of the job that arrived earliest.
[0123] Longest Processing Time (LPT): Select the operation with the next longest processing time
[0124] MOPR (Most Operation Remaining): The job with the most remaining operations is processed first.
[0125] Most Work Remaining (MWKR): The highest priority is given to the operation belonging to this work, and the total processing time required to complete this operation.
[0126] Flow Deadline to Most Work Remaining Ratio (FDD): The task with an earlier deadline has a higher priority.
[0127] Table 2 Solution of Priority Rules and E2APPO in Different Instances
[0128]
[0129] The comparison between the scheduling rules and the Direct High-Efficiency Exploration and Asynchronous Update Proximal Policy Optimization algorithm is shown in Table 2. Among the 25 test instances, the Direct High-Efficiency Exploration and Asynchronous Update Proximal Policy Optimization algorithm is superior to the regular scheduling solution in 18 cases, with an overrate of 72%, indicating that the Direct High-Efficiency Exploration and Asynchronous Update Proximal Policy Optimization algorithm is superior to the regular scheduling. To prove the advantage of the Direct High-Efficiency Exploration and Asynchronous Update Proximal Policy Optimization algorithm in terms of generalization ability, 70 large-scale instances are selected from the benchmark, and the generalization test is carried out on the well-trained 30*20 scale model, and the average value of the results is compared with the known rules. As Figure 9 shown, the curve of the Direct High-Efficiency Exploration and Asynchronous Update Proximal Policy Optimization algorithm is always at the lower left corner. Compared with the known rules, the 30*20 model can also quickly solve the optimal value of a similar scale. The Direct High-Efficiency Exploration and Asynchronous Update Proximal Policy Optimization algorithm has strong generalization ability and adaptive performance, and is more suitable for complex and uncertain production environments.
[0130] Comparison with Existing Reinforcement Learning (RL) Scheduling Algorithms: To further confirm the advantage of the Direct High-Efficiency Exploration and Asynchronous Update Proximal Policy Optimization algorithm (E2APPO) over traditional reinforcement learning algorithms, the basic Proximal Policy Optimization algorithm (PPO) and the Deep Q-Network algorithm (DQN) are selected for comparison. First, it can be observed that the scheduling algorithm proposed in the present invention can further improve the performance of the basic Proximal Policy Optimization algorithm (PPO) and obtain higher scheduling scores in most cases, such asFigure 10 As shown. Especially for the 30*20 example, the scheduling score increased by 5.6%, demonstrating the effectiveness of asynchronous updates and strategies. At the same time, Table 3 presents the test results of several well-known rules, the modified deep Q-network algorithm (MDQN), and the E2APPO algorithm on the DMU dataset. The best values are shown in bold; compared with the modified deep Q-network algorithm (MDQN), the completion time for all examples was significantly reduced, with an average reduction of 8.9%. The results for each example and their average demonstrate the superiority of the direct and efficient exploration and asynchronous update proximal policy optimization algorithm. From Figure 11 it can be seen that the direct and efficient exploration and asynchronous update proximal policy optimization algorithm has a uniform training distribution and has obvious advantages in terms of considering the stability of individual instance results.
[0131] Table 3 Comparison between MDQN and E2APPO on DMU examples
[0132]
[0133]
[0134] Simulation conclusion: For the job shop scheduling problem, a direct and efficient exploration and asynchronous update proximal policy optimization algorithm is proposed. This algorithm adopts a dynamic greedy search strategy and an asynchronous update mechanism to minimize the total completion time. The proposed search strategy improves the search efficiency and avoids unnecessary searches, and the asynchronous update mechanism makes the update of the actor network more stable. The actor network adaptively selects the current operation according to the environmental state. Based on the proposed direct and efficient exploration and asynchronous update proximal policy optimization algorithm, an adaptive scheduling strategy is proposed in the real-time scheduling system, including offline training and online implementation. The adaptive scheduling strategy improves the adaptability to complex workshop environments. The results show that the well-trained model based on the direct and efficient exploration and asynchronous update proximal policy optimization algorithm has better generalization performance than heuristic algorithms at different scales and can achieve an optimal balance between scheduling quality and scheduling speed.
[0135] Through numerical experiments on a large number of examples, including well-known benchmarks and randomly generated examples as a realistic reproduction of actual manufacturing, the advantages of the proposed direct and efficient exploration and asynchronous update proximal policy optimization algorithm are verified. By comparison with heuristic algorithms, the superiority of the direct and efficient exploration and asynchronous update proximal policy optimization algorithm is verified, especially its generalization performance at different scales. Compared with existing reinforcement learning algorithms, the direct and efficient exploration and asynchronous update proximal policy optimization algorithm achieves our goal.
[0136] In summary, in the modern, diverse, and complex manufacturing industry, due to the limitations of response time, traditional scheduling methods can no longer meet the requirements of high efficiency. Therefore, an optimized action policy and asynchronous update mechanism are designed in the proximal policy optimization algorithm (PPO) to form the explicit exploration and asynchronous update proximal policy optimization algorithm (E2APPO), which combines the advantages of more explicit exploration directions and more stable training processes. Based on the explicit exploration and asynchronous update proximal policy optimization algorithm (E2APPO), a graph neural network is combined with the hierarchical non-linear refinement of the original state information to design an end-to-end reinforcement learning method. On this basis, we implemented an adaptive scheduling system, which consists of two subsystems: one is an offline system that pre-trains and stores the trained model; the other is an online system that calls the model in real time. Under this system, extensive tests were conducted on the trained explicit exploration and asynchronous update proximal policy optimization algorithm (E2APPO), and comparisons were made with heuristic algorithms such as genetic algorithms, priority scheduling rules, and other existing reinforcement learning-based scheduling methods. Compared with the genetic algorithm, 75% of the cases obtained solutions that were better than or equivalent to the genetic algorithm. In the generalization test, all large instances outperformed the known scheduling rules, demonstrating the advanced robustness of the explicit exploration and asynchronous update proximal policy optimization algorithm (E2APPO). The scheduling score increased by 5.6% compared to the proximal policy optimization algorithm (PPO), and the minimum completion time decreased by 8.9% compared to the deep Q-network algorithm (DQN). The experimental results prove the effectiveness and generality of the proposed adaptive scheduling strategy.
[0137] The present invention has the following advantages: (1) By designing a dynamic optimization exploration strategy and an asynchronous update mechanism, an explicit exploration and asynchronous update proximal policy optimization algorithm (E2APPO) is developed to obtain an optimal operation sequence that maps states and action probability distributions. (2) An adaptive scheduling scheme is constructed for different instance states, especially different instance scales. (3) A real-time scheduling system is established to achieve offline training and online execution; this system can allocate trained models to handle unforeseen workshop environments to improve scheduling efficiency. (4) The numerical experimental results prove the effectiveness and generality of the proposed explicit exploration and asynchronous update proximal policy optimization algorithm (E2APPO).
[0138] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. An adaptive scheduling method for job shops based on deep reinforcement learning, characterized in that: The method comprises the following steps: (1)Construct a scheduling function model for the job shop scheduling problem: There are n jobs and m machines. Each job consists of m different operations. In job shop scheduling, n jobs J = {J1, J 2…… , J n must be processed on m machines m = {M1, M2 ……, M m} in a pre-known different order. Let O k,b denote the k-th operation of workpiece b. Each operation O k,b must be executed on a specific machine within a specific time period. The processing time of workpiece b on machine M k is marked with t b,k . t b,k is pre-determined. The actual completion time of workpiece b on machine M k is denoted by C bk , which is equal to A b,k + t b,k , where A b,k represents the start processing time of workpiece b on machine M k . A workpiece is completely finished after the completion of its last operation. All scheduling objectives depend on the completion times of all workpieces; the objective function of minimizing the makespan corresponds to the length of the schedule; the scheduling function model of the job shop scheduling problem is defined as: C max = minmax{C b,k} (1) where b = 1, 2 …… n; k = 1, 2 ……, m; C bk -t bk +M(1 - y bhk )≥C bh (2) where M is a maximum value, b = 1, 2,..., n; h, k = 1, 2,..., m; C bk represents the actual completion time of workpiece b on M k machine; t b,k represents the processing time of workpiece b on machine M k ; C bh represents the actual completion time of workpiece b on M h machine; y bhk represents the conditional function as in (4), if workpiece b is processed on machine h before machine k, y bhk equals 1, otherwise equals 0 C ak -C bk +M(1 - x bak ) ≥ t ak (3) Among them, M is a maximum value, a, b = 1, 2... n; k = 1, 2..., m; C ak represents the actual completion time of workpiece a on M k machine, C bk represents the actual completion time of workpiece b on M k machine; t a,k represents the processing time of workpiece a on machine M k ; x bhk represents a conditional function such as (5). If workpiece b is processed on machine k before workpiece a, x bhk is equal to 1, otherwise it is equal to 0 Equation (1) is the total objective function for minimizing the completion time of all workpieces; Equations (2)-(3) are the constraints in the scheduling process; Equation (2) indicates that workpiece b is processed before machine k on machine h, and Equation (3) indicates that the processing of workpiece b on machine k is prior to that of workpiece a; (2) After introducing an optimization strategy and an asynchronous update mechanism into the proximal policy optimization algorithm, a direct and efficient exploration and asynchronous update proximal policy optimization algorithm is formed; (3) Combine the graph neural network with the hierarchical non-linear refinement of the original state information, and based on the direct and efficient exploration and asynchronous update proximal policy optimization algorithm in step (2), present an end-to-end deep reinforcement learning method; (4) Based on the end-to-end deep reinforcement learning method in step (3), make an adaptive scheduling decision for the job shop in step (1); The steps of the direct and efficient exploration and asynchronous update proximal policy optimization algorithm are as follows: (2.1) Input: Network A π with training parameter θ θ ; Network C v with training parameter ω ω , clipping coefficient ∈, update frequency multiple K of Network C relative to Network A, discount factor λ, greedy factor ε; (2.2) Markov process modeling in the production environment, design of environmental states (s t ), action set (a t ), reward value (r t ); (2.3) Perform 1-N rounds of scheduling training; for the 1-J steps in this round of training; perceive the state s t , select the action a based on the action policy t ; obtain the immediate reward r t and the next state s t+1 ; collect the above parameters {s t, r t , a t} into the experience pool, and determine whether this round of scheduling is completed; After the scheduling in (2.4) is completed, the advantage function of this round of training is evaluated by inputting the experience pool data into the C network. (2.5) Backpropagation updates the C network (2.6) When the number of training times is an integer multiple of K, update the parameters θ of the A network according to the following formula, (2.7) Assign the updated parameters to the A network π old ← π θ .
2. The adaptive scheduling method for a job shop based on deep reinforcement learning according to claim 1, characterized in that: The action policy adopts a new exploration strategy 3. The adaptive scheduling method for a job shop based on deep reinforcement learning according to claim 1, wherein: Adopt the following loss function Among them, where x i , y i represent the target value and the predicted value respectively. The average value of the square of the difference between the target value and the predicted value is used in the region where the error is close to 0, and the average value of the absolute value of the difference between the target value and the predicted value is used in the region where the error is far from 0.
4. The adaptive scheduling method for a job shop based on deep reinforcement learning according to claim 1, characterized in that: Both the A network and the C network adopt activation functions f(x) = x.sigmoid(βx) (10) where x is the input of the network, f(x) is the output after the non-linear change of the network, and β is a trainable parameter.
5. The adaptive scheduling method for a job shop based on deep reinforcement learning according to claim 1, wherein: The A network and the C network are updated using an asynchronous update mechanism: K = 2 means that the A network is updated once after the C network is updated 2 times.
Citation Information
Patent Citations
Three-dimensional group exploration method based on multi-head attention asynchronous reinforcement learning
CN113283169A
Concept for Placing an Execution of a Computer Program
US20220107793A1
Cited By
Method for solving job shop scheduling problem based on generative adversarial imitation learning
CN116796964A