Distributed Blocking Flow Shop Scheduling Method and System Based on Deep Reinforcement Learning
By using deep reinforcement learning methods in the distributed blocking flow workshop, the agent's Actor network is trained, and the problems of low dynamic online scheduling efficiency and poor compatibility are solved, and efficient workpiece scheduling decisions and dynamic environment adaptation are achieved.
Patent Information
- Application Number
- CN202210864726.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-21
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-07-21
AI Technical Summary
The dynamic online scheduling method of the existing distributed blocking flow workshop is inefficient and has poor compatibility, making it difficult to adapt to changes in the dynamic environment.
Using a deep reinforcement learning method, the workshop is regarded as an agent, and four deep reinforcement learning networks, Actor, Critic, targetActor and targetCritic, are used to obtain the optimal network parameters through training, so that the Actor can make the optimal decisions and reduce the deviation of the total completion time of the artifact.
It realizes efficient online decision-making, can make local modifications based on the original scheduling plan, is suitable for various processing scenarios, has strong compatibility, and is adapted to dynamic production environments.
Smart Images

Figure CN115330028B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of distributed blocking pipeline scheduling, and more specifically, relates to a distributed blocking flow shop scheduling method and system based on deep reinforcement learning. Background Art
[0002] As a typical production scheduling problem, the flow shop scheduling problem (FSP) plays an important role in modern manufacturing systems. In the classical pipeline scheduling problem, there is an infinite buffer capacity between consecutive machines. However, in many real-world manufacturing environments, such as pharmaceutical workshops, due to technical requirements or work characteristics, the buffer capacity between machines is limited or even zero. In this case, the FSP becomes a blocking flow shop scheduling problem. At the same time, compared with the traditional single-factory production and processing mode, the distributed manufacturing system makes full use of the resources of multiple factories and realizes overall optimization through the effective allocation of resources, quickly achieving the lowest-cost manufacturing. Therefore, the distributed blocking flow shop scheduling problem (DBFSP) has become one of the active research topics in the field of manufacturing scheduling.
[0003] Typical scheduling methods can be divided into two categories: exact scheduling methods and approximate scheduling methods. Due to high computational complexity, exact scheduling methods are not efficient in solving large-scale scheduling problems. Compared with exact scheduling methods, approximate scheduling methods do not search the entire solution space but explore multiple directions based on specific strategies. Therefore, approximate scheduling methods have a lower computational complexity, can obtain feasible solutions more quickly, and have more advantages in solving large-scale scheduling problems. However, most approximate methods applied to dynamic rescheduling must consider all information and formulate a completely new schedule instead of adjusting the original plan. Therefore, they are not flexible enough, lack adaptability, and have poor compatibility when facing a changing environment and different situations.
[0004] Therefore, the problems of low efficiency and poor compatibility of the existing dynamic online scheduling methods for distributed blocking flow shops have become technical challenges in this field. Summary of the Invention
[0005] In view of the above defects or improvement requirements of the prior art, the present invention provides a distributed blocking flow shop scheduling method and system based on deep reinforcement learning, thereby solving the technical problems of low efficiency and poor compatibility of the existing dynamic online scheduling methods for distributed blocking flow shops.
[0006] To achieve the above object, according to one aspect of the present invention, the following technical solutions are provided:
[0007] An intelligent agent training method for distributed blocking flow shop scheduling based on deep reinforcement learning. In this method, a workshop is regarded as an intelligent agent, and each intelligent agent includes four deep reinforcement learning networks: Actor, Critic, targetActor, and targetCritic. The training objective is to obtain the optimal network parameters of the Actor so that the Actor can make an optimal decision on whether the intelligent agent receives a new workpiece to be scheduled, minimizing the total completion time deviation of all workpieces within all intelligent agents. The completion time deviation is the absolute value of the completion time of the workpiece minus the due date. The training method for any intelligent agent A includes the following steps:
[0008] (T1) Randomly select a set of data (s, a, s′, r) from the pre-obtained training set. Here, s is the set of current observation states of each intelligent agent, a is the set of actions decided by the Actor of each intelligent agent under the current observation state, s′ is the set of new observation states of each intelligent agent after executing the corresponding actions in a, and r is the sum of the reward values obtained by each intelligent agent after executing the corresponding actions in a. The smaller the total completion time deviation of all workpieces of the intelligent agent after inserting the new workpiece to be scheduled, the greater the obtained reward value;
[0009] (T2) Evaluate the scheduling plan of each intelligent agent executing the corresponding actions in a under the current observation state set s through an evaluation function to obtain an evaluation value. The greater the evaluation value, the better the scheduling plan; μ The greater the evaluation value, the better the scheduling plan;
[0010] (T3) Use the gradient ascent method to update the Actor network parameters θ of intelligent agent A; μ Update;
[0011] (T4) Use the gradient descent method of the TD error of the deep Q network to update the Critic network parameters θ of intelligent agent A, making the Critic of intelligent agent A closer to its targetCritic; Q Update, making the Critic of intelligent agent A closer to its targetCritic;
[0012] (T5) Copy the updated Actor network parameters θ after step (T3) to the targetActor of intelligent agent A generation by generation, and copy the updated Critic network parameters θ after step (T4) to the targetCritic of intelligent agent A generation by generation; μ Copy the updated Actor network parameters θ after step (T3) to the targetActor of intelligent agent A generation by generation, and copy the updated Critic network parameters θ after step (T4) to the targetCritic of intelligent agent A generation by generation; Q generation by generation;
[0013] (T6) Return to execute step (T1) until the set training times are reached or the gradient converges in step (T3).
[0014] Preferably, in step (T1), s = (o1,..., o i ,..., on ), where \(n\) is the total number of agents, \(o\) i is the current observation state of the \(i\)-th agent, \(o\) i =(b i1 , b i2 ), \(b\) i1 is the total completion time deviation of all workpieces in the \(i\)-th agent under the current execution scheduling scheme, \(b\) i2 is the total completion time deviation of all workpieces in the \(i\)-th agent after inserting the new workpiece to be scheduled; \(a=(a_1,...,a\) i ,...,a\) n ), \(a\) i is the action of the Actor of the \(i\)-th agent for making a decision on the new workpiece to be scheduled; \(s'=(o_1',...,o\) i ',...,o\) n '), \(o\) i ' is the new observation state after the \(i\)-th agent executes the action \(a\) i , \(o\) i '=(b i1 ', b i2 '), \(b\) i1 ' is the total completion time deviation of all workpieces of the \(i\)-th agent after executing the action \(a\) i , \(b\) i2 ' is the total completion time deviation of all workpieces of the \(i\)-th agent after inserting the next new workpiece to be scheduled after executing the action \(a\) i .
[0015] Preferably, in step (T2), the evaluation function is as follows:
[0016]
[0017] In the formula, \(J(\theta\) μ ) is the evaluation value, \(\theta\) μ is the Actor network parameter of agent A, \(\theta\) Q is the Critic network parameter of agent A, \(Q\) is the function of the Critic network of agent A, \(Q(s,a|\theta\) Q )|s = s t ,a = \(\mu(s\) t ) means that when the current observation state set s = s t , the Actor network parameter of agent A is \(\theta\) μ , the Q function values obtained by each agent executing the corresponding actions in the action set a are, \(\mu\) is the set of functions of the Actor networks of all agents, \(\mu(s|\theta\) μ )|s = s t means that when the current observation state set s = s tAt this time, the Actor network parameters of Agent A are θ μ Under this condition, the Actors of each agent respectively execute the corresponding actions in the action set a; p is the set of all current observation states, and E is the expected value.
[0018] Preferably, in step (T3), the gradient calculation formula of the gradient ascent method is as follows:
[0019]
[0020] Preferably, in step (T4), the formula for the TD error is as follows:
[0021]
[0022] In the formula, a' = μ'(s'|θ μ' ) is the action set decided by the target Actor of each agent under the observation state set s', γ is the attenuation factor, θ Q' is the network parameter of the target Critic of Agent A, θ μ' is the network parameter of the target Actor of Agent A, is the set of functions of all agent target Actor networks, μ θi' is the function of the target Actor network of the i-th agent, and E is the expected value.
[0023] Preferably, the method for obtaining the training data (s, a, s′, r) is as follows:
[0024] (T11) Initialize the number of new workpieces and the queuing order in the workpiece pool, the number of processes and process times of each workpiece, the current processing state in each agent, and the network parameters of the four networks of each agent;
[0025] (T12) Each agent obtains the total completion time deviation of all workpieces under the current scheduling plan, and the total completion time deviation of all workpieces in the agent if the new workpiece to be scheduled is inserted, that is, obtains the current observation state of the agent, so as to obtain the current observation state set s;
[0026] (T13) The Actor of each agent inputs its current observation state and outputs the decision-making action. If multiple agents all decide to execute the insertion of the new workpiece, a random agent is selected from the multiple agents to execute the insertion action, so as to obtain the action set a;
[0027] (T14) Each agent executes the decision-making action, calculates the reward value of each agent's executed action using the reward value function, and obtains the sum r of the reward values of each agent's executed action; removes the newly inserted workpiece from the workpiece pool;
[0028] (T15) Determine whether there are still new workpieces to be scheduled in the workpiece pool. If so, return to step (T12) to start the next scheduling, and use the current observation state set calculated in step (T12) as the new observation state set s' after each agent executes an action in this scheduling. If not, return to step (T11) until the data (s, a, s', r) of the set group is obtained.
[0029] Preferably, in step (T11), the steps of initializing the current processing state in each agent are as follows:
[0030] (T111) Randomly assign all initial workpieces to be processed to N workshops so that each workshop generates an initial sequence of workpieces to be processed;
[0031] (T112) In the sequence of workpieces to be processed in each workshop, randomly select two workpieces and re-insert them at all possible positions in this workshop to obtain a set of sequences of workpieces to be processed, that is, the first sequence set. Calculate the total completion time deviation of all sequences in the first sequence set. The sequence with the smallest deviation is the optimal one, and replace the initial sequence of workpieces to be processed in this workshop; thus, the optimization of the workpiece sequences in all workshops is completed;
[0032] (T113) Select the workshop fmax with the largest completion time deviation and the workshop fmin with the smallest completion time deviation. Move each workpiece in the sequence of fmax to all insertable positions in fmin to obtain a set of sequences, that is, the second sequence set. Each sequence contains two sequences from the workshops fmax and fmin. Calculate the total completion time deviation of all sequences in the second sequence set. The sequence with the smallest total completion time deviation is the optimal one, and replace the sequences in fmax and fmin according to this optimal sequence; Exchange the positions of each workpiece in the sequence of fmax with each workpiece in fmin to obtain a set of sequences, that is, the third sequence set. Each sequence contains two sequences from the workshops fmax and fmin. Calculate the total completion time deviation of all sequences in the third sequence set. The sequence with the smallest total completion time deviation is the optimal one, and replace the sequences in fmax and fmin according to this optimal sequence;
[0033] (T114) If the total completion time deviation of the sequences in all workshops after optimization in step (T113) is smaller than that before optimization, return to step (T112); otherwise, the initialization of the current processing state in each agent is completed.
[0034] Preferably, in step (T14), the calculation formula for the reward value of the agent is as follows:
[0035]
[0036] where, rewi is the reward value of the i-th agent, and is the total completion time deviation of all its workpieces after the i-th agent executes action a i where i is the total completion time deviation of all its workpieces after the i-th agent executes action a.
[0037] According to another aspect of the present invention, the following technical solutions are also provided:
[0038] A distributed blocking flow shop scheduling method based on deep reinforcement learning, which uses the agent trained by the above training method, inputs the current observation state of the agent into its Actor, and can output the optimal decision-making action of the agent that minimizes the total completion time deviation of all agents.
[0039] According to another aspect of the present invention, the following technical solutions are also provided:
[0040] A distributed blocking flow shop scheduling agent training system based on deep reinforcement learning, which regards a workshop as an agent, and each agent includes four deep reinforcement learning networks: Actor, Critic, targetActor, and targetCritic. The training objective is to obtain the optimal network parameters of Actor so that Actor can make an optimal decision on whether the agent receives a new workpiece to be scheduled, which minimizes the total completion time deviation of all workpieces in all agents. The completion time deviation is the absolute value of the completion time of the workpiece minus the due date. This system is used to train any agent A, including:
[0041] A data acquisition module, which is used to randomly select a set of data (s, a, s′, r) from a pre-acquired training set, where s is the set of current observation states of each agent, a is the set of actions decided by the Actor of each agent in the current observation state, s′ is the set of new observation states of each agent after executing the corresponding actions in a, and r is the sum of the reward values obtained by each agent after executing the corresponding actions in a. The smaller the total completion time deviation of all its workpieces after the agent inserts a new workpiece to be scheduled, the greater the obtained reward value;
[0042] An evaluation module, which is used to evaluate the scheduling scheme of the Actor network parameters θ of agent A under the current observation state set s μ and under the condition that each agent executes the corresponding actions in a, and obtain an evaluation value. The greater the evaluation value, the better the scheduling scheme;
[0043] An Actor network parameter update module, which is used to update the Actor network parameters θ of agent A by using the gradient ascent method μ for update;
[0044] The Critic network parameter update module is used to update the Critic network parameter θ of Agent A by using the gradient descent method of the TD error of the deep Q network, Q so that the Critic of Agent A is closer to its target Critic;
[0045] The network parameter replication module is used to replicate the updated Actor network parameter θ of the Actor network parameter update module across generations to the target Actor of Agent A, and replicate the updated Critic network parameter θ of the Critic network parameter update module across generations to the target Critic of Agent A; μ Q
[0046] The repeated training module repeatedly calls the data acquisition module, the evaluation module, the Actor network parameter update module, the Critic network parameter update module, and the network parameter replication module until the set number of training times is reached or the gradient converges in the Actor network parameter update module.
[0047] According to another aspect of the present invention, the following technical solutions are also provided:
[0048] A distributed blocking flow shop scheduling system based on deep reinforcement learning, comprising:
[0049] The scheduling module is used to input the current observation state of the agent trained by the above training method into its Actor, and then the optimal decision-making action of the agent that minimizes the total completion time deviation of all agents can be output.
[0050] Generally speaking, compared with the prior art through the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:
[0051] 1. The distributed blocking flow shop scheduling and agent training method and system based on deep reinforcement learning provided by the present invention regard a workshop as an agent. Each agent includes four deep reinforcement learning networks: Actor, Critic, targetActor, and targetCritic. By training the agent, the optimal network parameters of the Actor are obtained, enabling the Actor to make an optimal decision that minimizes the total completion time deviation of all workpieces in all agents regarding whether the agent receives a new workpiece to be scheduled. Thus, during online decision-making, only by inputting the current observation value of the agent into the Actor, the optimal decision action that minimizes the total completion time deviation of all workpieces in all agents can be output. The present invention is a data-driven scientific decision-making method with high decision-making efficiency. When a new workpiece arrives, it makes local modifications based on the original scheduling plan, can accurately assign priorities to newly inserted workpieces, and is applicable to various processing scenarios with strong compatibility.
[0052] 2. The distributed blocking flow shop scheduling and agent training method and system based on deep reinforcement learning provided by the present invention improves the VND algorithm. Before the agent schedules a new workpiece, an optimized original scheduling plan is arranged for each workshop through the improved VND algorithm. That is, for a batch of workpieces to be processed randomly assigned to each workshop, the internal workpiece sorting in each workshop is the optimal original scheduling plan with the minimum total completion time deviation, and it is combined with the multi-agent deep deterministic policy gradient algorithm (MADDPG) of the present invention. On the basis of static scheduling optimization, it learns workpiece selection in a dynamic order-insertion environment, which is closer to the actual production dynamic scheduling problem. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 is the flowchart of the distributed blocking flow shop scheduling method based on deep reinforcement learning in the preferred embodiment of the present invention;
[0054] Figure 2 is the structural diagram of the Actor network in the preferred embodiment of the present invention;
[0055] Figure 3 is the structural diagram of the Critic network in the preferred embodiment of the present invention;
[0056] Figure 4 is the schematic diagram of the offline training method of Agent A in the preferred embodiment of the present invention;
[0057] Figure 5 is the schematic diagram of the online decision-making method of the executing agent in the preferred embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0059] As Figure 1 shown, the present invention provides a distributed blocking flow shop scheduling method based on deep reinforcement learning. The method includes the following steps:
[0060] S1. Establish a multi-shop scheduling simulation environment for distributed blocking flow shop scheduling.
[0061] In this multi-shop scheduling simulation environment, the workpiece pool contains new workpieces waiting in line for scheduling. Each new workpiece is regarded as the same processing operation sequence to be provided by the next shop, but the processing duration of the same operation of different new workpieces in the intelligent body can be different. Each shop is regarded as an intelligent body and has the same processing capacity. Each shop is exactly the same, including the machines and their quantities, and there is no time limit between consecutive operations. For example, each new workpiece requires a total of 5 operations from 1 to 5. The time of each operation may be the same or different, and each intelligent body can provide the continuous processing process of these 1-5 operations.
[0062] The processing state of the intelligent body includes: the set of workpieces to be processed in the intelligent body, and the processing state of the set of workpieces being processed. The processing states of each intelligent body may be different. For example, the first working station (corresponding to the first operation) of some intelligent bodies is currently idle, while the first working station of some intelligent bodies is currently processing, etc.; each intelligent body can only receive new workpieces and start processing when the first working station is vacant, and cannot quit the current workpiece and start processing a new workpiece when the current workpiece on the first working station is half processed; there is no buffer between consecutive operations, and the workpiece cannot leave the current working station until the next working station is available for processing. All new workpieces waiting for scheduling in the workpiece pool are urgent workpieces. Once scheduled to a certain intelligent body, the intelligent body inserts the new workpiece, and starts processing the cut-in workpiece immediately after the first working station becomes vacant. Then, the processing time of all unstarted workpieces in the intelligent body is postponed accordingly.
[0063] The multi-shop scheduling simulation environment takes the arrival of new workpieces as a dynamic event, and triggers each intelligent body to make a decision between the new workpiece and the originally planned workpiece when the current workpiece of the corresponding first operation is processed: whether to insert the new workpiece; when a new workpiece is selected by an intelligent body, it is removed from the workpiece pool until all new workpieces in the workpiece pool are removed, completing the distributed blocking flow shop scheduling.
[0064] To better conform to the current production and manufacturing mode, the objective of the present invention is to minimize the total completion time deviation, that is, to minimize the sum of the absolute values of the completion time of all workpieces (including the workpieces being processed and the workpieces scheduled for processing on all agents) minus the due time, which is defined as follows:
[0065]
[0066] C j is the completion time of workpiece j, DT j is the due time of workpiece j, and N is the total number of workpieces.
[0067] S2. Obtain multiple sets of data on the interaction between each agent and the multi-shop scheduling simulation environment as the training set for subsequent training of each agent.
[0068] Each agent includes four networks: Actor, Critic, targetActor (Actor target network), and targetCritic (Critic target network).
[0069] Actor is used to make a decision on whether to insert a new workpiece. Its inputs are two data: the total completion time deviation of the current scheduling plan executed by the agent, and the total completion time deviation of the scheduling plan after inserting the new workpiece. Its output is the action of the decision, that is, whether to insert the new workpiece. The structure of the Actor network in the present invention is as Figure 2 shown. The input of the Actor network is two observation values obs1 and obs2 in the observation state of the agent. After passing through three fully connected layers and a softmax function, the probability distribution of the action is calculated, and the action with a relatively large probability is taken as the action of the decision.
[0070] Critic is used to evaluate the action decided by Actor, and optimize the parameters of the Actor network according to the evaluation value. The optimization objective of the parameters of the Actor network is to make Actor make an optimal decision, so that the total completion time deviation of the scheduling plan after the decision is minimized. The structure of the Critic network is as Figure 3 shown. The input of the Critic network is the observation values and actions of all agents. After passing through three fully connected layers, the Q value, that is, the evaluation value, is calculated.
[0071] targetActor is used to copy the updated network parameters of Actor across generations to assist in calculating the target value for updating the parameters of the Actor network. The network structure of targetActor is the same as that of Actor.
[0072] targetCritic is used to copy the updated network parameters of Critic across generations to assist in calculating the target value for updating the parameters of the Critic network. The network structure of targetCritic is the same as that of Critic.
[0073] A set of data for each agent interacting with the multi - workshop scheduling simulation environment is denoted as (s, a, s′, r), where the current observation state set s of each agent is s=(o1,...,o i ,...,o n ), n is the total number of agents, o i is the current observation state of the i - th agent, o i =(b i1 ,b i2 ), b i1 is the total completion - time deviation of all workpieces in the i - th agent under the current execution scheduling scheme, b i2 is the total completion - time deviation of all workpieces in the i - th agent after inserting the new workpiece to be scheduled; the action set a of each agent is a=(a1,...,a i ,...,a n ), a i is the action of the Actor of the i - th agent for making a decision on the new workpiece to be scheduled; the new observation state set s′ of each agent after executing the action is s′=(o1′,...,o i ′,...,o n ′), o i ′ is the new observation state of the i - th agent after executing the action a i , o i ′=(b i1 ′,b i2 ′), b i1 ′ is the total completion - time deviation of all workpieces of the i - th agent after executing the action a i , b i2 ′ is the total completion - time deviation of all workpieces of the i - th agent after inserting the next new workpiece to be scheduled after executing the action a i ; r is the sum of the reward values of each agent after executing the action.
[0074] The method for obtaining the above data (s, a, s′, r) is as follows:
[0075] S21, Initialize the multi - workshop scheduling simulation environment and each agent, including initializing the number and queuing order of new workpieces in the workpiece pool, the number of processes and process times of each workpiece, the current processing state in each agent, and the network parameters of the four networks of each agent.
[0076] Through the improved VND (variable neighborhood descent) algorithm, the present invention obtains the initial scheduling plan for each workshop after initialization (i.e., the current processing state within each intelligent body). The initial scheduling plan for each workshop is represented by the processing sequences of all workpieces in each workshop initially. The improved VND algorithm includes the following steps:
[0077] S211 Randomly assign all initially unprocessed workpieces to N workshops, so that each workshop generates an initial sequence of unprocessed workpieces {J1, J2, …, Ji};
[0078] S212 In the sequence of unprocessed workpieces in each workshop, randomly select two workpieces and re-insert them at all possible positions within the workshop to obtain a set of sequences of unprocessed workpieces, i.e., the first sequence set. Calculate the total completion time deviation of all sequences in the first sequence set. The sequence with the minimum deviation is the optimal one, and replace the initial sequence of unprocessed workpieces in this workshop; thus, the optimization of the workpiece sequences in all workshops is completed;
[0079] S213 After the optimization within each workshop is completed, optimization operations also need to be carried out among the workshop sequences. Select the workshop fmax with the largest completion time deviation and the workshop fmin with the smallest completion time deviation. Move each workpiece in the sequence of fmax to all insertable positions in fmin to obtain a set of sequences, i.e., the second sequence set. Each sequence contains the sequences of the fmax and fmin workshops. Calculate the total completion time deviation of all sequences in the second sequence set. The sequence with the minimum total completion time deviation is the optimal one, and replace the sequences in fmax and fmin according to this optimal sequence; Exchange the positions of each workpiece in the sequence of fmax with each workpiece in fmin to obtain a set of sequences, i.e., the third sequence set. Each sequence contains the sequences of the fmax and fmin workshops. Calculate the total completion time deviation of all sequences in the third sequence set. The sequence with the minimum total completion time deviation is the optimal one, and replace the sequences in fmax and fmin according to this optimal sequence;
[0080] S214 If the total completion time deviation of the sequences after optimization in all workshops in step (T113) is smaller than that before optimization, return to step S212; otherwise, the VND algorithm ends.
[0081] Furthermore, before each decision of the intelligent body, the present invention re-prioritizes the remaining new workpieces in the workpiece pool. The specific steps are as follows:
[0082] S215 Calculate the remaining time (LT) of all workpieces according to the processing time and due date DT of all new workpieces and the decision time deciT. The LT of workpiece j is calculated as follows:
[0083]
[0084] Among them, p j,k is the processing time of workpiece j at stage k, and S is the set of processing stages.
[0085] S216 sorts all workpieces in ascending order of LT to form a new sequence, and starts the insertion operation in the same way as mentioned above in the VND algorithm to obtain the optimal processing order. The first new workpiece is the next decision-making new workpiece.
[0086] S22, each agent obtains the total completion time deviation of all workpieces within the agent under its current execution scheduling scheme, and the total completion time deviation of the scheduling scheme if the new workpiece to be scheduled is inserted, that is, obtains the current observation state of the agent, so as to obtain the current observation state set s.
[0087] S23, the Actor of each agent inputs its current observation state and outputs the decision-making action. If multiple agents all decide to execute the insertion of the new workpiece, a random agent is selected from these multiple agents to execute the insertion action, so as to ensure that only one agent executes the insertion action for a new workpiece, so as to obtain the action set a.
[0088] S24, each agent executes the decision-making action, calculates the reward value of each agent executing the action by using the reward value function, and obtains the sum r of the reward values of each agent executing the action. Remove the newly inserted workpiece from the workpiece pool.
[0089] The objective function of the dynamic job shop scheduling of the present invention is to minimize the total completion time deviation. The completion deviation of a workpiece can only be determined after all processes of the workpiece are completed. Therefore, the designed reward function is that the more rewards obtained, the better the decision of the Actor. All agents need to work together, so they share all rewards. The reward value calculation formula for each agent is as follows:
[0090]
[0091] Among them, rew i is the reward value of the i-th agent, is the total completion time deviation of all workpieces of the i-th agent after executing the action a i . The reward value is the negative of the total completion time deviation of the scheduling scheme after the agent executes the action . The larger , the smaller the negative reward value rew i .
[0092] r is the sum of the reward values of each agent executing the action, that is
[0093] S25. Determine whether there are still new workpieces to be scheduled in the workpiece pool. If so, return to step S22 to start the next scheduling, and use the current observation state set calculated in step S22 as the new observation state set s' after each agent executes an action in this scheduling. If not, return to step S21 until the data (s, a, s', r) of the set is obtained. Put the obtained data (s, a, s', r) into the experience replay pool D. The obtained data (s, a, s', r) of the set serves as the training set for subsequent training of each agent.
[0094] S3. Use multi-agent deep reinforcement learning and adopt the training set data obtained in step S2 to perform offline training on each agent to obtain the network parameters that can make optimal decisions.
[0095] As Figure 4 shown, taking an arbitrary agent A as an example, the offline training method for agent A includes the following steps:
[0096] S31. Randomly select a set of data from the training set, input the current observation state set s and the action sets a of each agent into the evaluation function, and evaluate the scheduling plan in which each agent executes action a under the current observation state set s through the evaluation function, that is, calculate the evaluation value. The larger the evaluation value, the better the scheduling plan. μ The evaluation function formula is as follows:
[0097] The evaluation function formula is as follows:
[0098]
[0099] In the formula, J(θ μ ) is the evaluation value, θ μ is the Actor network parameter of agent A, θ Q is the Critic network parameter of agent A, Q is the function of the Critic network of agent A, Q(s, a|θ Q )|s = s t , a = μ(s t ) means that when the current observation state set s = s t , with the Actor network parameter of agent A being θ μ , the Q function value obtained when each agent executes the corresponding action in the action set a is μ, which is the set of functions of all agents' Actor networks, and μ(s|θμ)|s = s t means that when the current observation state set s = s t , with the Actor network parameter of agent A being θ μUnder this condition, the Actors of each agent each execute the corresponding action in the action set a. p is the set of all current observation states. Refers to the expected value under the conditions of s t ~p, a~μ, that is, the evaluation value.
[0100] S32. Use the gradient ascent method to update the Actor network parameters θ of agent A μ The update objective, that is, the optimization objective, is to make the Actor make a decision on whether the agent receives a newly scheduled workpiece, so that the total completion time deviation of all workpieces in all agents is minimized.
[0101] The gradient calculation formula is as follows:
[0102]
[0103] S33. Use the gradient descent method of the TD error of DQN (Deep Q Network) to update the Critic network parameters θ of agent A Q To make the Critic network of agent A closer to its target Critic, that is, to minimize the error.
[0104] The TD error is as follows:
[0105]
[0106] In the formula, a' = μ'(s'|θ μ' ) is the action set decided by the targetActor of each agent under the observation state set s', γ is the attenuation factor, θ Q' is the network parameter of the targetCritic of agent A, θ μ' is the network parameter of the targetActor of agent A, is the set of functions of the networks of the targetActors of all agents, μ θi' is the function of the targetActor network of the i-th agent, E s,a,r,s' is the expected value under the conditions of s, a, r, s′.
[0107] The order of step S32 and step S33 can be interchanged.
[0108] S34. Copy the Actor network parameters θ of agent A updated in step S32 μ across generations to the targetActor of agent A, and copy the Critic network parameters θ of agent A updated in step S33 Q across generations to the targetCritic of agent A.
[0109] Intergenerational replication directly copies the updated network parameters to the target network after one or more loops, that is, directly copies the updated Actor network parameters to the targetActor, or copies them after looping a specified number of times, or directly copies the updated Critic network parameters to the targetCritic, or copies them after looping a specified number of times.
[0110] S35, return to execute step S31 until the set number of training times is reached or the gradient calculation formula converges.
[0111] Use the above offline training method of steps S31 - S35 to perform offline training on all agents.
[0112] S4, the agent directly inherits the network parameters of the Actor after offline training and makes a quick decision on the new scheduling instance.
[0113] The process of online decision-making is as Figure 5 shown. Input the current observation state into the Actor, that is, the total completion time deviation of the scheduling plan before and after inserting the new workpiece. The Actor can then output the optimal decision action that minimizes the total completion time deviation of all agents. Moreover, the executing agent can also learn from the new scheduling instance, and then continuously update the network parameters, that is, the scheduling strategy, to improve the decision-making performance.
[0114] An embodiment of the present invention also provides a distributed blocking flow shop scheduling agent training system based on deep reinforcement learning. This system regards a workshop as an agent. Each agent includes four deep reinforcement learning networks: Actor, Critic, targetActor, and targetCritic. The training objective is to obtain the optimal network parameters of the Actor, so that the Actor can make an optimal decision that minimizes the total completion time deviation of all workpieces in all agents on whether the agent receives a new workpiece to be scheduled. The completion time deviation is the absolute value of the completion time of the workpiece minus the due date. This system is used to train any agent A, including:
[0115] A data acquisition module, which is used to randomly select a set of data (s, a, s′, r) from a pre-acquired training set. Among them, s is the set of current observation states of each agent, a is the set of actions decided by the Actor of each agent in the current observation state, s′ is the set of new observation states of each agent after executing the corresponding action in a, and r is the sum of the reward values obtained by each agent after executing the corresponding action in action a. The smaller the total completion time deviation of all workpieces of the agent after inserting the new workpiece to be scheduled, the greater the obtained reward value;
[0116] An evaluation module, which is used to evaluate the scheduling scheme of each agent performing the corresponding action in a through an evaluation function for the Actor network parameters θ of agent A under the current observation state set s, and obtain an evaluation value. The larger the evaluation value, the better the scheduling scheme. μ Under, the scheduling scheme of each agent performing the corresponding action in a is evaluated to obtain an evaluation value. The larger the evaluation value, the better the scheduling scheme.
[0117] An Actor network parameter update module, which is used to update the Actor network parameters θ of agent A by using the gradient ascent method. μ For updating;
[0118] A Critic network parameter update module, which is used to update the Critic network parameters θ of agent A by using the gradient descent method of the TD error of the deep Q network, so that the Critic of agent A is closer to its target Critic. Q For updating, making the Critic of agent A closer to its target Critic;
[0119] A network parameter replication module, which is used to replicate the updated Actor network parameters θ of the Actor network parameter update module to the target Actor of agent A across generations, and replicate the updated Critic network parameters θ of the Critic network parameter update module to the target Critic of agent A across generations. μ Replicate the updated Actor network parameters θ of the Actor network parameter update module to the target Actor of agent A across generations, and replicate the updated Critic network parameters θ of the Critic network parameter update module to the target Critic of agent A across generations. Q For replication across generations to the target Critic of agent A;
[0120] A repeated training module, which repeatedly calls the data acquisition module, the evaluation module, the Actor network parameter update module, the Critic network parameter update module, and the network parameter replication module until the set number of training times is reached or the gradient converges in the Actor network parameter update module.
[0121] An embodiment of the present invention also provides a distributed blocking flow shop scheduling system based on deep reinforcement learning, including:
[0122] A scheduling module, which is used to use the agents trained by the above offline training method, input the current observation state of the agent into its Actor, and then the optimal decision-making action of the agent that minimizes the total completion time deviation of all agents can be output.
[0123] Among them, the specific implementation manners of each module of each system can refer to the description in the method embodiment, and the embodiment of the present invention will not be repeated.
[0124] Those skilled in the art can easily understand that the above are only the preferred embodiments of the present invention, and are not used to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A training method for an intelligent agent of a distributed blocking flow shop scheduling based on deep reinforcement learning, characterized in that Regarding a workshop as an agent, each agent includes four deep reinforcement learning networks: Actor, Critic, targetActor, and targetCritic. The training objective is to obtain the optimal network parameters of Actor so that Actor can make an optimal decision on whether the agent receives a new workpiece to be scheduled, minimizing the total completion time deviation of all workpieces within all agents. The completion time deviation is the absolute value of the difference between the completion time of the workpiece and the due date. The training method for any agent A includes the following steps: (T1) Randomly select a set of data (s, a, s′, r) from the pre-acquired training set. Here, s is the set of current observation states of each agent, a is the set of actions decided by the Actor of each agent under the current observation state, s′ is the set of new observation states of each agent after executing the corresponding actions in a, and r is the sum of the reward values obtained by each agent after executing the corresponding actions in action a. The smaller the total completion time deviation of all workpieces of the agent after inserting the new workpiece to be scheduled, the greater the obtained reward value. (T2)Evaluate the scheduling scheme of each agent performing the corresponding action in a under the Actor network parameters θ of agent A in the current observation state set s through the evaluation function, and obtain the evaluation value; the larger the evaluation value, the better the scheduling scheme; μ Under the conditions, evaluate the scheduling scheme of each agent performing the corresponding action in a, and obtain the evaluation value; the larger the evaluation value, the better the scheduling scheme; (T3) Use the gradient ascent method to update the Actor network parameters θ of Agent A μ ; (T4) Use the gradient descent method of the TD error of the deep Q network to update the parameters θ of the Critic network of agent A Q to make the Critic of agent A closer to its target Critic; (T5) Copy the updated Actor network parameters θ in step (T3) μ to the targetActor of agent A with a generation gap, and copy the updated Critic network parameters θ in step (T4) Q to the targetCritic of agent A with a generation gap; (T6) Return to execute step (T1) until the set number of training times is reached or the gradient converges in step (T3). In step (T1), s = (o1,..., o i ,..., o n ), where n is the total number of agents, o i is the current observation state of the i-th agent, o i = (b i1 , b i2 ), b i1 is the total completion time deviation of all workpieces in the i-th agent under the current execution scheduling scheme, and b i2 is the total completion time deviation of all workpieces in the i-th agent after inserting the new workpiece to be scheduled; a = (a1,..., a i ,..., a n ), a i is the action of the Actor of the i-th agent for making a decision on the new workpiece to be scheduled; s′ = (o1′,..., o i ′,..., o n ′), o i ′ is the new observation state of the i-th agent after executing the action a i , o i ′ = (b i1 ′, b i2 ′), b i1 ′ is the total completion time deviation of all workpieces of the i-th agent after executing the action a i , and b i2 ′ is the total completion time deviation of all workpieces of the i-th agent after inserting the next new workpiece to be scheduled after executing the action a i .
2. The intelligent agent training method for distributed blocking flow shop scheduling based on deep reinforcement learning according to claim 1, wherein In step (T2), the evaluation function is as follows: where, J(θ μ ) is the evaluation value, θ μ are the Actor network parameters of agent A, θ Q are the Critic network parameters of agent A, Q is the function of the Critic network of agent A, Q(s,a|θ Q )|s = s t ,a = μ(s t ) means that when the current observation state set is s = s t , the Actor network parameters of agent A are θ μ , the Q function values obtained by each agent performing the corresponding actions in the action set a respectively; μ is the set of functions of the Actor networks of all agents, μ(s|θ μ )|s = s t means that when the current observation state set is s = s t , the Actor network parameters of agent A are θ μ , each agent's Actor performs the corresponding actions in the action set a respectively; p is the set of all current observation states, and E is the expected value.
3. The intelligent agent training method for distributed blocking flow shop scheduling based on deep reinforcement learning according to claim 2, wherein In step (T3), the gradient calculation formula of the gradient ascent method is as follows:
4. The intelligent agent training method for distributed blocking flow shop scheduling based on deep reinforcement learning according to claim 3, wherein In step (T4), the formula for the TD error is as follows: L(θ Q ) = E s,a,r,s' [(Q(s, a|θ Q ) - y) 2 , y = r + γQ'(s', a'|θ Q' )|a' = μ'(s'|θ μ' ) where a' = μ'(s'|θ μ' ) is the set of actions decided by each agent's targetActor under the observation state set s', γ is the attenuation factor, θ Q' are the network parameters of the targetCritic of agent A, θ μ' are the network parameters of the targetActor of agent A, is the set of functions of all agents' targetActor networks, μ θi' is the function of the targetActor network of the i-th agent, and E is the expected value.
5. The intelligent agent training method for distributed blocking flow shop scheduling based on deep reinforcement learning according to claim 1, characterized in that, The method for obtaining the training data (s, a, s′, r) is as follows: (T11) Initialize the number of new workpieces and the queuing order in the workpiece pool, the number of processes and process times of each workpiece, the current processing state within each agent, and the network parameters of the four networks of each agent. (T12) Each agent obtains the total completion time deviation of all workpieces under the current scheduling plan, as well as the total completion time deviation of all workpieces within the agent if the new workpiece to be scheduled is inserted, that is, obtains the current observation state of the agent, thereby obtaining the set of current observation states s. (T13) The Actor of each agent inputs its current observation state and outputs the decided action. If multiple agents all decide to execute inserting the new workpiece, randomly select one agent from these multiple agents to execute the insertion action, thereby obtaining the set of actions a. (T14) Each agent executes the decided action, calculates the reward value of each agent's executed action using the reward value function, and obtains the sum of the reward values of each agent's executed action r; remove the newly inserted workpiece from the workpiece pool. (T15) Determine whether there are still new workpieces to be scheduled in the workpiece pool. If so, return to step (T12) to start the next scheduling, and use the set of current observation states calculated in step (T12) as the set of new observation states s′ of each agent after executing the action in this scheduling. If not, return to step (T11) until a set of data (s, a, s′, r) is obtained.
6. The intelligent agent training method for distributed blocking flow shop scheduling based on deep reinforcement learning according to claim 5, wherein In step (T11), the steps for initializing the current processing state within each agent are as follows: (T111) Randomly assign all initial workpieces to be processed to N workshops, so that each workshop generates an initial sequence of workpieces to be processed. (T112) In the sequence of workpieces to be processed in each workshop, randomly select two workpieces and re-insert them at all possible positions in this workshop to obtain a set of sequences of workpieces to be processed, namely the first sequence set. Calculate the total completion time deviation of all sequences in the first sequence set. The sequence with the smallest deviation is the optimal one, and replace the initial sequence of workpieces to be processed in this workshop; thus, the optimization of the workpiece sequences in all workshops is completed. (T113) Select the workshop with the largest completion time deviation fmax and the workshop with the smallest completion time deviation fmin. Move each workpiece in the sequence of fmax to all insertable positions in fmin to obtain a set of sequences, namely the second sequence set. Each sequence contains the sequences of the two workshops fmax and fmin. Calculate the total completion time deviation of all sequences in the second sequence set. The sequence with the smallest total completion time deviation is the optimal one, and replace the sequences in fmax and fmin according to this optimal sequence; Exchange the positions of each workpiece in the sequence of fmax with each workpiece in fmin to obtain a set of sequences, namely the third sequence set. Each sequence contains the sequences of the two workshops fmax and fmin. Calculate the total completion time deviation of all sequences in the third sequence set. The sequence with the smallest total completion time deviation is the optimal one, and replace the sequences in fmax and fmin according to this optimal sequence. (T114) If the total completion time deviation of the sequences after optimization in all workshops in step (T113) is smaller than that before optimization, return to step (T112); otherwise, the initialization of the current processing state of each intelligent body is completed.
7. A distributed blocking flow shop scheduling method based on deep reinforcement learning, characterized in that, Using the intelligent body trained by the training method described in any one of claims 1-6, input the current observation state of the intelligent body into its Actor, and the optimal decision action of the intelligent body that minimizes the total completion time deviation of all intelligent bodies can be output.
8. A distributed blocking flow shop scheduling agent training system based on deep reinforcement learning, characterized in that The system regards a workshop as an intelligent body. Each intelligent body includes four deep reinforcement learning networks: Actor, Critic, targetActor, and targetCritic. The training objective is to obtain the optimal network parameters of Actor, so that Actor can make an optimal decision on whether the intelligent body receives a new workpiece to be scheduled, which minimizes the total completion time deviation of all workpieces in all intelligent bodies. The completion time deviation is the absolute value of the difference between the completion time of the workpiece and the due date. The system is used to train any intelligent body A, including: A data acquisition module, which is used to randomly select a set of data (s, a, s′, r) from a pre-acquired training set. Among them, s is the set of current observation states of each intelligent body, a is the set of actions decided by the Actor of each intelligent body under the current observation state, s′ is the set of new observation states of each intelligent body after executing the corresponding actions in a, and r is the sum of the reward values obtained by each intelligent body after executing the corresponding actions in action a; the smaller the total completion time deviation of all workpieces of the intelligent body after inserting the new workpiece to be scheduled, the greater the obtained reward value. An evaluation module for evaluating, through an evaluation function, the scheduling scheme of each agent executing the corresponding action in a under the current observation state set s and the Actor network parameters θ of the agent A μ to obtain an evaluation value; the larger the evaluation value, the better the scheduling scheme; The Actor network parameter update module is used to update the Actor network parameters θ of Agent A by using the gradient ascent method μ for updating; The Critic network parameter update module is used to update the Critic network parameter θ of agent A by using the gradient descent method of the TD error of the deep Q network, Q so that the Critic of agent A is closer to its target Critic; A network parameter replication module, which is used to replicate the Actor network parameters θ updated by the Actor network parameter update module μ across generations to the target Actor of Agent A, and replicate the Critic network parameters θ updated by the Critic network parameter update module Q across generations to the target Critic of Agent A; The repeated training module repeatedly calls the data acquisition module, the evaluation module, the Actor network parameter update module, the Critic network parameter update module, and the network parameter copying module until the set number of training times is reached or the gradient converges in the Actor network parameter update module; In the data acquisition module, s = (o1,...,o i ,...,o n ), where n is the total number of agents, o i is the current observation state of the i-th agent, o i = (b i1 ,b i2 ), b i1 is the total completion time deviation of all workpieces in the i-th agent under the current execution scheduling scheme, b i2 is the total completion time deviation of all workpieces in the i-th agent after inserting the new workpiece to be scheduled; a = (a1,...,a i ,...,a n ), a i is the action of the Actor of the i-th agent for making a decision on the new workpiece to be scheduled; s′ = (o1′,...,o i ′,...,o n ′), o i ′ is the new observation state of the i-th agent after executing the action a i , o i ′ = (b i1 ′,b i2 ′), b i1 ′ is the total completion time deviation of all workpieces of the i-th agent after executing the action a i , and b i2 ′ is the total completion time deviation of all workpieces of the i-th agent after inserting the next new workpiece to be scheduled after executing the action a i .
9. A distributed blocking flow shop scheduling system based on deep reinforcement learning, characterized in that, It includes: The scheduling module is used to use the trained agent of the training method described in any one of claims 1-6. When the current observation state of the agent is input to its Actor, the optimal decision-making action of the agent that minimizes the total completion time deviation of all agents can be output.
Citation Information
Patent Citations
Intelligent agent training method and device, computer equipment and storage medium
CN113919482A
Multi-agent adaptive sampling strategy generation method
CN113952733A