Multi-agent dynamic flexible job shop scheduling solving method based on D3QN network

By adopting a multi-agent scheduling system based on D3QN network in the dynamic flexible operation workshop, the problem of poor scheduling performance in dynamic environments is solved, real-time and efficient scheduling and resource allocation are achieved, and production efficiency and intelligence are improved.

CN120146146AActive Publication Date: 2025-06-13KUNMING UNIV OF SCI & TECH

Patent Information

Application Number
CN202510630588.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-06-13
Estimated Expiration
2045-05-16

AI Technical Summary

Technical Problem

The existing dynamic flexible operation workshop scheduling scheme is difficult to ensure scheduling performance when facing variability and random dynamic environments, and the traditional deep reinforcement learning method of individual agents has problems such as poor scalability and difficulty in distributing reward credits when dealing with multi-agent environments.

Method used

Using a multi-agent scheduling system based on D3QN network, by constructing a mathematical model of a dynamic flexible work workshop, processing information is obtained in real time, rescheduling strategies are generated, and problem-solving steps are distributed and parallelized through a multi-agent architecture to achieve global optimization.

Benefits of technology

It realizes real-time and efficient scheduling of workpieces and machines in a dynamic environment, reduces production time and resource costs, and meets the efficiency of production processing and the intelligence of production resource allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146146A_ABST
    Figure CN120146146A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-agent dynamic flexible job shop scheduling solving method based on a D3QN network, and belongs to the technical field of agent deep reinforcement learning and production scheduling. According to the method, the dynamic change of the machined workpiece and the machine can be monitored in real time, a real-time scheduling decision is given timely and efficiently according to different emergencies such as workpiece emergency insertion and machine damage, and the time cost and the resource cost of production are effectively reduced. The method comprises the steps that in the actual production process of a dynamic flexible workshop, machining information in the dynamic flexible workshop is obtained in real time, and the machining information comprises machining workpiece information and machining machine information; when it is determined that a rescheduling event occurs based on the processing information, a rescheduling strategy is generated according to the real-time states of workpieces and machines at the rescheduling point of the dynamic event based on the mathematical model of the dynamic flexible job shop scheduling problem, and rescheduling is executed until the production process is completed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of agent deep reinforcement learning and production scheduling, and particularly relates to a method for solving dynamic flexible job shop scheduling by multiple agents based on a D3QN network. Background Art

[0002] Production scheduling is an essential core link in the production management of manufacturing enterprises, directly affecting production costs and productivity. It is an important part of modern supply chains and manufacturing systems because a suitable production scheduling plan can improve machine utilization, ensure on-time delivery, and reduce inventory costs. The flexible job shop scheduling problem (FJSP) is one of the most widely applied problems in this field. FJSP is an extension of the job shop scheduling problem, where the processing machine for each operation is uncertain, and each operation of each workpiece can be processed on multiple alternative processing machines, and the processing times required for selecting different processing machines are different. For decades, FJSP has been widely studied, and many existing methods for solving FJSP are based on a static shop floor production environment, giving a definite scheduling plan. However, there are still differences from the actual situation. In actual processing production, typical internal disturbances, such as machine breakdowns and operator absences, need to be dealt with promptly to ensure the high efficiency of production. In addition to these internal disturbances, in a highly competitive market environment, companies enhance their competitiveness by offering customers a high degree of flexibility, which also leads to a rapid increase in the number of new orders and cancellations. Therefore, the existing scheduling plans for FJSP may no longer be suitable for production systems in a dynamic environment and may even result in poor scheduling performance. Therefore, it is necessary to develop dynamic scheduling to handle these interference events so that the production system can operate continuously and efficiently.

[0003] In recent years, the research on dynamic flexible job shop scheduling (DFJSP) has received extensive attention, and many solutions have been proposed. Among them, the most widely used are scheduling rules and meta-heuristics. Scheduling rules immediately handle dynamic events to achieve the best time efficiency. However, for the decisions at the time points of dynamic events, scheduling rules cannot guarantee their local optimality, and due to the variability of the dynamic environment, it is very difficult for decision-makers to select the best rules at a specific time point, which highly depends on professional knowledge. The basic idea of meta-heuristics is to decompose the dynamic scheduling problem into corresponding static scheduling sub-problems, and then use meta-heuristic algorithms to solve these static problems. For example, meta-heuristic algorithms such as genetic algorithm (GA), ant colony optimization (ACO), particle swarm optimization (PSO), and artificial bee colony algorithm (ABC) are used to solve them respectively. The solutions (new job insertions) they obtain often have higher quality, but usually come with high time consumption and infeasible real-time scheduling. For the continuous changes in the dynamic environment in the job shop, at each decision time point, while ensuring the efficiency of real-time performance, how to select suitable scheduling rules from numerous scheduling rules is the main motivation of this paper.

[0004] To select the most suitable scheduling rule at each scheduling decision time point, the dynamic scheduling process is regarded as a Markov decision process (MDP), and a dynamic selection strategy is formulated through a deep reinforcement learning (DRL) algorithm.

[0005] Traditional deep reinforcement learning only considers a single agent and controls the agent through the DRL algorithm. However, with the continuous complexity of the job shop scheduling environment, traditional deep reinforcement learning algorithms will be troubled by poor scalability. When multiple agents explore the environment simultaneously, as the dimension of the action space increases, the training will become chaotic and inefficient. This challenge is even more difficult in the dynamic job shop scheduling problem because the results (rewards) of the agents' actions usually have latency and sparsity, making it very difficult to allocate the credit of rewards to formulate cooperation strategies for all agents. Summary of the Invention

[0006] To overcome the above technical defects, the present invention provides a method for solving dynamic flexible job shop scheduling by multi-agent based on the D3QN network. The main solutions are as follows:

[0007] During the actual production process of the dynamic flexible job shop, the processing information in the dynamic flexible job shop is obtained in real time, and the processing information includes processing workpiece information and processing machine information;

[0008] When it is determined based on the processing information that a rescheduling event occurs, a rescheduling strategy is generated for the real-time states of workpieces and machines at the rescheduling point for the dynamic event based on the mathematical model of the dynamic flexible job shop scheduling problem, and rescheduling is executed until the production process is completed;

[0009] Among them, the mathematical model of the dynamic flexible job shop scheduling problem is constructed based on the initial parameter information of the dynamic flexible job shop, including scheduling objectives and constraints. The scheduling objectives are: minimizing the total tardiness of all jobs and minimizing the variance of the utilization rate of all machines. The constraints are: there are a total of O operations for n jobs and m machines. The operations of each job are processed in a fixed order, and each operation has a fixed set of processing machines, including the general constraints and dynamic constraints of FJSP.

[0010] Rescheduling events are dynamic / uncertain events including the insertion of emergency jobs and unexpected machine failures. The rescheduling point is the time point when an emergency job is inserted or a machine fails.

[0011] Optionally, in some possible implementation manners, the specific operations of generating a rescheduling strategy and performing rescheduling include:

[0012] 1) Update the status of jobs and machines at the rescheduling point:

[0013] Based on the obtained processing information, initially, the job and machine status update module at the rescheduling point updates the action status information, obtains the set of optional jobs according to whether the predecessor operation of the current operation of the job is completed; obtains the set of optional machines according to the set of processing machines of the current operation of the job and the current machine status set; when a rescheduling event occurs, the job and machine status update module at the rescheduling point updates the status information again until all jobs are processed.

[0014] 2) The job agent generates job scheduling information:

[0015] After the status information at the rescheduling point is updated, the job agent assigns the next job to be processed according to the set of optional jobs and generates job scheduling information. The job scheduling information includes the job number and the operation number.

[0016] 3) The machine agent generates machine scheduling information:

[0017] Based on the job scheduling information and the set of rescheduling optional machines, the machine agent arranges the processing machines for the jobs and generates machine scheduling information. The machine scheduling information includes the machine number.

[0018] Optionally, in some possible implementation manners, the general constraints of FJSP include: ① Each machine is available at time zero. ② All arrived jobs can be processed at time zero. ③ There is no precedence relationship between the operations of different jobs. ④ Each machine can process at most one operation at a time. ⑤ Each operation should be processed in a non-preemptive manner without interruption. ⑥ All operations belonging to the same job should be processed one by one in a fixed order, i.e., precedence constraints.

[0019] Optionally, in some possible implementation manners, the dynamic constraint conditions include: ①) the transportation time and the setup time can be ignored, ② the buffer between machines is not restricted, ③ if the operation is interrupted due to a machine failure, the remaining processing time is equal to the total processing time minus the completed processing time, and ④ each job contains different types of operations, and each job has a due date that must be met, otherwise it will be considered overdue.

[0020] Optionally, in some possible implementation manners, minimizing the total tardiness time of all jobs is expressed as: 。

[0021] Optionally, in some possible implementation manners, minimizing the variance of the utilization rate of all machines is expressed as: 。

[0022] In addition, the present invention also provides a multi-agent scheduling system, including: a rescheduling point workpiece and machine state update module, a workpiece agent module, a machine agent module, and a joint reward mapping agent module, which are used to execute the above-mentioned method for solving the dynamic flexible job shop scheduling based on the D3QN network.

[0023] The present invention can achieve the following technical effects by adopting the above technical solutions:

[0024] By introducing a multi-agent scheduling system based on D3QN, the problem of dynamic flexible job shop scheduling is solved. For the variability and randomness existing in this scheduling problem, a multi-agent and distributed agent architecture is adopted to distribute and parallelize the problem-solving steps, realizing the coherence from step optimization to global optimization. In the actual factory processing process, the invention herein can monitor the dynamic changes of the processed workpieces and machines in real time, and give real-time scheduling decisions in a timely and efficient manner for different emergencies, such as the emergency insertion of workpieces and machine damage, effectively reducing the production time cost and resource cost. Different from the traditional single-agent deep reinforcement learning method, the multi-agent scheduling system of the present invention can not only meet the high efficiency of production processing but also meet the intelligence of production resource allocation when dealing with production scheduling problems with a large amount of data and many sudden events. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The drawings exemplarily show the embodiments and form a part of the specification, and are used together with the written description of the specification to explain the exemplary implementation manners of the embodiments. The shown embodiments are only for illustrative purposes and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0026] Figure 1 It is a schematic diagram of the basic structure of the interaction between the reinforcement learning agent and the environment in the present invention;

[0027] Figure 2 It is a process schematic diagram of the method for multi-agent solving of dynamic flexible job shop scheduling based on the D3QN network in the present invention;

[0028] Figure 3 It is a process schematic diagram of the rescheduling process in the method for multi-agent solving of dynamic flexible job shop scheduling based on the D3QN network in the present invention;

[0029] Figure 4 It is a process schematic diagram of the optional machine set update method in the present invention;

[0030] Figure 5 It is a system diagram of the network framework for multi-agent solving of dynamic flexible job shop scheduling based on the D3QN network in the present invention;

[0031] Figure 6 It is a scheduling Gantt chart for an example in the present invention with 14 initial workpieces, 8 newly inserted workpieces, and 8 machines;

[0032] Figure 7 It is a schematic diagram of the change in the total delay time during the training phase of the multi-agent solving of dynamic flexible job shop scheduling based on the D3QN network in the present invention. Detailed implementation manners

[0033] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0034] It should be noted that the descriptions involving "first", "second", etc. in the embodiments of the present invention are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.

[0035] In the description of the present invention, it should be understood that the numerical labels before the steps do not identify the order of execution of the steps, but are only used to facilitate the description of the present invention and distinguish each step, and thus cannot be understood as a limitation to the present invention.

[0036] As one of the most powerful sequential decision-making tools in the Markov decision process (MDP), reinforcement learning (RL) is suitable for solving dynamic optimization problems in workshop production and can achieve relatively high average performance. RL algorithms drive agents to learn through rewards. Agents interact with the environment, perform operations on the environment according to the perceived state, and then obtain rewards from the environment. Referring to Figure 1 , Figure 1 in denotes the action at time step denotes the reward from time step to

[0037] The purpose of the present invention is to solve the problems existing at present, such as the difficulty in multi-agent reward credit assignment in the dynamic flexible job shop, resulting in disordered training among agents, poor scheduling efficiency, and waste of production resources. In this invention patent, the workshop is represented as a multi-agent system including workpiece agents, machine agents, and joint reward agents. Among them, the workpiece agents select appropriate workpieces to be processed, the machine agents select appropriate processing machines, and the joint reward agents generate mappings of joint rewards. This invention patent uses the monotonic reward function decomposition algorithm and uses neural networks to approximate the joint rewards of all agents' actions, which can represent the corresponding relationship between joint rewards and each agent's rewards, and finally form a scheduling production strategy.

[0038] In the technical solution of the embodiment of the present invention, first, a mathematical model for solving the dynamic flexible job shop scheduling problem based on the D3QN network is constructed. The construction scheme and the model are as follows:

[0039] 1) Obtain the initial parameter information of the dynamic flexible job shop.

[0040] The parameter information includes: the number of workpieces, the number of machines, the number of newly inserted workpieces, the status of the machines, the number of processes, the processing attribution relationship between workpieces and machines, the process sequence of workpiece processing, and the processing time of workpieces on machines.

[0041] It should be noted that the parameter information here is the initial parameter information, which represents the overall workpiece and machine information obtained before the scheduling starts; the information of workpieces and machines during the scheduling process in the following text refers to the information obtained at that moment during each rescheduling.

[0042] 2) According to the obtained parameter information of the dynamic flexible job shop, construct a mathematical model for the dynamic flexible job shop scheduling problem, including constraint conditions and scheduling objectives.

[0043] Due to the insertion of emergency workpieces and unexpected machine failures, the goal of this study is to minimize the total tardiness of all jobs and minimize the variance of the system machine utilization rate in the case of such dynamic events. Therefore, the scheduling objective of the mathematical model in the present invention is: to minimize the total tardiness of all workpieces and minimize the variance of all machine utilization rates; where minimizing the total tardiness of all workpieces is defined as the sum of the extension times when the actual processing completion times of all workpieces are greater than the due date after the scheduling is completed; the variance of machine utilization rate is defined as the average of the squared differences between the utilization rates of each machine and the overall average utilization rate after the scheduling is completed.

[0044] Furthermore, the constraint conditions are: there are a total of O processes for n workpieces and m machines, and the process processing of each workpiece has a fixed order, and each process has a fixed set of processing machines. n is the total number of workpieces, and m is the total number of machines.

[0045] Specifically, the constraint conditions include the general constraint conditions of FJSP and dynamic constraint conditions.

[0046] Among them, the general constraint conditions of FJSP are as follows:

[0047] ① Each machine is available at time zero.

[0048] ② All workpieces can be processed at time zero.

[0049] ③ There is no prior priority between the processes of different workpieces.

[0050] ④ Each machine can process at most one operation at a time.

[0051] ⑤ Each operation should be processed in a non-preemptive manner without interruption.

[0052] ⑥ All operations belonging to the same job should be processed one after another in a fixed order (precedence constraint).

[0053] Correspondingly, the dynamic constraint conditions are as follows:

[0054] ① The transportation time and setup time can be ignored.

[0055] ② The buffer between machines is not restricted.

[0056] ③ If an operation is interrupted due to a machine failure, the remaining processing time is equal to the total processing time minus the completed processing time.

[0057] ④ Each job contains different types of operations, and each job has a due date that must be met, otherwise it will be considered overdue.

[0058] Construct a mathematical model for solving the dynamic flexible job shop scheduling problem based on the D3QN network. The dynamic flexible job shop scheduling problem is briefly described as follows:

[0059] There are jobs being processed on machines . Each job contains operations, which is represented as a set of O processes. denotes the th operation of job . Any machine that can process belongs to the set of machines that can process . The actual processing time of machine for is , denoted as the completion time of operation . The arrival time and due date of workpiece and are

[0060] Based on the data model constructed from the operations described in 1) and 2) above, it is specifically as follows:

[0061] (1)

[0062] In the formula, Represents the total number of jobs, Represents the index of the job, ∈ ; Represents The completion time of the workpiece after a certain number of operations; Represents the workpiece 's processing period; Represents the part of the completion time that exceeds the processing period;

[0063] (2)

[0064] In the formula, Represents the variance of the machine utilization rate in the whole system; m represents the total number of machines; Represents the total processing time of the k-th machine, k ∈ m; Represents the processing time of all machines, l ∈ m, k ∈ l;

[0065] (3)

[0066] In the formula, Represents the number of operations belonging to the workpiece , Represents the index of the number of operations, ∈ ; Represents the operation 's completion time;

[0067] (4)

[0068] In the formula, = 0 represents that the completion time of the initial operation is 0;

[0069] (5)

[0070] In the formula, Represents the set of machines that can process the operation ; Represents a binary variable, that is, whether the operation is assigned to the k-th machine;

[0071] (6)

[0072] In the formula, Represents the start time of the operation ; Represents the operation 's processing time on the machine , where Represents the th machine considering the processing time;

[0073] (7)

[0074] In the formula, when = 1, represents the completion time under the first operation; represents the processing time on the machine under the first operation ; represents the arrival time of the workpiece ; represents the binary variable under the first operation;

[0075] (8)

[0076] In the formula, represents the completion time of the predecessor operation of the same workpiece;

[0077] (9)

[0078] In the formula, represents the completion time of the operation of other workpieces, that is, the machine conflict constraint; represents the processing time of the operation of other workpieces; represents the binary variable of the operation of other workpieces; represents the operation and represents the sequence of operations of other workpieces on the same machine;

[0079] (10)

[0080] In the formula, represents the non-processing time of the workpiece before the completion time; represents the time at the rescheduling point after the occurrence of the dynamic time; represents the due date of the workpiece ;

[0081] (11)

[0082] In the formula, represents the average processing time of the operation on the machine where it can be processed;

[0083] (12)

[0084] In the formula, represents the remaining operation preprocessing time of the workpiece ; represents the sequence number of the operation to be performed at the rescheduling point after the occurrence of the dynamic event;

[0085] (13)

[0086] In the formula, represents the idle delay time of the workpiece;

[0087] (14)

[0088] In the formula, represents the delay time of the workpiece; represents the final completion time;

[0089] (15)

[0090] In the formula, represents the total tardiness rate generated by the scheduling;

[0091] (16);

[0092] Formula (1) represents the total delay of all jobs. Formula (2) represents the variance of the machine utilization rate in the whole system. Formula (3) represents the sum of the completion times of all jobs. Formula (4) represents that the completion time of each operation must be non - negative. Formula (5) represents that each operation can only be assigned to one machine. Formula (6) represents that the completion time of an operation minus the start time equals the processing time. Formula (7) ensures that a job can only be processed after its arrival time. Formula (8) guarantees the precedence constraint of the operations of the same workpiece. Formula (9) guarantees the execution order constraint of the operations of the workpieces on each machine. Formula (10) represents the remaining time of the workpiece from the current time to the due date. Formula (11) represents the average processing time of the operation on one of the machines in the set of alternative machines. Formula (12) represents the processing time of the remaining operations of the workpiece when rescheduling occurs Formula (13) represents the idle time of the workpiece Formula (14) represents the delay time of the workpiece Formula (15) represents the tardiness rate of the scheduling system. Formula (16) represents the total processing time on the machine

[0093] The decision variables corresponding to the above mathematical model are:

[0094] , and, .

[0095] ​Secondly, based on the established mathematical model of the dynamic flexible job shop scheduling problem, schedule the rescheduling events that occur in the actual production process of the dynamic flexible job shop. Refer to Figure 2 , and the specific solution is as follows:

[0096] 201. During the actual production process of the dynamic flexible job shop, obtain the processing information in the dynamic flexible job shop in real time.

[0097] The processing information includes processing workpiece information and processing machine information; the processing workpiece information includes: the number of processed workpieces, the number of newly inserted workpieces, the number of processes of the workpiece, the process sequencing of the workpiece, the processing time of each process on different machines, the average processing time of the process, the arrival time of the workpiece, and the completion time of the workpiece; the processing machine information includes: machine status information, and the processing attribution relationship between the workpiece and the machine.

[0098] 202. When it is determined that a rescheduling event occurs based on the processing information, generate a rescheduling strategy for the real-time status of the workpiece and the machine at the rescheduling point for the dynamic event based on the mathematical model of the dynamic flexible job shop scheduling problem, and execute the rescheduling until the production process is completed.

[0099] The problem studied in this paper is how to generate a rescheduling strategy in real time and efficiently to handle dynamic events. Usually, the dynamic / uncertain events in the actual production process mainly include the insertion of emergency workpieces and unexpected machine failures. Therefore, the rescheduling events in the solution of the present invention are dynamic / uncertain events including the insertion of emergency workpieces and unexpected machine failures. Correspondingly, the rescheduling point is the time point when the workpiece is inserted urgently or the machine fails.

[0100] Therefore, the embodiment of the present invention focuses on the generation of scheduling at the rescheduling point. The rescheduling process of the dynamic flexible job shop scheduling problem is as follows: At the initial stage of scheduling, all workpieces can be processed on all machines. After workpieces have been arranged to start processing on all machines, the production process proceeds. When a workpiece is inserted urgently or a machine fails, it is a rescheduling point. At this time, the remaining workpiece set and the optional machine set are updated, new scheduling information is generated, and the production process continues to proceed. By analogy, every time a rescheduling event occurs and a rescheduling point is generated, new scheduling information is generated until the production process is completed.

[0101] The present invention solves the dynamic flexible job shop scheduling problem by introducing a multi-agent scheduling system based on D3QN. A multi-agent and distributed agent architecture is adopted for the variability and randomness existing in the scheduling problem, distributing and parallelizing the problem-solving steps, and realizing the coherence from step optimization to global optimization. During the actual factory processing, the invention of this paper can monitor the dynamic changes of the processed workpieces and machines in real time, and give real-time scheduling decisions in a timely and efficient manner for different emergencies, such as the emergency insertion of workpieces and machine failures, effectively reducing the production time cost and resource cost. Different from the traditional single-agent deep reinforcement learning method, the multi-agent scheduling system of the present invention can meet both the high efficiency of production processing and the intelligence of production resource allocation when dealing with production scheduling problems with a large amount of data and many sudden events.

[0102] The rescheduling process in 202 above will be described in detail. Refer to Figure 3 , and its scheduling process specifically includes:

[0103] S1. Obtain the processed workpiece information and processing machine information.

[0104] S2. Update the status of workpieces and machines at the rescheduling point.

[0105] Based on the obtained processing information, initially, the status information update module of workpieces and machines at the rescheduling point updates the action status information, and obtains the set of optional workpieces according to whether the predecessor process of the current process of the workpiece is completed; obtains the set of optional machines according to the set of processing machines of the current process of the workpiece and the current machine status set, referring to Figure 4 . During the production process, when a rescheduling event occurs, the status information update module of workpieces and machines at the rescheduling point updates the status information again until all workpieces are processed.

[0106] S3. The workpiece agent generates workpiece scheduling information.

[0107] After the status information of the rescheduling point is updated, the workpiece agent allocates the next processed workpiece according to the set of optional workpieces, generates workpiece scheduling information, and the workpiece agent configures the trained workpiece D3QN evaluation network. The workpiece scheduling information includes the workpiece number and the process number.

[0108] S4. The machine agent generates machine scheduling information.

[0109] Based on the workpiece scheduling information and the set of rescheduling optional machines, the machine agent arranges the processing machine for the workpiece, generates machine scheduling information, and the machine agent configures the trained machine D3QN evaluation network to evaluate whether the workpiece is processed. If it is not processed, return to S2 to update the status again. If the processing is completed, output the optimal scheduling plan.

[0110] Reinforcement learning can be modeled as a Markov decision model represented as a five-tuple . represents the state space, represents the finite action space, is represented as the state transition distribution, . represents the discount factor. represents the reward function, . In reinforcement learning, the agent follows a specific policy and interacts with the surrounding environment. For each decision point , the agent observes the current state , and according to the policy selects an action , and then transitions to the next state by sampling , and immediately obtains an immediate reward . In the embodiment of the present invention, there are three agents. Taking the workpiece agent as an example, whenever a rescheduling event occurs, the environmental state space is updated to , at this time the agent extracts state features according to the new state. The state features extracted by each agent are different, but these state features are interrelated. Because the state features are interrelated, the subsequent action decisions between agents are correlated.

[0111] The state features of the workpiece agent are shown in Table 1:

[0112] Table 1: State features of the workpiece agent

[0113]

[0114] The state features of the machine agent are as follows in Table 2:

[0115] Table 2: State features of the machine agent

[0116]

[0117] The state features of the joint reward agent are as follows in Table 3:

[0118] Table 3: State features of the joint reward agent

[0119]

[0120] After the state features at a certain moment are updated, the rescheduling point workpiece and machine state update module updates the action state information. The workpiece agent and the machine agent combine their respective obtained state features and action state information, and select different actions and , the actions are jointly executed, and the environment promptly provides the combined reward . The combined reward agent, based on the state features obtained at a certain moment, executes actions , obtaining different rewards for the workpiece agent and the machine agent . The environmental state features are updated to , and the rewards obtained . The goal of deep reinforcement learning is to learn decision-making problems by maximizing the discounted cumulative reward. For any policy , the state-action value function ( function) is defined as:

[0121] Q π ( s , a ) = E π [ ∑ t = 0 T γ t r ( s t , a t ) | s 0 = s , a a = a ] ;

[0122] where represents the expectation under policy , represents the time horizon, represents the initial state, represents the initial action;

[0123] The ultimate goal of the three agents in the present invention is to learn the decision-making problem of dynamic job shop scheduling through the maximum discounted cumulative reward, such that the final scheduling policy meets the scheduling objectives.

[0124] As Figure 5 shown is the system diagram of the multi-agent solution for dynamic flexible job shop scheduling network framework based on the D3QN network. At each rescheduling point, the workpiece agent and the machine agent execute joint actions according to the workpiece and machine states in the current environment, and the environment gives the total reward. The reward agent maps the total reward to the workpiece agent and the machine agent respectively based on the current environment and the total reward as the state, and the two agents evaluate the action effects based on the obtained mapped rewards. The pseudo-code for the training phase steps of the multi-agent solution for dynamic flexible job shop scheduling network framework based on the D3QN network is as follows:

[0125] Input:

[0126] Shop parameters: number of workpieces , number of newly inserted workpieces , number of operations , number of machines ; Number of iterations: ; Experience replay buffer and buffer lower limit: ; ; Batch size: ;

[0127] Output:

[0128] Workpiece agent evaluation network parameters , machine agent evaluation network parameters , combined reward agent evaluation network parameters ;

[0129] Initial workpiece agent evaluation network parameters And assign the evaluation network parameters to the target network parameters ;

[0130] Initialize the machine agent evaluation network parameters And assign the evaluation network parameters to the target network parameters ;

[0131] Combined reward agent evaluation network parameters And assign the evaluation network parameters to the target network parameters ;

[0132] Initialize the experience replay buffers of the workpiece agent, machine agent, and reward combined agent respectively , and set the buffer size to ;

[0133] for episode = 1: do:

[0134] Generate an instance according to the workshop parameter information ;

[0135] Reset the scheduling environment and generate the initial state , and ;

[0136] for do:

[0137] Update the action status information in the rescheduling point workpiece and machine status update module;

[0138] Form the legal action space of the workpiece and machine

[0139] if the action space is not empty then;

[0140] According to Calculate the action of the workpiece agent ;

[0141] According to Calculate the action of the machine agent ;

[0142] Execute the action , and the scheduling environment feedbacks the combined reward R

[0143] Form the initial state of the reward - combined agent ;

[0144] According to Calculate the action of the reward - combined agent ;

[0145] Execute the action , schedule the reward feedback from the environment , and obtain ;

[0146] Schedule the environment to transfer to ;

[0147] Store the experience tuple( , , , ) into the workpiece - agent experience replay buffer ;

[0148] Store the experience tuple( , , , ) into the machine - agent experience replay buffer ;

[0149] Store the experience tuple( , , , ) into the reward - combined agent experience replay buffer ;

[0150] if the size of the workpiece - agent experience replay buffer > the lower limit of the top - level experience buffer then:

[0151] Prioritize and replay experience tuples of size in the workpiece - agent experience replay buffer( , , , ) to update the network parameters (see Network Update 1);

[0152] Prioritize and replay experience tuples of size in the machine - agent experience replay buffer;

[0153] Prioritize and replay experience tuples of size in the machine - agent experience replay buffer( , , , ) Update network parameters (see Network Update 1);

[0154] Endif.

[0155] if the size of the combined reward agent experience replay buffer > the lower limit of the underlying experience buffer then:

[0156] Prioritize the replay of experience tuples of size in the combined reward agent experience replay buffer ( , , , ) Update network parameters (see Network Update 2);

[0157] Endif.

[0158] Endif.

[0159] Endfor.

[0160] Endfor.

[0161] The pseudo - code for Network Update 1 is as follows:

[0162] 1: Input: Experience tuple: ( , , , ); Evaluation network parameters: ; Target network parameters: ; Target network update frequency: ; Reward discount factor: ;

[0163] 2: Calculate the maximum Q - value of the next state according to the target network parameters ;

[0164] 3:if scheduling ends then;

[0165] 4: TD error ;

[0166] 5: Else;

[0167] 6: TD error ;

[0168] 7:Endif;

[0169] 8: Gradient descent Update the evaluation network;

[0170] Every Next updated target network .

[0171] The pseudo-code for network update two is as follows:

[0172] 1: Input: Experience tuple: ( , , , ); Evaluation network parameters: ; Target network parameters: ; Target network update frequency: ; Reward discount factor: ;

[0173] 2: Calculate the maximum Q-value of the next state according to the target network parameters V ( s ; θ , β ) + [ A ( s , a ; θ , α ) − 1 | A | ∑ a ' A ( s , a ' ; θ , α ) ] , where represents the action space, represents the candidate actions in the action space, ∈ ; is the value function, related to the state characteristics, is the advantage function, related to both the state characteristics and the actions;

[0174] 3: if the scheduling ends then;

[0175] 4: TD error ;

[0176] 5: Else;

[0177] 6: TD error ;

[0178] 7: Endif;

[0179] 8: Gradient descent Update the evaluation network every times of updating the target network .

[0180] Training is completed. The scheduling Gantt chart for an instance with an initial number of workpieces of 14, a newly inserted number of workpieces of 8, and a number of machines of 8 is referred to Figure 6 .

[0181] During the system training phase, the change in the total delay time. In the Figure 7 shown line chart, it can be clearly seen that after the number of training times reaches 150, the total delay time can reach the lowest.

[0182] In addition, an embodiment of the present invention further provides a multi-agent scheduling system, which includes four modules, namely, a rescheduling point workpiece and machine status update module, a workpiece agent module, a machine agent module, and a joint reward mapping agent module. The rescheduling point workpiece and machine status update module is used to update the status of workpieces and machines when a rescheduling event occurs in the production process; the workpiece agent module is used to generate scheduling information for workpieces; the machine agent module is used to generate scheduling information for machines; and the joint reward mapping agent module is used to map and allocate joint rewards to corresponding agents.

[0183] The above multi-agent scheduling system is used to perform all operations in the method for solving the dynamic flexible job shop scheduling based on the D3QN network described in the above method embodiment.

[0184] Obviously, those skilled in the art should understand that the above modules or steps of the embodiments of the present invention can be implemented by general computer devices. They can be concentrated on a single computer device or distributed on a network composed of multiple computer devices. Optionally, they can be implemented by program codes executable by computer devices, so that they can be stored in a storage device and executed by computer devices. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. Thus, the embodiments of the present invention are not limited to any specific combination of hardware and software.

[0185] It should be noted that the above is only a preferred embodiment of the present invention, and thus does not limit the patent protection scope of the present invention. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied to other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A method for solving dynamic flexible job shop scheduling based on a multi-agent D3QN network, characterized in that: include: During the actual production process of the dynamic flexible job shop, the processing information in the dynamic flexible job shop is obtained in real time, and the processing information includes processing workpiece information and processing machine information; When a rescheduling event is determined to occur based on processing information, a rescheduling strategy is generated based on the mathematical model of the dynamic flexible job shop scheduling problem according to the real-time status of the workpiece and the machine at the rescheduling point where the dynamic event occurs, and the rescheduling is executed until the production process is completed; Among them, the mathematical model of the dynamic flexible job shop scheduling problem is constructed based on the initial parameter information of the dynamic flexible job shop, including scheduling objectives and constraints. Its scheduling objectives are: minimizing the total delay time of all workpieces and minimizing the variance of all machine utilization. Its constraints are: n workpieces have a total of O processes and m machines. The process processing of each workpiece has a fixed order, and each process has a fixed set of processing machines, including the general constraints and dynamic constraints of FJSP. Rescheduling events are dynamic / uncertain events including urgent workpiece insertion and unexpected machine failure. The rescheduling point is the time point when the workpiece is urgently inserted or the machine fails.

2. The method for solving dynamic flexible job shop scheduling based on D3QN network according to claim 1 is characterized in that: Generate a rescheduling strategy. The specific operations of rescheduling include: 1) Update the status of the rescheduling point workpiece and machine: Based on the obtained processing information, at the beginning, the rescheduling point workpiece and machine status update module updates the action status information, and obtains the optional workpiece set according to whether the predecessor process of the current process of the workpiece has been processed; obtains the optional machine set according to the processing machine set of the current process of the workpiece and the current machine status set; when a rescheduling event occurs, the rescheduling point workpiece and machine status update module updates the status information again until all workpieces are processed; 2) The artifact agent generates artifact scheduling information: After the status information of the rescheduling point is updated, the workpiece agent allocates the next processing workpiece according to the optional workpiece set and generates workpiece scheduling information, which includes the workpiece serial number and the process serial number; 3) The machine agent generates machine scheduling information: Based on the workpiece scheduling information and the rescheduling optional machine set, the machine agent arranges the processing machine of the workpiece and generates machine scheduling information, which includes the machine serial number.

3. The method for solving dynamic flexible job shop scheduling based on D3QN network according to claim 1 is characterized in that: The general constraints of FJSP include: ① Every machine is available at time zero, ② All arrived workpieces can be processed at time zero, ③ There is no priority between the processes of different workpieces, ④ Each machine can only process at most one operation at a time, ⑤ Each operation should be processed in a non-preemptive manner without interruption, ⑥ All operations belonging to the same job should be processed one by one in a fixed order, i.e., priority constraints.

4. The method for solving dynamic flexible job shop scheduling based on D3QN network according to claim 1 is characterized in that: The dynamic constraints include: ① transportation time and setup time are negligible, ② buffers between machines are unlimited, ③ if an operation is interrupted due to machine failure, the remaining processing time is equal to the total processing time minus the completed processing time, ④ each job contains different types of operations, and each job has a delivery deadline that must be met, otherwise it will be considered timed out.

5. The method for solving dynamic flexible job shop scheduling based on D3QN network according to claim 1 is characterized in that: Minimizing the total delay time of all workpieces is expressed as: ; In the formula, Indicates the total number of jobs. Indicates the index of the job. ∈ ; express The completion time of the workpiece after each operation; Representation of workpiece duration of the project; Indicates the part of the completion time that exceeds the construction period.

6. The method for solving dynamic flexible job shop scheduling based on D3QN network according to claim 1 is characterized in that: Minimizing the variance of all machine utilization is expressed as: ; In the formula, represents the variance of machine utilization in the whole system; m represents the total number of machines; represents the total processing time of the kth machine, k∈m; represents the processing time of all machines, l∈m, k∈l.

7. A multi-agent scheduling system, characterized in that: include: The rescheduling point workpiece and machine status update module, the workpiece agent module, the machine agent module, and the joint reward mapping agent module are used to execute the method for solving dynamic flexible job shop scheduling based on a multi-agent D3QN network as described in any one of claims 1-6 above.

Citation Information

Patent Citations

  • Dynamic workshop scheduling method based on Conv-Dueling and generalization representation

    CN116562584A

  • Multi-resource flexible job shop scheduling method based on multi-agent reinforcement learning

    CN118761572A

  • Distributed flexible job shop scheduling method and system based on double deep reinforcement learning and multi-layer intelligent agent

    CN119494504A

  • Dynamic flexible workshop scheduling method based on event-driven DDQN algorithm

    CN119596869A

Cited By

  • AGV flexible line body scheduling method and system suitable for digital intelligent factory

    CN122390411A