A production and logistics collaborative scheduling method for dynamic flexible job shop

By employing deep reinforcement learning and a multi-agent nested hierarchical framework, the problems of disturbance factors and production logistics coordination scheduling in flexible work workshops were solved, thereby optimizing production and logistics and improving the production efficiency and logistics cost-effectiveness of flexible work workshops.

CN119338156BActive Publication Date: 2026-02-06TONGJI UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411255437.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-09
Publication Date
2026-02-06
Estimated Expiration
2044-09-09

AI Technical Summary

Technical Problem

Existing research has failed to effectively address disturbances in flexible work workshops and has failed to simultaneously optimize the goals pursued by production and logistics activities, making it difficult to balance production efficiency and logistics costs.

Method used

A multi-agent nested hierarchical framework is constructed using deep reinforcement learning, including a target agent, a production agent, and a logistics agent. By designing reward functions and state features, and using multi-agent proximal policy optimization algorithms for collaborative training, the maximum completion time and total logistics cost are minimized. Considering logistics equipment failure scenarios and constraints, a dynamic flexible workshop production and logistics collaborative scheduling model is established.

Benefits of technology

It enables the simultaneous optimization of production and logistics goals in a dynamic and flexible workshop, improving production efficiency and logistics cost-effectiveness, adapting to disturbances in the actual manufacturing environment, and obtaining a scientific and reasonable scheduling scheme.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119338156B_ABST
    Figure CN119338156B_ABST
Patent Text Reader

Abstract

The application relates to a production and logistics collaborative scheduling method for a dynamic flexible job shop, comprising the following steps: obtaining real-time information of the dynamic flexible job shop, inputting a production and logistics collaborative scheduling problem model for the dynamic flexible job shop, solving by using a deep reinforcement learning method, and obtaining a workshop scheduling scheme; wherein the production and logistics collaborative scheduling problem model for the dynamic flexible job shop sets respective optimization targets for production activities and logistics activities, considers a fault scene of logistics equipment and a corresponding processing strategy, and the solving process is as follows: S301, designing a multi-agent nested hierarchical framework; S302, designing key elements of a Markov decision process of each agent; and S303, cooperatively training each agent by using a multi-agent proximal policy optimization algorithm. Compared with the prior art, the application can obtain a more reasonable and effective flexible job scheduling scheme.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of flexible job shop scheduling, and particularly relates to a production and logistics collaborative scheduling method for a dynamic flexible job shop. BACKGROUND

[0002] Currently, the increasingly diversified and customized market demand forces manufacturing industry to adopt a more flexible production mode. In this context, the flexible job shop with flexible processing advantage has become the first choice for the development of various manufacturing industries. Unlike the traditional job shop, the flexible job shop is equipped with a series of multi-functional machines to manufacture products. The use of multi-functional machines makes the processing path of each product in the workshop have great flexibility and diversity, which directly leads to the urgent need for frequent and flexible transfer of semi-finished products between machines. Especially in some flexible job shops with high machine versatility and high processing efficiency, such as aircraft parts production workshop and automobile parts production workshop, this transfer demand is more urgent and complex. It can be said that in the flexible job shop, the completion of manufacturing tasks is closely related to production activities and logistics activities.

[0003] More specifically, for production activities and logistics activities, the former mainly involves allocating multi-functional machines to form diversified processing paths to manufacture products, and the latter needs to arrange logistics resources according to the diversified processing paths to provide necessary semi-finished product transfer. Obviously, there is a complex coupling relationship between production and logistics, that is, flexible production arrangement derives complex transfer requirements, thereby restricting and challenging the effective organization of logistics; and the logistics arrangement determines the production rhythm, thereby affecting the flexibility and efficiency of production. In view of the necessity of production activities and logistics activities in the flexible job shop and the complex coupling relationship between them, it is necessary to collaboratively schedule them to ensure effective production.

[0004] At present, many studies have focused on the production and logistics collaborative scheduling problem in flexible job shop. For example, in the Chinese invention patent "Intelligent scheduling decision method for flexible job shop combined with transportation equipment constraints" (publication number CN112949077A), for the intelligent scheduling problem of flexible job shop considering transportation equipment constraints, a genetic algorithm is used to schedule the corresponding processing equipment and transportation equipment, and the optimization of the completion time is realized. In the Chinese invention patent "Multi-objective distributed flexible job shop scheduling method considering limited transportation resources" (publication number CN116187093A), a multi-objective comprehensive decision method is combined with an intelligent optimization algorithm to solve the multi-objective distributed flexible job shop scheduling problem considering limited transportation resources and low carbon, and the simultaneous optimization of the maximum completion time, total processing energy consumption and total processing quality is realized. In the Chinese invention patent "Flexible job shop scheduling method considering transportation time and adjustment time" (publication number CN116109084A), for the flexible job shop scheduling problem considering transportation time and adjustment time, an ant colony algorithm integrated with reinforcement learning is provided to optimize the completion time, machine total load, machine critical load and workpiece delivery period penalty value.

[0005] However, existing research mainly focuses on static environment and does not consider disturbance factors. In actual production, disturbance is inevitable, and ignoring disturbance will greatly limit the application value of the research. In addition, existing research mainly focuses on optimizing production activity-related targets (such as completion time, machine load, etc.) and ignores logistics-related targets. In fact, production and logistics are often delivered to different managers within the workshop for separate management, and they each pursue specific targets related to their own tasks or resources. Therefore, to achieve coordination between the two separately managed activities, it is essentially necessary to optimize the targets of both activities simultaneously, not just one of them. In summary, there is still a need to design a new production and logistics collaborative scheduling method to effectively cope with disturbances in actual production and simultaneously optimize the targets pursued by production and logistics activities. SUMMARY

[0006] The purpose of the present application is to overcome the defects of the prior art and provide a production and logistics collaborative scheduling method for dynamic flexible job shop to effectively cope with disturbances in actual production and simultaneously optimize the targets pursued by production and logistics activities.

[0007] The purpose of the present application can be achieved by the following technical solutions:

[0008] The present application provides a production and logistics collaborative scheduling method for dynamic flexible job shop, comprising the following steps:

[0009] Obtaining real-time information of the dynamic flexible job shop, inputting a production and logistics collaborative scheduling problem model facing the dynamic flexible job shop, and solving to obtain a workshop scheduling scheme;

[0010] The production and logistics collaborative scheduling problem model facing the dynamic flexible job shop is constructed in the following manner:

[0011] S1, setting optimization objectives of production activities and logistics activities respectively, setting task processing and resource allocation constraint conditions, and establishing a production and logistics collaborative scheduling problem model facing a static flexible job shop, wherein the optimization objective of the production activities is to minimize the maximum completion time, and the optimization objective of the logistics activities is to minimize the total logistics cost;

[0012] S2, updating the optimization objective of the logistics activities according to the failure scenarios of the logistics equipment and the corresponding processing strategies, further setting the task processing and resource allocation constraint conditions under disturbance, and adding them to the production and logistics collaborative scheduling problem model facing the static flexible job shop to establish a production and logistics collaborative scheduling problem model facing the dynamic flexible job shop;

[0013] The production and logistics collaborative scheduling problem model facing the dynamic flexible job shop is solved by using a deep reinforcement learning method, and the specific process is as follows:

[0014] S301, designing a multi-agent nested hierarchical framework adapted to the production and logistics collaborative scheduling problem model facing the dynamic flexible job shop, wherein the multi-agent nested hierarchical framework comprises a target agent, a production agent and a logistics agent, the production agent and the logistics agent are independent of each other and are respectively used for making decisions on production scheduling and logistics scheduling, and the target agent is used for guiding the decision direction of the production agent and the logistics agent;

[0015] S302, designing a reward function, a state feature and an action space according to the functions and management authorities of the agents in the multi-agent nested hierarchical framework;

[0016] S303, using a multi-agent proximal policy optimization algorithm to collaboratively train the agents, inputting the real-time information of the dynamic flexible job shop into the trained agents, and obtaining a workshop scheduling scheme.

[0017] Further, in step S1, the expression of the optimization objective of the production activities is as follows:

[0018]

[0019] Wherein, MS represents the maximum completion time, ct ij represents the completion time of operation j of manufacturing task i, I represents the total number of manufacturing tasks, and J idenotes the total number of operations required to complete manufacturing task i;

[0020] The expression of the optimization objective of the logistic activity is as follows:

[0021]

[0022] where TC denotes the total logistic cost, c denotes the commissioning cost required to activate an automated guided vehicle, uc denotes the logistic cost required for each unit of transportation time of the automated guided vehicle, y v is a binary variable, and is equal to 1 if and only if automated guided vehicle v is activated in the scheduling scheme, V denotes the total number of automated guided vehicles, nst ij and nct ij denote the start time and the end time of the non-load phase in the transportation task derived from operation j of manufacturing task i, respectively, lst ij and lct ij denote the start time and the end time of the load phase in the transportation task derived from operation j of manufacturing task i, respectively.

[0023] Further, the task processing and resource allocation constraints include:

[0024] The processing time of each operation of each manufacturing task is the sum of its start time and machine processing time, and the expression is as follows:

[0025]

[0026] where K denotes the total number of machines, st ij denotes the start time of operation j of manufacturing task i, pt ijk denotes the processing time of operation j of manufacturing task i on machine k, λ ijk is a binary variable, and is equal to 1 if and only if operation j of manufacturing task i is allocated to machine k for processing;

[0027] Each operation of each manufacturing task must start processing after the task is transported to the corresponding machine, and the expression is as follows:

[0028]

[0029] Each machine device can only process one operation at a time, and the expression is as follows:

[0030] st ij + M(1 - σ iji′j′k ) ≥ ct i′j′

[0031]

[0032] where σ iji′j′kis a binary variable equal to 1 if and only if operation j of manufacturing task i is processed on machine k immediately followed by operation j' of manufacturing task i', ct i′j′ denotes the completion time of operation j' of manufacturing task i', δ ijk is a binary variable equal to 1 if and only if operation j of manufacturing task i is processed on machine k first;

[0033] Each automated guided vehicle can transport only one manufacturing task at a time, and the expressions are as follows:

[0034] nst ij + M(1 - ω iji′j′v ) ≥ lct i′j′

[0035]

[0036] where ω iji′j′v is a binary variable equal to 1 if and only if operation j of manufacturing task i is processed on automated guided vehicle v immediately followed by operation j' of manufacturing task i', η ij is a binary variable equal to 1 if and only if operation j of manufacturing task i is processed on the assigned automated guided vehicle first;

[0037] Each automated guided vehicle has sufficient time to complete the non-load phase of the transportation task without being disturbed, and the expressions are as follows:

[0038]

[0039]

[0040]

[0041] where tt ll′ denotes the transportation time from location l to l', x ijv is a binary variable equal to 1 if and only if operation j of manufacturing task i is completed by automated guided vehicle v, β kl is a binary variable equal to 1 if and only if machine k is located at location l, λ i′j′k is a binary variable equal to 1 if and only if operation j' of manufacturing task i' is assigned to machine k for processing, λ ij-1k is a binary variable equal to 1 if and only if operation j-1 of manufacturing task i is assigned to machine k for processing;

[0042] The load phase of operation j of manufacturing task i corresponding to the transportation task can only start after the completion of the previous operation of this operation and after the assigned automated guided vehicle has completed its non-load phase, and the expressions are as follows:

[0043]

[0044]

[0045] wherein, lst ij denotes the start time of the load phase of the transport task derived from operation j of manufacturing task i;

[0046] Each automated guided vehicle has sufficient time to complete the load phase of the transport task without being disturbed, which is expressed as follows:

[0047]

[0048]

[0049] wherein, lct ij denotes the end time of the load phase of the transport task derived from operation j of manufacturing task i;

[0050] Each operation of each manufacturing task can only be assigned to one machine to be completed, which is expressed as follows:

[0051]

[0052] The transport task derived from each operation of each manufacturing task can only be assigned to one automated guided vehicle to be completed, which is expressed as follows:

[0053]

[0054] Any two operations can only have a tight-precedence-tight-succession relationship on one processing machine, which is expressed as follows:

[0055]

[0056] Any two operations can only have a tight-precedence-tight-succession relationship on one automated guided vehicle, which is expressed as follows:

[0057]

[0058] Further, in step S2, the logistics equipment takes the automated guided vehicle as the object, and the failure scenarios of the logistics equipment and the corresponding processing strategies are as follows:

[0059] If the automated guided vehicle fails when it is idle, it is repaired at the current position to restore availability;

[0060] If the automated guided vehicle fails in the non-load phase, it is moved to the position closer to the failure position between the start position and the destination position for repair;

[0061] If the automated guided vehicle fails during the loading phase, it will be moved to the location closer to the start location and the destination location for repair, and its loading task will be moved synchronously and rescheduled.

[0062] Further, the expression of the optimization objective of the updated logistics activities is as follows:

[0063]

[0064] Wherein, TC represents the total logistics cost, c represents the debugging cost required for enabling an automated guided vehicle, uc represents the logistics cost required for the automated guided vehicle per unit transportation time, y v is a binary variable, and is 1 only when the automated guided vehicle v is enabled in the scheduling scheme, V represents the total number of automated guided vehicles, nst ij and nct ij respectively represent the start time and the end time of the non-loading phase of the transportation task derived from the operation j of the manufacturing task i, lst ij and lct ij respectively represent the start time and the end time of the loading phase of the transportation task derived from the operation j of the manufacturing task i, at n represents the transportation time consumed by the logistics task being processed by the faulty vehicle when the fault n occurs, and N represents the total number of faults.

[0065] Further, the task processing and resource allocation constraint conditions under the disturbance include:

[0066] The automated guided vehicle after the fault can only start processing a new logistics task after being repaired, and the expression is as follows:

[0067]

[0068] Wherein, tr n represents the time when the fault n occurs, rt n represents the repair time required for repairing the fault n, is a binary variable, and is 1 only when the automated guided vehicle assigned to process the logistics task derived from the operation j of the manufacturing task i has experienced a fault before processing the task, ξ ij is a binary variable, and is 1 only when the logistics task derived from the operation j of the manufacturing task i is the first task processed by the assigned automated guided vehicle after repairing the fault, ξ ijn is a binary variable, and is 1 only when the logistics task derived from the operation j of the manufacturing task i is the first task processed by the assigned automated guided vehicle after repairing the fault n;

[0069] The transportation task interrupted due to the fault can be immediately restarted after the interruption, and the expression is as follows:

[0070]

[0071] where μijnis a binary variable equal to 1 if and only if the operation j of the manufacturing task i is interrupted by the failure n affecting the logistic task derived from it;

[0072] each automated guided vehicle still satisfies the constraints related to transportation in the task handling and resource allocation constraints when not directly affected by the disturbance;

[0073] each automated guided vehicle starts from the repair location to handle the first logistic task after repair, expressed as follows:

[0074]

[0075]

[0076]

[0077] where z ij represents the interruption of the loading phase of the logistic task derived from the operation j of the manufacturing task i by the failure, represents the transportation time from the location rj n to L+1, rj n represents the repair location of the failed automated guided vehicle after the n-th failure, λ ij-1k is a binary variable equal to 1 if and only if the operation j-1 of the manufacturing task i is assigned to the machine k for handling, β kl is a binary variable equal to 1 if and only if the machine k is located at the location l;

[0078] for the logistic tasks interrupted in the loading phase, each automated guided vehicle goes to the corresponding repair location to complete its non-loading phase tasks, expressed as follows:

[0079]

[0080]

[0081]

[0082]

[0083] where ω iji′j′v is a binary variable equal to 1 if and only if the operation j of the manufacturing task i corresponding transportation task is handled immediately after the operation j' of the manufacturing task i' corresponding transportation task on the automated guided vehicle v, η ij is a binary variable equal to 1 if and only if the operation j of the manufacturing task i corresponding transportation task is handled first on the assigned automated guided vehicle, xijv is a binary variable that is 1 if and only if the operation j of the manufacturing task i corresponds to a transport task that is completed by the automated guided vehicle v;

[0084] For a logistics task that is interrupted due to a failure occurring in the load phase, each automated guided vehicle starts from the corresponding repair location to complete the load phase task, and the expression is specifically as follows:

[0085]

[0086] Further, in step S301, the design of the multi-agent nested hierarchical framework is specifically as follows:

[0087] The target agent is deployed on the upper layer of the production agent and the logistics agent, and controls the production agent and the logistics agent;

[0088] The production agent is deployed on the upper layer of the logistics agent, and controls the logistics agent together with the target agent.

[0089] Further, in step S302, the state features of the target agent include resource utilization information, task processing information, and optimization preferences, the resource utilization information includes machine utilization information and logistics equipment utilization information, the task processing information includes production task processing information and logistics task processing information, and the action space of the target agent includes selecting optimization minimization maximum completion time and selecting minimum total logistics cost;

[0090] The state features of the production agent include machine utilization information, production task processing information, optimization preferences, and target agent decision information, and the action space of the production agent includes a plurality of combined scheduling rules for selecting the earliest to-be-processed operation of an unfinished manufacturing task and assigning it to a machine for processing;

[0091] The state features of the logistics agent include logistics equipment utilization information, logistics task processing information, optimization preferences, and target agent decision information, and the action space of the logistics agent includes a plurality of scheduling rules for assigning the derivative logistics task of the selected operation to an automated guided vehicle for processing.

[0092] Further, in step S302, the reward function includes a minimum maximum completion time target related reward function and a minimum total logistics cost related reward function, and the expression is specifically as follows:

[0093]

[0094]

[0095] wherein, RMS tto minimize the maximum completion time, w1 and w2 are weights, nms t denotes a normalized value of the change value of the minimum maximum completion time before and after the state transition, ms t denotes the change value of the minimum maximum completion time before and after the state transition, pms and maxp respectively represent the optimization preference for the minimum maximum completion time target and the maximum possible preference for one target, ms max and ms min are respectively the maximum value and the minimum value of the change value of the minimum maximum completion time before and after the state transition;

[0096]

[0097]

[0098] wherein RTC t is a reward function related to minimizing the total logistics cost, ntc t denotes a normalized value of the logistics cost generated by the current state transition, tc t denotes the logistics cost generated by the current state transition, ptc and maxp respectively represent the optimization preference for the minimum total logistics cost target and the maximum possible preference for one target, tc max and tc min are respectively the maximum value and the minimum value of the logistics cost generated by the current state transition.

[0099] Further, in step S303, the specific process of collaborative training of each agent by using the multi-agent proximal policy optimization algorithm is as follows:

[0100] a policy network and a value network are randomly initialized for each agent in the multi-agent nested hierarchical framework, the policy network is used to output a selected action according to a current state, and the value network is used to output a value estimate according to a state-action pair;

[0101] a random environment required for single training is initialized, at each decision point in the environment, an e-greedy strategy combined with noise is adopted, the interaction and decision transmission between the environment and the agent and the agent and the agent are sequentially executed according to the structure of the multi-agent nested hierarchical framework, and the interaction trajectory is accumulated into the corresponding experience replay pool, if the number of interaction trajectories in the experience replay pool reaches a fixed batch size, the parameters of the policy network and the value network are updated by using the Adam optimizer based on the update logic of the proximal policy optimization algorithm combined with the stochastic gradient descent, until the manufacturing tasks in the environment are all completed;

[0102] If the maximum number of training is not reached after all manufacturing tasks in the environment are completed, the random environment required for single training is reinitialized, the foregoing training process is repeated, the network parameters are updated, and the maximum number of training is reached.

[0103] Compared with the prior art, the present application has the following beneficial effects:

[0104] 1. The present application provides a production and logistics collaborative scheduling method for dynamic flexible job shop, which obtains a workshop scheduling scheme by solving a production and logistics collaborative scheduling problem model for dynamic flexible job shop, wherein the construction process of the production and logistics collaborative scheduling problem model for dynamic flexible job shop is as follows: first, the optimization objective of production activity is to minimize the maximum completion time, the optimization objective of logistics activity is to minimize the total logistics cost, the task processing and resource allocation constraint conditions are set, and a production and logistics collaborative scheduling problem model for static flexible job shop is established; second, according to the failure scenarios of logistics equipment and the corresponding processing strategies, the optimization objective of logistics activity is updated, the disturbance task processing and resource allocation constraint conditions are further set, and the production and logistics collaborative scheduling problem model for static flexible job shop is added to establish a production and logistics collaborative scheduling problem model for dynamic flexible job shop, which can optimize the respective goals pursued by the two activities of production and logistics, at the same time, considers the possible failure scenarios and influences of logistics equipment, extracts new constraint conditions, is more consistent with the actual manufacturing environment, and obtains a more scientific, reasonable and effective flexible job shop scheduling scheme.

[0105] 2、In the present application, a deep reinforcement learning method is used to solve the production and logistics collaborative scheduling problem model of a dynamic flexible job shop, and the specific process is as follows: first, a multi-agent nested hierarchical framework suitable for the production and logistics collaborative scheduling problem model of a dynamic flexible job shop is designed, the multi-agent nested hierarchical framework includes a target agent, a production agent and a logistics agent, wherein the production agent and the logistics agent are independent of each other and are used for decision-making of production scheduling and logistics scheduling respectively, the target agent is arranged in the upper layer of the production agent and the logistics agent and is used for guiding the decision direction of the production agent and the logistics agent, the production agent is arranged in the upper layer of the logistics agent and controls the logistics agent together with the target agent, the framework fully considers the functions and management authority of each agent and is consistent with the actual manufacturing environment; secondly, according to the functions and management authority of each agent in the multi-agent nested hierarchical framework, a reward function, a state feature and an action space are designed, the above process clearly defines the Markov key elements suitable for the actual decision-making process and provides a practical and specific solution for the problem; finally, the multi-agent near-optimal strategy optimization algorithm is used to cooperatively train each agent, the algorithm can more effectively utilize the collected data to accelerate the learning process, is suitable for processing complex multi-agent optimization problems and has good scalability and flexibility. BRIEF DESCRIPTION OF DRAWINGS

[0106] Figure 1 is a flowchart of the method of the present application;

[0107] Figure 2 is a structural schematic diagram of the multi-agent nested hierarchical framework;

[0108] Figure 3 is a workshop layout diagram of the embodiment,

[0109] wherein L / U represents a loading / unloading station, M m represents a machine m. DETAILED DESCRIPTION

[0110] The present application will be described in detail below in combination with the drawings and specific embodiments. The present embodiment is implemented on the premise of the technical solution of the present application, and detailed implementation modes and specific operation processes are given, but the protection scope of the present application is not limited to the following embodiments.

[0111] Embodiment:

[0112] The present embodiment provides a production and logistics collaborative scheduling method for a dynamic flexible job shop, as shown in Figure 1 , including the following steps:

[0113] S1, set optimization objectives of production activities and logistics activities respectively, set task processing and resource allocation constraints, and establish a production and logistics collaborative scheduling problem model for a static flexible job shop.

[0114] The optimization objective of production activities is to minimize the maximum completion time, which is expressed as follows:

[0115]

[0116] where MS represents the maximum completion time, ct ij represents the completion time of operation j of manufacturing task i, I represents the total number of manufacturing tasks, j i represents the total number of operations required to complete manufacturing task i.

[0117] The optimization objective of logistics activities is to minimize the total logistics cost, which is expressed as follows:

[0118]

[0119] where TC represents the total logistics cost, c represents the commissioning cost required to enable one automated guided vehicle, uc represents the logistics cost required for each unit of transportation time of the automated guided vehicle, y v is a binary variable, which is 1 only when the automated guided vehicle v is enabled in the scheduling scheme, V represents the total number of automated guided vehicles, nst ij and nct ij respectively represent the start time and end time of the non-loading phase in the transportation task derived from operation j of manufacturing task i, lst ij and lct ij respectively represent the start time and end time of the loading phase in the transportation task derived from operation j of manufacturing task i.

[0120] The task processing and resource allocation constraints include:

[0121] (1) The completion time of each operation of each manufacturing task is the sum of its start time and machine processing time, which is expressed as follows:

[0122]

[0123] where K represents the total number of machines, st ij represents the start time of operation j of manufacturing task i, pt ijk represents the processing time of operation j of manufacturing task i on machine k, λ ijk is a binary variable, which is 1 only when operation j of manufacturing task i is assigned to machine k for processing.

[0124] (2) Each operation of each manufacturing task can only start processing after the task is transported to the corresponding machine, which is expressed as follows:

[0125]

[0126] (3) Each machine device can only process one operation at a time, which is expressed as follows:

[0127] st ij + M(1 - σ iji′j′k ) ≥ ct i′j′

[0128]

[0129] where σ iji′j′k is a binary variable, and is 1 if and only if operation j of manufacturing task i is processed on machine k immediately after operation j' of manufacturing task i' is processed, ct i′j′ denotes the completion time of operation j' of manufacturing task i', and δ ijk is a binary variable, and is 1 if and only if operation j of manufacturing task i is processed on machine k first.

[0130] (4) Each automated guided vehicle can only transport one manufacturing task at a time, which is expressed as follows:

[0131] nst ij + M(1 - ω iji′j′v ) ≥ lct i′j′

[0132]

[0133] where ω iji′j′v is a binary variable, and is 1 if and only if operation j of manufacturing task i corresponding to the transportation task is processed on automated guided vehicle v immediately after operation j' of manufacturing task i' corresponding to the transportation task is processed, η ij is a binary variable, and is 1 if and only if operation j of manufacturing task i corresponding to the transportation task is processed on the assigned automated guided vehicle first.

[0134] (5) Each automated guided vehicle has sufficient time to complete the non-load phase of the transportation task without being disturbed, which is expressed as follows:

[0135]

[0136]

[0137]

[0138] where tt ll′denotes the transportation time from location l to l', x ijv is a binary variable that is equal to 1 if and only if operation j of manufacturing task i corresponds to a transportation task that is performed by automated guided vehicle v, β kl is a binary variable that is equal to 1 if and only if machine k is located at location l, λ i′j′k is a binary variable that is equal to 1 if and only if operation j' of manufacturing task i' is assigned to machine k for processing, λ ij-1k is a binary variable that is equal to 1 if and only if operation j-1 of manufacturing task i is assigned to machine k for processing.

[0139] (6) Operation j of manufacturing task i can start its load phase of the corresponding transportation task only after the completion of its preceding operation and after the assigned automated guided vehicle has finished its non-load phase, which is expressed as follows:

[0140]

[0141]

[0142] wherein, lst ij denotes the start time of the load phase of the transportation task derived from operation j of manufacturing task i.

[0143] (7) Each automated guided vehicle has sufficient time to complete the load phase of the transportation task without being disturbed, which is expressed as follows:

[0144]

[0145]

[0146] wherein, lct ij denotes the end time of the load phase of the transportation task derived from operation j of manufacturing task i.

[0147] (8) Each operation of each manufacturing task can be assigned to only one machine for completion, which is expressed as follows:

[0148]

[0149] (9) The transportation task derived from each operation of each manufacturing task can be assigned to only one automated guided vehicle for completion, which is expressed as follows:

[0150]

[0151] (10) Any two operations can have a tight precedence relationship only on one processing machine, which is expressed as follows:

[0152]

[0153] (11) Any two operations can only have a tight front-tight back relationship on one automated guided vehicle, which is expressed as follows:

[0154]

[0155] S2, according to the failure scenarios of the logistics equipment and the corresponding processing strategies, update the optimization objective of the logistics activities, and further set the disturbance task processing and resource allocation constraint conditions, add the production and logistics collaborative scheduling problem model for static flexible job shop, and establish the production and logistics collaborative scheduling problem model for dynamic flexible job shop.

[0156] Taking the automated guided vehicle as the logistics equipment object, in the flexible job shop, each automated guided vehicle may fail at any time, and the failure scenarios of the logistics equipment and the corresponding processing strategies are as follows:

[0157] (1) If the automated guided vehicle fails when it is idle, it is repaired at the current position to restore its availability.

[0158] (2) If the automated guided vehicle fails during the non-load stage, it is moved to the position closer to the failure position between the start position and the destination position for repair, avoiding the interference of in-road repair on other activities.

[0159] (3) If the automated guided vehicle fails during the load stage, it is moved to the position closer to the failure position between the start position and the destination position for repair, and the loaded task is moved synchronously and rescheduled, the logistics activities affected by the failure are rescheduled, and the failed vehicle is restored for use after repair.

[0160] The expression of the updated optimization objective of the logistics activities is as follows:

[0161]

[0162] Where TC represents the total logistics cost, c represents the debugging cost required to enable one automated guided vehicle, uc represents the logistics cost required by the automated guided vehicle per unit of transportation time, y v is a binary variable, which is 1 only when the automated guided vehicle v is enabled in the scheduling scheme, V represents the total number of automated guided vehicles, nst ij and nct ij respectively represent the start time and end time of the non-load stage in the transportation task derived from the operation j of the manufacturing task i, lst ij and lct ij respectively represent the start time and end time of the load stage in the transportation task derived from the operation j of the manufacturing task i, at nThis represents the transportation time consumed by the malfunctioning vehicle in handling the logistics task when malfunction n occurs, and N represents the total number of malfunctions.

[0163] The constraints on task processing and resource allocation under disturbances include:

[0164] (1) An automated guided vehicle that has experienced a malfunction can only begin processing new logistics tasks after it has been repaired. The specific expression is as follows:

[0165]

[0166] Among them, tr n rt represents the time when fault n occurs. n This represents the repair time required to fix fault n. A binary variable, ξ is 1 if and only if the automated guided vehicle (AGV) assigned to handle the logistics task derived from operation j of manufacturing task i has experienced a malfunction before handling that task. ij For binary variables, ξ is 1 if and only if the logistics task derived from operation j of manufacturing task i is the first task handled by the assigned automated guided vehicle after a fault is repaired. ijn It is a binary variable, and is 1 if and only if the logistics task derived from the operation j of manufacturing task i is the first task handled by the assigned automated guided vehicle after the fault n is repaired.

[0167] (2) Transportation tasks interrupted by faults can be restarted immediately after the interruption, as shown in the following expression:

[0168]

[0169] Wherein, μijn is a binary variable, which is 1 if and only if the logistics task derived from operation j of manufacturing task i is interrupted by fault n.

[0170] (3) Each automated guided vehicle still satisfies the transportation-related constraints in the task processing and resource allocation constraints when it is not directly affected by disturbances (i.e., the load process of the logistics task it is responsible for is not affected or it is not the first task to be processed after the fault is repaired).

[0171] (4) Each automated guided vehicle departs from the repair location to handle the first logistics task after the repair, as shown in the following expression:

[0172]

[0173]

[0174]

[0175] Among them, z ijdenotes the disruption of the load phase of the logistics task derived from operation j of manufacturing task i due to the failure, denotes the transportation time from location rl n to L+1, rl n denotes the repair location of the automated guided vehicle after the nth failure, λ ij-1k is a binary variable, which is 1 if and only if operation j-1 of manufacturing task i is assigned to machine k for processing, β kl is a binary variable, which is 1 if and only if machine k is located at location l;

[0176] (5) For the logistics task whose load phase is disrupted, each automated guided vehicle will go to the corresponding repair location to complete its non-load phase task, and the expression is as follows:

[0177]

[0178]

[0179]

[0180]

[0181] wherein ω iji′j′v is a binary variable, which is 1 if and only if the operation j of manufacturing task i corresponding to the transportation task is processed immediately after the operation j' of manufacturing task i' corresponding to the transportation task on the automated guided vehicle v, η ij is a binary variable, which is 1 if and only if the operation j of manufacturing task i corresponding to the transportation task is processed first on the assigned automated guided vehicle, x ijv is a binary variable, which is 1 if and only if the operation j of manufacturing task i corresponding to the transportation task is completed by the automated guided vehicle v.

[0182] (6) For the logistics task whose load phase is disrupted due to the failure occurring in the load phase, each automated guided vehicle will start from the corresponding repair location to complete the load phase task, and the expression is as follows:

[0183]

[0184] S3, a deep reinforcement learning method is used to solve the production and logistics collaborative scheduling problem model of a dynamic flexible job shop, and a real-time job shop scheduling scheme is obtained. The specific process is as follows:

[0185] S301, a multi-agent nested hierarchical framework suitable for the production and logistics collaborative scheduling problem model of a dynamic flexible job shop is designed.

[0186] Specifically, the multi-agent nested hierarchical framework includes a target agent, a production agent, and a logistics agent. Considering that the above problems require separate decision-making for production scheduling and logistics scheduling, and that the two are usually managed separately in actual manufacturing, a production agent and a logistics agent with independent state spaces and action spaces are designed to make decisions for production scheduling and logistics scheduling, respectively. Secondly, considering that the above problems require simultaneous optimization of two objectives, an adapted hierarchical reinforcement learning idea is combined to design a target agent to guide the decision direction of the production agent and the logistics agent, guide them to achieve a good compromise between different objectives in long-term decision-making, and achieve the required multi-objective optimization.

[0187] As shown in Figure 2 , the target agent is deployed at the upper layer of the production agent and the logistics agent, and makes decisions to control the production agent and the logistics agent. For the lower-layer production agent and logistics agent, considering the sequential decision-making relationship between production scheduling and logistics scheduling in practice, a hierarchical structure is also used for deployment between them, i.e., the production agent is deployed at the upper layer of the logistics agent, and controls the logistics agent together with the target agent.

[0188] S302, according to the functions and management authorities of each agent in the multi-agent nested hierarchical framework, design the reward function, state features and action space.

[0189] The state features of the target agent include resource (including machine and logistics equipment) utilization information, task processing information (including production task processing information and logistics task processing information), and optimization preference; the state features of the production agent include machine utilization information, production task processing information, optimization preference, and target agent decision information; the state features of the logistics agent include logistics equipment utilization information, logistics task processing information, optimization preference, and target agent decision information. The state features of each agent are shown in Table 1:

[0190] Table 1 Agent state features

[0191]

[0192]

[0193] The action space of the target agent includes two possible target selections, i.e., selecting optimization to minimize maximum completion time and selecting minimum total logistics cost; the action space of the production agent includes nine combined scheduling rules for selecting the earliest pending operation of an unfinished manufacturing task and assigning it to a suitable machine for processing; the action space of the logistics agent includes three scheduling rules for assigning the derivative logistics task of the selected operation to an automated guided vehicle for processing. The action space of each agent is shown in Table 2:

[0194] Table 2 Action space of agent

[0195]

[0196] For the reward function, two reward functions corresponding to two optimization objectives are designed to evaluate the decision quality of the agent, including a reward function related to the objective of minimizing the maximum completion time and a reward function related to the objective of minimizing the total logistics cost. Both of the reward functions are composed of two parts, i.e., the optimization degree of the objective and the satisfaction degree of the manager's preference, and are invoked according to the temporary objective decision of the target agent.

[0197] The reward function related to the objective of minimizing the maximum completion time at decision point t is denoted as RMS t , and its expression is as follows:

[0198]

[0199]

[0200] where w1 and w2 are weights, nms t represents the normalized value of the change value of the objective of minimizing the maximum completion time before and after state transition, ms t represents the change value of the objective of minimizing the maximum completion time before and after state transition, pms and maxp represent the optimization preference for the objective of minimizing the maximum completion time and the maximum possible preference for an objective, respectively, ms max and ms min are the maximum value and the minimum value of the change value of the objective of minimizing the maximum completion time before and after state transition, respectively.

[0201] The reward function related to the objective of minimizing the total logistics cost at decision point t is denoted as RTC t , and its expression is as follows:

[0202]

[0203]

[0204] where ntc t represents the normalized value of the logistics cost generated by this state transition, tc t represents the logistics cost generated by this state transition, ptc and maxp represent the optimization preference for the objective of minimizing the total logistics cost and the maximum possible preference for an objective, respectively, tc max and tc min are the maximum value and the minimum value of the logistics cost generated by this state transition, respectively.

[0205] S303, the multi-agent proximal policy optimization algorithm is used to collaboratively train each agent, and the trained agent is directly used for subsequent decision making.

[0206] The specific process of model training is as follows:

[0207] A strategy network and a value network are randomly initialized for each agent in the multi-agent nested hierarchical framework, and an experience replay pool with a specific capacity is constructed for each agent, the strategy network is used to output a selected action according to the current state, and the value network is used to output a value estimate according to the state-action pair, and the structures of the two are shown in Table 3:

[0208] Table 3 Structure of strategy network and value network

[0209]

[0210]

[0211] A random environment required for single training is initialized, at each decision point in the environment, an ε-greedy strategy combined with noise is adopted, the interaction between the environment and the agent and the decision transmission between the agents are sequentially executed according to the structure of the multi-agent nested hierarchical framework, the interaction trajectory is accumulated into the corresponding experience replay pool, if the number of interaction trajectories in the experience replay pool reaches a fixed batch size, the update logic based on the proximal policy optimization algorithm is combined with the stochastic gradient descent, and the Adam optimizer is used to update the parameters of the strategy network and the value network, until the manufacturing tasks in the environment are completed. The ε-greedy strategy combined with noise exists in the training process to balance exploration and exploitation. It uses a certain possibility ε to add random noise to the action probability distribution output by the strategy network to affect the action selection of the agent. As the training proceeds, ε will gradually decrease until it reaches a minimum value.

[0212] If the manufacturing tasks in the environment are completed without reaching the maximum number of training L, the random environment required for single training is reinitialized, the above training process is repeated, and the network parameters are updated until the maximum number of training L is reached.

[0213] The parameters related to the above training process are shown in Table 4:

[0214] Table 4 Training parameters

[0215] Parameter Value Number of training iterations 1000 Target agent / production agent / logistics agent replay buffer size 32 / 32 / 500 Batch size 32 Epsilon parameter 0.9 Lower bound on epsilon 0.1 Learning rate for target agent / production agent / logistics agent 0.0001 / 0.0001 / 0.001 Discount factor 0.99 Pruning parameter 0.2

[0216] This embodiment takes an aviation parts production workshop (a typical complex flexible job shop) as an object to verify the effectiveness of the above method. The workshop is composed of 18 multifunctional machines and 6 automatic guided vehicles, and can process 7 types of workpieces with a certain process route, including compressor disc, turbine disc, fan disc, integral drum, compressor front shaft neck, compressor rear shaft neck and long shaft. The workshop layout is shown in Figure 3 , where L / U represents a loading / unloading station, Mm The machines m. The manufacturing task and fault related information in the workshop are randomly adjusted to generate different instances for training and validation, including the number of jobs to be processed is randomly set to 10 to 40, and the type of each job is also randomly allocated. The processing time of each job on each available machine is randomly set to 2 to 10 time units. The faults are randomly assigned to the automated guided vehicles. The interval time between two consecutive automated guided vehicle failures is exponentially distributed, with a mean value between 50 and 100. The repair time of each failure is randomly generated between 5 and 10. In addition, the manager's optimization preference for the two objectives is also randomly set within the corresponding range. Each instance is named as "Exp(a / b)", indicating that it involves a number of jobs and b operations.

[0217] The method of the present application is compared with the random scheduling method (i.e. the method of randomly selecting one objective, one production scheduling rule and one logistics scheduling rule at each decision point), the scheduling method based on scheduling rules (i.e. 27 kinds of composite scheduling rules composed of the actions of production agents and logistics agents) and two popular deep reinforcement learning based scheduling methods, namely the Double Deep Q network (DDQN) based scheduling method and the Deep Deterministic Policy Gradient (DDPG) based scheduling method. Table 5 and Table 6 respectively list the diversity indicators (Spacing, SP) and hypervolume measures (Hypervolume, HV) results of the Pareto solutions obtained after running each method 20 times on different instances. In order to effectively display the results, only the results of the top 3 composite scheduling rules among all scheduling rules are listed under each instance, and an additional winning rate is added to represent the overall comparison between the method of the present application and all rules, which is calculated by the number of rules that are better than or equal to the method of the present application divided by the total number of rules.

[0218] Table 5 SP index comparison results

[0219]

[0220]

[0221] Table 6 HV index comparison results

[0222]

[0223]

[0224] As can be seen from Table 5 (‘-’ represents that only one solution is obtained by this method, resulting in that the SP or winning rate cannot be measured) and Table 6, compared with the random scheduling method, the method of the application can obtain better results on the two indicators in almost all cases, which shows that the method of the application has learned an effective agent behavior strategy, rather than using a random strategy to select the feasible action at each decision point. This comparison preliminarily verifies the learning effect of the method of the application and the effectiveness of the subsequent experiments. Compared with the scheduling rule-based and other deep reinforcement learning-based scheduling methods, the effective results in Table 5 show that the method of the application obtains better or equivalent SP in most cases, indicating that the Pareto solution set obtained by the method of the application maintains better or equal distribution. Table 6 shows that the HV value obtained by the method of the application is better than that of the scheduling rule-based or deep reinforcement learning-based method in almost all cases, which shows that the approximate Pareto solution set obtained by the method of the application has better convergence and diversity than that obtained by the benchmark method. These results further confirm that the method of the application is superior to other benchmark methods in solving the proposed multi-objective problem. In addition, it can also be seen that the method of the application performs well in solving different instances, which also confirms that it has excellent versatility in dealing with different production environments. In summary, it can be concluded that the method of the application can effectively solve the production and logistics collaborative scheduling problem in different dynamic flexible job shop production environments.

[0225] If the above method is realized in the form of a software function unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product, which is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in various embodiments of the application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various program code storage media.

[0226] The above description of the embodiments is for the convenience of those skilled in the art to understand and use the application. Those skilled in the art can easily make various modifications to these embodiments, and apply the general principles described herein to other embodiments without having to go through creative labor. Therefore, the application is not limited to the above embodiments, and those skilled in the art can make improvements and modifications to the application without departing from the scope of the application.

Claims

1. A production and logistics collaborative scheduling method for dynamic flexible job shop, characterized in that, The method comprises the following steps: acquiring real-time information of a dynamic flexible job shop, inputting a production and logistics collaborative scheduling problem model for the dynamic flexible job shop, and solving to obtain a job shop scheduling scheme; The production and logistics collaborative scheduling problem model for the dynamic flexible job shop is constructed as follows: S1, setting optimization objectives of production activities and logistics activities respectively, setting task processing and resource allocation constraint conditions, and establishing a production and logistics collaborative scheduling problem model for a static flexible job shop, wherein the optimization objective of the production activities is to minimize the maximum completion time, and the optimization objective of the logistics activities is to minimize the total logistics cost; S2, according to the failure scenarios of the logistics equipment and the corresponding processing strategies, updating the optimization objective of the logistics activities, further setting the task processing and resource allocation constraint conditions under disturbance, and adding them to the production and logistics collaborative scheduling problem model for the static flexible job shop to establish a production and logistics collaborative scheduling problem model for the dynamic flexible job shop; The production and logistics collaborative scheduling problem model for the dynamic flexible job shop is solved by using a deep reinforcement learning method, and the specific process is as follows: S301, designing a multi-agent nested hierarchical framework adapted to the production and logistics collaborative scheduling problem model for the dynamic flexible job shop, wherein the multi-agent nested hierarchical framework comprises a target agent, a production agent and a logistics agent, the production agent and the logistics agent are independent of each other and are respectively used for making decisions on production scheduling and logistics scheduling, and the target agent is used for guiding the decision direction of the production agent and the logistics agent; S302, designing a reward function, a state feature and an action space according to the functions and management authorities of the agents in the multi-agent nested hierarchical framework; S303, using a multi-agent proximal policy optimization algorithm to collaboratively train the agents, inputting the real-time information of the dynamic flexible job shop into the trained agents, and obtaining a job shop scheduling scheme; In step S301, the multi-agent nested hierarchical framework is designed as follows: The target agent is deployed in the upper layer of the production agent and the logistics agent, and controls the production agent and the logistics agent; The production agent is deployed in the upper layer of the logistics agent, and controls the logistics agent together with the target agent; In step S302, the state feature of the target agent comprises resource utilization information, task processing information and optimization preference, the resource utilization information comprises machine utilization information and logistics equipment utilization information, the task processing information comprises production task processing information and logistics task processing information, and the action space of the target agent comprises selecting optimization to minimize the maximum completion time and selecting optimization to minimize the total logistics cost; The state feature of the production agent comprises machine utilization information, production task processing information, optimization preference and target agent decision information, and the action space of the production agent comprises a plurality of combined scheduling rules for selecting an earliest to-be-processed operation of an unfinished manufacturing task and allocating it to a machine for processing. The state features of the logistics agent include logistics equipment utilization information, logistics task processing information, optimization preference and target agent decision information, and the action space of the logistics agent includes a plurality of scheduling rules for assigning the derivative logistics task of the selected operation to an automated guided vehicle for processing. In step S302, the reward function includes a minimum maximum completion time target related reward function and a minimum total logistics cost related reward function, and the expression is specifically as follows: wherein, is a minimization of maximum completion time objective related reward function, and are weights, denotes a normalized value of the change in the minimization of maximum completion time before and after the state transition, denotes the change in the minimization of maximum completion time before and after the state transition, and denote an optimization preference for the minimization of maximum completion time objective and a maximum possible preference for one objective, respectively, and are maximum and minimum values of the change in the minimization of maximum completion time before and after the state transition, respectively. wherein, is a reward function associated with minimizing the total logistics cost, denotes a normalized value of the logistics cost resulting from the present state transition, denotes the logistics cost resulting from the present state transition, and denote an optimization preference for the goal of minimizing the total logistics cost and a maximum preference for one goal, respectively, and are a maximum and a minimum value of the logistics cost resulting from the present state transition, respectively. In step S303, the specific process of collaborative training of each agent by using the multi-agent proximal policy optimization algorithm is as follows: A policy network and a value network are randomly initialized for each agent in the multi-agent nested hierarchical framework, the policy network is used to output a selected action according to a current state, and the value network is used to output a value estimate according to a state-action pair. Initialize the random environment required for a single training iteration. At each decision point within this environment, use a combination of noise... - The greedy strategy sequentially executes the interaction and decision transmission between the environment and the agent, and between agents, according to the structure of the multi-agent nested hierarchical framework. It accumulates interaction trajectories into the corresponding experience replay pool. If the number of interaction trajectories in the experience replay pool reaches a fixed batch size, it updates the parameters of the policy network and the value network based on the update logic of the near-end policy optimization algorithm, combined with stochastic gradient descent, using the Adam optimizer until all manufacturing tasks in the environment are completed. If the maximum number of training is not reached after all manufacturing tasks in the environment are completed, the random environment required for single training is reinitialized, the foregoing training process is repeated, the network parameters are updated, and the maximum number of training is reached.

2. The production and logistics co-scheduling method for dynamic flexible job shop according to claim 1, characterized in that, In step S1, the expression of the optimization target of the production activity is specifically as follows: wherein, denotes the maximum completion time, denotes a manufacturing task an operation of the manufacturing task, denotes the total number of manufacturing tasks, denotes the total number of operations required to complete the manufacturing task . The expression of the optimization target of the logistics activity is specifically as follows: wherein denotes the total logistics costs, denotes the commissioning costs required for activating an automated guided vehicle, denotes the logistics costs required per unit of transport time of an automated guided vehicle, is a binary variable, which is 1 if and only if the automated guided vehicle is activated in the dispatching scheme, denotes the total number of automated guided vehicles, denote the start time and the end time of the non-load phase in the transport task derived from the operation of the manufacturing task , respectively, and denote the start time and the end time of the load phase in the transport task derived from the operation of the manufacturing task , respectively.

3. The production and logistics co-scheduling method for dynamic flexible job shop according to claim 2, characterized in that, The task processing and resource allocation constraint conditions include: The operation completion time of each manufacturing task is the sum of the start time and the machine processing time, and the expression is specifically as follows: in, Indicates the total number of machines. Indicates manufacturing task Operation The start time, Indicates manufacturing task Operation In the machine On the processing time, It is a binary variable if and only if the manufacturing task Operation Assigned to machine The value is 1 during processing; Each operation of each manufacturing task can only start processing after the task is transported to the corresponding machine, and the expression is specifically as follows: Each machine device can only process one operation at a time, and the expression is specifically as follows: in, It is a binary variable if and only if the manufacturing task Operation In the machine Immediately following the manufacturing task Operation When processed, it is 1. Indicates manufacturing task Operation The completion time, It is a binary variable if and only if the manufacturing task Operation In the machine The value is 1 when it is the first one processed. Each automated guided vehicle can only transport one manufacturing task at a time, and the expression is specifically as follows: in, It is a binary variable if and only if the manufacturing task Operation The corresponding transportation task is carried out in the automated guided vehicle. Immediately following the manufacturing task Operation The value is 1 when the corresponding transportation task is processed. It is a binary variable if and only if the manufacturing task Operation The value is 1 when the corresponding transport task is the first one processed on the assigned automated guided vehicle; Each automated guided vehicle has sufficient time to complete the non-load phase of the transportation task without being disturbed, and the expression is specifically as follows: in, Indicates from position arrive The delivery time It is a binary variable if and only if the manufacturing task Operation The corresponding transportation task is handled by an automated guided vehicle. When the completion value is 1, A variable is a binary variable if and only if it represents a machine... Located in The value is 1 when it is above. It is a binary variable if and only if the manufacturing task Operation Assigned to machine The value during processing is 1. It is a binary variable if and only if the manufacturing task Operation Assigned to machine The value is 1 during processing; Manufacturing tasks Operation of the manufacturing tasks The load phase of a transport task can only start after the previous operation of the operation is finished and the assigned automated guided vehicle has completed its non-load phase, expressed as follows: wherein, indicates a manufacturing task operation derives the start time of the load phase in the transport task; Each automated guided vehicle has sufficient time to complete the load phase of the transportation task without being disturbed, and the expression is specifically as follows: wherein, indicates a manufacturing task operation derives an end time of a load phase of a transport task; Each operation of each manufacturing task can only be assigned to one machine for completion, and the expression is specifically as follows: The transportation task derived from each operation of each manufacturing task can only be assigned to one automated guided vehicle for completion, and the expression is specifically as follows: Any two operations can only have a tight front and back relationship on one processing machine, and the expression is specifically as follows: Any two operations can only have a tight front and back relationship on one automated guided vehicle, and the expression is specifically as follows: 。 4. The method according to claim 1, wherein, In step S2, the logistics equipment is taken as the object of the automated guided vehicle, and the failure scenarios of the logistics equipment and the corresponding processing strategies are as follows: If the automated guided vehicle fails when it is idle, it is repaired at the current location to restore its availability; If the automated guided vehicle fails in the non-load phase, it is moved to the location closer to the failure location among the start location and the destination location for repair; If the automated guided vehicle fails in the load phase, it is moved to the location closer to the failure location among the start location and the destination location for repair, and the loaded task is moved synchronously and rescheduled.

5. The production and logistics co-scheduling method for dynamic flexible job shop according to claim 4, characterized in that, The expression of the updated optimization target of the logistics activity is specifically as follows: wherein, denotes the total logistics cost, denotes the commissioning cost required to activate one automated guided vehicle, denotes the logistics cost required per unit of transport time of an automated guided vehicle, is a binary variable, which is 1 if and only if the automated guided vehicle is activated in the dispatching scheme, denotes the total number of automated guided vehicles, denote the start time and the end time of the non-load phase in the transport task derived from the operation of the manufacturing task respectively, and denote the start time and the end time of the load phase in the transport task derived from the operation of the manufacturing task respectively, denotes the transport time consumed by the logistics task being processed by the malfunctioning vehicle when a malfunction occurs, denotes the total number of malfunctions.

6. The production and logistics co-scheduling method for dynamic flexible job shop according to claim 5, characterized in that, The disturbance task processing and resource allocation constraints include: The automated guided vehicle after a failure can only start processing a new logistics task after being repaired, and the expression is as follows: wherein, represents a malfunction occurred, represents a repair of the malfunction required, is a binary variable that is 1 if and only if the assignment of the operation deriving the logistics task experienced a malfunction before handling the task, is a binary variable that is 1 if and only if the assignment of the operation deriving the logistics task is the task that is first handled by the assigned automated guided vehicle after repairing the malfunction, is a binary variable that is 1 if and only if the assignment of the operation deriving the logistics task is the task that is first handled by the assigned automated guided vehicle after repairing the malfunction ; The interrupted transport task affected by the failure can be immediately reprocessed after the interruption, and the expression is as follows: wherein, is a binary variable that is 1 if and only if the manufacturing task is operated derives from a logistics task that is interrupted due to a failure ; Each automated guided vehicle still satisfies the transport-related constraints in the task processing and resource allocation constraints when not directly affected by the disturbance; Each automated guided vehicle departs from the repair location to process the first logistics task after being repaired, and the expression is as follows: wherein denotes a manufacturing task operation deriving a load phase of a logistic task is interrupted by a fault, denotes a transport time from a location to , denotes a repair location of a fault autonomous guided vehicle after a fault occurrence, is a binary variable that is 1 if and only if a manufacturing task operation is assigned to a machine handling, is a binary variable that is 1 if and only if a machine is located at a location ; For a logistics task interrupted during the loading process, each automated guided vehicle will go to the corresponding repair location to complete its non-loading phase task, and the expression is as follows: wherein is a binary variable that is 1 if and only if the operation of manufacturing task corresponds to a transport task that is processed immediately after the manufacturing task is a binary variable that is 1 if and only if the operation of manufacturing task corresponds to a transport task that is processed first on the assigned automated guided vehicle, is a binary variable that is 1 if and only if the operation of manufacturing task corresponds to a transport task that is processed by the automated guided vehicle, is a binary variable that is 1 if and only if the operation of manufacturing task corresponds to a transport task that is processed first on the assigned automated guided vehicle, is a binary variable that is 1 if and only if the operation of manufacturing task corresponds to a transport task that is processed by the automated guided vehicle, is a binary variable that is 1 if and only if the operation of manufacturing task corresponds to a transport task that is processed by the automated guided vehicle, is a binary variable that is 1 if and only if the operation of manufacturing task For a logistics task interrupted due to a failure in the loading phase, each automated guided vehicle departs from the corresponding repair location to complete the loading phase task, and the expression is as follows: 。

Citation Information

Patent Citations

  • Flexible job shop intelligent scheduling decision-making method combined with transportation equipment constraint

    CN112949077A

  • Flexible job shop scheduling method considering transportation time and adjustment time

    CN116109084A

  • Distributed flexible job shop scheduling method considering limited transportation resources and multiple targets

    CN116187093A

  • Intelligent workshop production scheduling method, electronic equipment and scheduling system

    CN116934038A

  • Complete vehicle manufacturing stamping resource scheduling method based on deep reinforcement learning

    CN117557016A