Mathematical model and reinforcement learning integrated inconsistent batch-batch flow flexible job shop scheduling method and system, and storage medium

Through the integration of mathematical models and reinforcement learning methods, batch division and machine selection are optimized, and the problems of low efficiency and resource utilization in inconsistent batch flexible operation workshop scheduling problems are solved, achieving high-quality scheduling and improvement of production efficiency.

CN119940801AActive Publication Date: 2025-05-06HUAZHONG UNIV OF SCI & TECH

Patent Information

Application Number
CN202411981879.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-06
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

The prior art is difficult to quickly solve the problem of high-quality flexible operation workshop scheduling in the case of inconsistent batching, resulting in low processing efficiency, insufficient resource utilization and difficulty in production organization management.

Method used

Using integrated mathematical model and reinforcement learning methods, we use flexible work workshop scheduling problems as heterogeneous graphs, and design action space, reward function and decision network, and combine near-end strategy optimization algorithm for reinforcement learning training to optimize batch division and machine selection.

Benefits of technology

Quickly solve high-quality workshop scheduling methods in batches in a consistent manner, which improves processing efficiency and resource utilization, and reduces the difficulties in production organization management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940801A_ABST
    Figure CN119940801A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of workshop scheduling, and particularly discloses an inconsistent batch-batch flow flexible job workshop scheduling method and system integrating a mathematical model and reinforcement learning, and a storage medium, and the method comprises the steps: firstly introducing strict workshop state representation which comprises decision-related process, machine, batch and other information; extracting state feature information by using a double-attention network; secondly, designing a scheduling decision framework based on a near-section strategy optimization algorithm, solving a process sorting sub-problem and obtaining a batch division adaptive coefficient; and finally, considering the processing preparation time, constructing a process-level batch machine collaborative optimization mixed integer linear programming model, and realizing comprehensive optimization of a batch division sub-problem and a machine selection sub-problem. The scheduling method provided by the invention has remarkable advantages in solving the problem of batch flow flexible job shop scheduling, and a high-quality shop scheduling method can be obtained by quickly solving under the condition of inconsistent batches.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of workshop scheduling, and more specifically, relates to an inconsistent batch flow flexible job shop scheduling method, system and storage medium integrating mathematical model and reinforcement learning. Background Art

[0002] As one of the cores of intelligent manufacturing, intelligent optimization decision-making drives the execution of production activities and has an important impact on the production efficiency of enterprises. Production scheduling is to allocate limited resources to different tasks within a reasonable time under certain conditions to meet the daily goals of optimization decision-making. It is a key core technology for production planning and control of manufacturing enterprises. With the widespread application of multifunctional equipment such as CNC machine tools and machining centers and the rise of personalized and customized production models, Flexible Job Shop Scheduling Problems (FJSP) have emerged.

[0003] In traditional flexible job shop scheduling, if a workpiece has not completed the processing of this step, it cannot be transferred to the next machine for the next stage of processing. However, in manufacturing enterprises, the actual production is order-oriented, and orders often contain a number of the same workpieces. If the entire order is processed according to one workpiece, the processing capacity of the equipment in the flexible job shop cannot be fully utilized, resulting in a long equipment idle time; but if all the workpieces in an order are split into separate workpieces for scheduling and production, the expansion of the problem scale poses a severe challenge to the scheduling algorithm, which may lead to the scheduling algorithm being unable to find a better solution. What is more serious is that in actual production, the switching time is increased due to the frequent switching of different workpieces, the transportation cost is increased, the processing efficiency is reduced, the processing quality is difficult to guarantee, and the production organization and management of various aspects are difficult (including material preparation, workshop logistics, delivery management, etc.). Therefore, the usual way to deal with it is to split the order into sub-batches for processing, that is, the batch flow production method. Batch flow is classified according to two criteria: (1) Based on whether the batch size of each sub-batch of the same type of workpiece is the same, it can be divided into equal batching and unequal batching; (2) Based on whether the batching strategies between different processes of the same type of workpiece are the same, it can be divided into consistent batching and inconsistent batching.

[0004] Batch flow flexible job shops are very common in real production. At present, the research on this problem mainly focuses on the case of consistent batching. This type of method has the characteristics of insufficient batch flexibility and low machine utilization. There are fewer studies on the case of inconsistent batching. Zhang Chunjiang, Fan Jiaxin and others fixed the machine allocation of each sub-batch in the elite solution and the value of the processing order variable of each sub-batch on the machine, and constructed a mixed integer linear programming model with the sub-batch size as one of the decision variables. The mixed integer linear programming model was iteratively solved using an intelligent optimization algorithm. Although fixing some variables in the elite solution can accelerate local optimization, it may limit the search space and affect global optimality. In addition, the intelligent optimization algorithm based on encoding and decoding requires a long time to iterate to find a satisfactory solution, especially when the problem scale increases, its time complexity will increase significantly.

[0005] Reinforcement learning (RL) allows the agent to learn strategies through interaction with the environment. When the state changes, reinforcement learning can quickly generate decisions through the trained strategy, which is more real-time and adaptable. With the development of computer science, deep neural networks have also been introduced into the field of reinforcement learning to deal with continuous state space and action space, namely deep reinforcement learning (DRL). By introducing neural networks, not only the challenges of continuous state space and action space are effectively handled, but also the generalization ability and learning effect of the model are significantly improved, and more flexible and intelligent decision-making capabilities are added. Li Junqing et al. optimized the flexible job shop scheduling problem of batch flow based on the deep reinforcement learning framework of graph neural network. The number of materials in each sub-batch in all batches remains unchanged. Their invention limits the adjustment ability of sub-batches in state representation and action decision, which will lead to insufficient resource utilization of some machines, and there are defects in flexibility and resource utilization. Therefore, it is of great significance to design a fast solution to high-quality solutions in the flexible job shop with inconsistent batch flow. Summary of the invention

[0006] In response to the above defects or improvement needs of the prior art, the present invention provides a flexible job shop scheduling method, system and storage medium for inconsistent batch flow that integrates mathematical models and reinforcement learning. The purpose is to quickly solve a high-quality shop scheduling method under inconsistent batch conditions and achieve improved shop processing efficiency.

[0007] To achieve the above object, according to a first aspect of the present invention, a method for inconsistent batch flow flexible job shop scheduling integrating mathematical model and reinforcement learning is proposed, comprising the following steps:

[0008] The flexible job shop scheduling problem is represented as a heterogeneous graph, in which processes and machines are different types of nodes, and process-machine processing relationships are edges; the flexible job shop scheduling problem is transformed into a Markov decision process, and the action space, reward function and decision network are designed; the action space is a combination of the process to be processed and the batch partition coefficient γ, and the batch partition coefficient γ is used to allocate the process to be processed to the machine with lower processing time cost; the decision network is constructed based on the actor-critic network;

[0009] A mixed integer linear programming model for process scheduling is established, and the objective function of the mixed integer linear programming model is constructed based on the batch partition coefficient;

[0010] Use the proximal policy optimization algorithm to perform reinforcement learning training in a simulated environment, including:

[0011] S1, extracting environmental state features based on the current heterogeneous graph;

[0012] S2, the actor network makes decisions based on the characteristics of the environment state and determines the current processing steps and batch division coefficients;

[0013] S3. Solve the mixed integer linear programming model for the current process to be processed to obtain a scheduling solution, i.e., batch division and machine selection for the process to be processed;

[0014] S4, update the heterogeneous graph based on the scheduling scheme and optimize the decision network parameters according to the reward function;

[0015] S5. Repeat S1 to S4 until all decisions on pending processes are completed, and a flexible job shop scheduling method is obtained.

[0016] As a further preferred embodiment, the action space A(t) is expressed as:

[0017] A(t)={(O i,j ,γ)|O i,j ∈J c (t),γ∈Γ}

[0018] Among them, O i,j represents the jth process of the i-th type of operation, i.e., the process to be processed, J c (t) represents the set of candidate processes at time t; Γ represents the discretized set of batch partition coefficients γ, Γ={γ1,γ2,…,γ l}, v=1,2,…,l, where l represents the number of discretizations, γ v ∈[0,1].

[0019] As a further preference, the mixed integer linear programming model uses batch partitioning variables, machine selection variables and completion time variables in coordination, and takes minimizing maximum completion time and processing preparation time as the optimization goal.

[0020] As a further preferred embodiment, for the j-th process O of the i-th type of operation i,j , the objective function of the mixed integer linear programming model is:

[0021]

[0022] Among them, C i,j Indicates O i,j The maximum completion time of the sub-batch, B i,j Indicates O i,j Total number of sub-batches, S i,j,k Indicates O i,j In M i,j,k Preparation time for processing, x i,j,k Indicates O i,j Is it in M i,j,k Upper processing, M i,j,k Indicates O i,j The kth processable machine.

[0023] As a further preferred embodiment, when the proximal strategy optimization algorithm is used for reinforcement learning training, minimizing the total completion time of the workpiece is used as the objective function.

[0024] As a further preferred embodiment, the reward function r t It is expressed as:

[0025]

[0026] Among them, C((O i,j ,γ),s t )、C((O i,j ,γ),s t+1 ) represent process O respectively i,j The estimated completion time at time step t and time step t+1, s t 、s t+1 Represent the states at time step t and time step t+1 respectively.

[0027] As a further preferred embodiment, the environmental state characteristics include process characteristics Machine Features Sub-batch characteristics Process Machine Pair Features and the global feature h G .

[0028] As a further preferred embodiment, the environmental state characteristics are determined in the following manner:

[0029] Based on the current heterogeneous graph, obtain the initial process characteristics Initial Machine Characteristics Sub-batch characteristics and process machine pair features

[0030] Initial Operation Features Initial Machine Characteristics Input graph attention network to get process characteristics Machine Features

[0031] Operation Features Machine Features and sub-batch characteristics After average pooling, the global feature h is obtained. G .

[0032] According to a second aspect of the present invention, a flexible job shop scheduling system for inconsistent batch flow integrating mathematical models and reinforcement learning is provided, comprising a processor for executing the above-mentioned flexible job shop scheduling method for inconsistent batch flow integrating mathematical models and reinforcement learning.

[0033] According to a third aspect of the present invention, there is provided a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned inconsistent batch flow flexible job shop scheduling method integrating mathematical model and reinforcement learning.

[0034] In general, the above technical solution conceived by the present invention has the following technical advantages compared with the prior art:

[0035] The present invention introduces a batch partition coefficient, and based on the scheduling decision framework of the near-stage strategy optimization algorithm, solves the process sequencing sub-problem and obtains the batch partition adaptive coefficient, and then combines the mixed integer linear programming model to achieve comprehensive optimization of the batch partition sub-problem and the machine selection sub-problem. The method of the present invention has significant advantages in solving the batch flow flexible job shop scheduling problem. It can quickly solve the high-quality shop scheduling method in the case of inconsistent batching, and realize the improvement of shop processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 A flow chart of a method for flexible job shop scheduling of inconsistent batch flows integrating mathematical models and reinforcement learning according to an embodiment of the present invention;

[0037] Figure 2 This is a schematic diagram of feature extraction of attention network according to an embodiment of the present invention;

[0038] Figure 3Schematic diagram of the action space of the present invention. DETAILED DESCRIPTION

[0039] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0040] An embodiment of the present invention provides an inconsistent batch flow flexible job shop scheduling method that integrates mathematical models and reinforcement learning.

[0041] For the inconsistent batch flow flexible job shop scheduling problem: In the shop, there are n types of jobs J = {J1, J2, …, J n} and m machines M={M1,M2,…,M m}, the i-th job J i The Chinese Communist Party has D i The same job, J i A single job contains N i Process To be processed in a given sequence, a sub-batch is the smallest unit of processing.

[0042] Each process O i,j Divided into B i,j Sub-batch, B i,j,k Indicates process O i,j The kth sub-batch of O. i,j The set of processable machines is expressed as x i,j,k Indicates O i,j Is it in M i,j, Upper processing, O i,j In M i,j,k The sub-batch size for processing is denoted as σ i,j,k , O i,j In M i,j,k The sub-batch processed on i,j,k Indicates that O i,j In M i,j,k The processing time is represented by P i,j,k , O i,j In M i,j,k The earliest available time for processing is indicated as ST i,j,k , M i,j,k The preparation time is denoted as S i,j,k , c i,j,k Indicates O i,j In M i,j,k The processing time, C ijIndicates O i,j complete processing time.

[0043] The scheduling problem in this embodiment has the following characteristics:

[0044] 1) All devices are available at time 0;

[0045] 2) All tasks can be started at time 0;

[0046] 3) The processing order of the processing tasks cannot be changed, and there is no sequence constraint on the processes of different sub-batches;

[0047] 4) The workpieces in each sub-batch of the task process are processed continuously without interruption;

[0048] 5) A sub-batch can only be transferred to the next process after all the workpieces have completed the processing of one process;

[0049] 6) A sub-batch of a certain process of a task can come from multiple sub-batches in the previous process. In this case, the processing of this sub-batch can only start after the multiple sub-batches of the previous process are processed.

[0050] 7) The device buffer is infinite and the transit time is ignored;

[0051] 8) The same equipment can only process one process of workpieces in a certain sub-batch at the same time;

[0052] 9) A workpiece in the same sub-batch can only be processed by one device in one process at the same time.

[0053] In order to realize the inconsistent batch flow flexible job shop scheduling method integrating mathematical model and reinforcement learning, the present invention is designed as follows:

[0054] 1. Introducing state representation and extracting state feature information based on graph attention network

[0055] The state space includes information such as process, machine, batch, and equipment availability, which is used to describe the current scheduling status of the workshop. Based on the process-machine structure of the scheduling problem, a heterogeneous graph is constructed, with processes and machines as different types of nodes and process-machine processing relationships as edges. Edge weights represent process-machine information (such as processing time, etc.). A graph attention network (GAT) is used to capture the interaction between nodes and extract global state information. GAT aggregates the features of adjacent nodes through the attention mechanism to generate high-dimensional feature representations of each node for subsequent process sorting decisions.

[0056] 1. Status indication

[0057] The design of the state space in the present invention comprehensively covers the dynamic information of the workshop scheduling, including the elements such as the process, workpiece, machine and batch.

[0058] The process characteristics describe O i,j The status of a job, including its progress, remaining workload and other information, including: whether the process has been scheduled, the lower bound of the earliest completion time of the process, the span of the processing time, the number of remaining processes in the job, the waiting time of the process, the remaining workload of the process, the number of remaining processes in the job, the remaining workload in the job, and the number of available machines.

[0059] Machine characteristics describe the status and resource conditions of the machine, mainly including: the number of jobs available to the current machine, the number of processes available to the current machine, the minimum processing time of the machine, the remaining workload of the machine, the idle time of the machine, and whether the machine is working.

[0060] The sub-batch characteristics describe the properties of each sub-batch, mainly including: the lower bound of the earliest completion time of the sub-batch, the average processing time of the sub-batch, the span of the sub-batch processing time, the maximum processing time of the sub-batch, the minimum processing time of the sub-batch, the average preparation time of the sub-batch, the maximum preparation time of the sub-batch, the minimum preparation time of the sub-batch, and the preparation time span of the sub-batch.

[0061] The process-machine pair feature represents the relationship between each process and the machine, mainly including: the processing time of the current candidate subbatch, the maximum subbatch processing time, the maximum processing time of the current candidate subbatch, the maximum preparation time of the current candidate subbatch, the minimum preparation time of the current candidate subbatch, and the preparation time span of the current candidate subbatch.

[0062] The above four types of features (process features, machine features, sub-batch features, and process-machine pair features) comprehensively describe the status of the process in job scheduling, the availability of the machine, the matching of the process and the machine, and the processing characteristics of the sub-batch. By effectively combining and normalizing these features, the model can better understand and predict the scheduling behavior of the system, optimize scheduling decisions, and improve resource utilization and job completion efficiency. In addition, feature normalization helps to avoid the impact of scale differences on model training, thereby improving the stability and generalization ability of the model.

[0063] 2. Construction of Heterogeneous Graph of Flexible Job Shop

[0064] A heterogeneous graph structure is used to model the processes and machines in the flexible job shop scheduling problem (FJSP), and the processes and machines are treated as different types of nodes. The process node represents the specific process of each job, including its process number, demand, processing time and other information; the machine node contains the machine number, the remaining workload of the machine, whether the machine is working and other information.

[0065] In a heterogeneous graph, the characteristics of the edge are used to represent the relationship between the process and the machine, and the weight of the edge usually reflects the processing time of the process on a specified machine or the matching correlation between the two. In this way, a clear basis can be provided for the optimization of the batch partitioning coefficient. Specifically, when the edge weight is small, it means that the match between the machine and the process is high and the processing time is short, which helps to determine the optimal batch partitioning coefficient and ensure that the processing of sub-batches can be more efficiently allocated to resources. When the edge weight is large, the match between the process and the machine is low or the processing time is long, suggesting that the machine should be avoided for processing during the batch partitioning process.

[0066] Edge types can be divided into two categories: connecting edges and scheduling edges. Connecting edges are used to represent the potential processing relationship between processes and machines, while scheduling edges represent the order in which a process is actually scheduled and executed on a specific machine. The update of scheduling edges reflects the process of job scheduling, where the scheduled processes will reflect the execution order of the processes by updating the direction of the edges, thereby converting the graph structure into a directed acyclic graph (DAG) to ensure that no circular dependencies occur during the scheduling process. With such a graph structure, the scheduling of the job shop can be flexibly managed and optimized, taking into account factors such as machine availability, process dependencies, and job requirements.

[0067] 3. Feature extraction

[0068] Graph Attention Networks (GAT) dynamically extracts the interaction features between process nodes and machine nodes through the attention mechanism, focusing on capturing the complex dynamic interaction relationship between process and machine, and designs a dual attention network, namely the process attention block and the machine adaptation attention block to process the specific features of the process and machine respectively. Figure 2 As shown, the specific implementation is as follows:

[0069] (1) Model input

[0070] Initial process characteristics Each process node The input features are expressed as Including processing time, priority and other information;

[0071] Initial Machine Characteristics Each machine node The input features are expressed as Including information such as current load, idle time, etc.

[0072] Sub-batch characteristics Each B i,j,k The sub-batch characteristics are recorded as in Including information such as sub-batch processing time and preparation time.

[0073] Process Machine Pair Features Including information such as the processing time of the current candidate sub-batch and the preparation time span of the candidate sub-batch.

[0074] (2) Process attention block

[0075] This module captures the characteristics of processes within the same job through the sequence relationship between processes and models their priorities.

[0076] Attention coefficient calculation: for process O i,j Input features and its neighbor nodes O i,j-1 and O i,j+1 , calculate the attention coefficient:

[0077]

[0078]

[0079] in is a linear transformation, is the attention weight.

[0080] Dynamic masking and normalization: If a node O i,j-1 or i,j+1 If missing, the corresponding e i,j-1 or e i,j+1 Set the mask. Perform Softmax normalization on the attention score to get the attention coefficient α i,j-1 , α i,j+1 .

[0081] Feature update: using the normalized attention coefficient α i,j-1 , α i,j+1 The neighbor features are weighted and summed, and the output features are updated through the linear activation function LeakyReLU:

[0082]

[0083] If a neighbor node is missing, its corresponding attention score and feature weighting item are directly set to 0. In this way, the attention scores of predecessors and successors and their impact on the features of the current node can be directly calculated. This method is more compact while maintaining effective modeling of the relationship between nodes.

[0084] (3) Machine Adaptability Attention Block

[0085] The machine fitness attention module aims to model the adaptation relationship and dynamic competition between machines, thereby providing support for generating high-quality machine features. In production scheduling, different machines may have different machinability, current status, and processing efficiency for unscheduled processes. Therefore, the machine fitness attention block captures the dynamic adaptability between machines by explicitly modeling this competitive relationship.

[0086] Definition of fitness: Assume C k,q Indicates machine M k and M q The adaptation set between them contains the unscheduled processes that can be processed by the two machines together; c k,q Represents the adaptability weight between them, which is defined as follows: Among them, J C represents the set of currently unscheduled processes. i,j ∈C k,q When it is empty, c k,q is 0.

[0087] Attention coefficient calculation: for each machine M k Computing and other machines M q The attention weight u between k,q , and its calculation formula is: in, It's Machine M k and M q The characteristics of c k,q are their fitness weights, and is the linear transformation matrix, is the weight vector.

[0088] Normalization and feature update: The attention weight is normalized by softmax to β k,q , indicating that machine M k For other machines M q Adaptability, machine M k The updated features of are obtained by weighted summation: Among them, N k It is with M k A collection of machines that have an adaptation relationship.

[0089] (4) Global features

[0090] Process O i,j 、Machine M k The initial features are The features of the process and machine are updated after passing through the L-layer attention network as right AveragingPooling is applied respectively, and the results of pooling are respectively represented These global features are then concatenated to form a complete global feature of the FJSPLS example. The specific expression is as follows:

[0091]

[0092] 2. Process Sequencing Decision Based on Proximal Strategy Optimization

[0093] The present invention designs a process sequencing decision network based on proximal policy optimization to optimize the batch flow flexible job shop scheduling problem (FJSPLS). A deep reinforcement learning framework based on proximal policy optimization (PPO) is adopted. This framework utilizes the size-independence characteristics of the attention model and combines the proximal policy optimization algorithm to achieve efficient decision-making. Specifically, the current processing process is selected according to the feature information reinforcement learning strategy, and an adaptive batch partitioning coefficient is generated. The objective function is to minimize the total completion time of the workpiece. Reinforcement learning training is performed in a simulation environment, and the policy network is continuously optimized through the feedback loop of state-action rewards. After training, the policy model can be deployed for real-time decision scheduling.

[0094] Specifically, Figure 3 As shown in Figure 2, the flexible job shop scheduling problem is transformed into a Markov decision process:

[0095] 1. Action Space

[0096] In a batch flow flexible job shop, the processing time of different sub-batches on different machines may be different. How to dynamically adjust the batch strategy to ensure more efficient task processing on different machines and avoid resource waste? The introduction of the batch partition coefficient helps to optimize the overall production scheduling. The batch partition coefficient γ represents the ratio of the weight of the process completion time to the weight of the process preparation time, which is used to allocate the processes to be processed to machines with lower processing time costs.

[0097] Specifically, the discretization set of batch partition coefficients γ is Γ = {γ1,γ2,…,γ l}, v=1,2,…,l, where l represents the number of discretizations, γ v ∈[0,1]. The larger the discretization number l is, the higher the batch refinement level is, but the corresponding reinforcement learning network convergence difficulty is also greater. Therefore, the value of l is determined according to the actual situation and specific needs.

[0098] The action space A(t) is designed as a combination of the processing steps to be processed and the batch division coefficient, expressed as:

[0099] A(t)={(Oi,j ,γ)|O i,j ∈J c (t),γ∈Γ}

[0100] Among them, J c (t) is the set of candidate processes at time t.

[0101] 2. State Transfer

[0102] At each time step t, an action a is performed t , which corresponds to selecting a process O to be processed i,j and the corresponding batch partition coefficient γ for processing. t , the system will update the status s t The process set in Machine Collection Sub-batch collection and action set A t , and get the new state s according to the new production state t+1 .

[0103] 3. Reward design

[0104] Reward t It should be designed to guide the agent to choose the maximum completion time (i.e., makespan, denoted by C) that helps to reduce all task processes. max ) action. Estimating process O through heuristic rules i,j The completion time C((O i,j ,γ),s t ). Reward t It is defined as the difference between the estimated completion time at step t and t+1, and is standardized and can be expressed as:

[0105]

[0106] 4. Strategy

[0107] The present invention adopts a random strategy π(a t |s t ), whose distribution is generated by a deep reinforcement learning algorithm (in this case, an actor-critic network) and has trainable parameters θ. t In the case of , the distribution will return to select each action a t ∈A(t).

[0108] 5. Decision Network

[0109] The present invention designs a decision network based on actor-critic, maintains the size-independent property of the attention model, and uses two multi-layer perceptrons (MLPs) as actors and critics, with parameters θ and

[0110] The actor network aims to generate a random policy π in two steps θ (a t |s t ): First, for each t ∈A(t) related features are concatenated. In order to reduce the size of the action space, the features are selected in the candidate area according to the "remaining maximum workload" (Feature Selection, FS), and the selected feature vector is input into the MLP θ In the equation, we get the vector μ(a t |s t ), expressed as: The softmax function is then used to output the desired distribution, expressed as: The critic network is based on the global feature h G As input, generate a vector As an estimate of the state value.

[0111] MILP model for process-level batch machine collaborative optimization

[0112] A mixed integer linear programming (MILP) model for process-level batch-machine collaborative optimization is constructed to solve the batch partitioning and machine selection sub-problems at the same time. The model takes minimizing the maximum completion time and processing preparation time (the line change preparation time needs to be considered when processing different workpieces on the equipment) as the optimization goal, and adopts batch partitioning variables, machine selection variables, and completion time variables for collaborative modeling. Constraints are used to ensure the logical consistency of processing requirements, batch partitioning, and machine selection. The model achieves comprehensive optimization of batch partitioning sub-problems and machine selection sub-problems through dynamic batching and flexible allocation strategies, thereby improving workshop processing efficiency and equipment utilization.

[0113] Specifically, for O i,j The MILP model for process-level batch-machine collaborative optimization is established, which can optimize batch partitioning (number of sub-batches, sub-batch size), machine allocation and processing time under given job and machine constraints, taking into account O i,j The batch adjustment coefficient in the model allows the optimization target to be flexibly adjusted according to actual needs.

[0114] 1. Variable definition

[0115] Table 1 Variable definitions

[0116]

[0117] 2. Objective Function

[0118] The introduction of batch partitioning coefficients can effectively reduce the idleness and uneven load of machine resources by affecting the number of sub-batches, and reduce the additional preparation time costs, such as machine adjustment and task switching. The introduction of batch partitioning coefficients makes the scheduling framework more flexible, and can dynamically adjust the batching strategy according to different production requirements and resource constraints to ensure the optimal production scheduling decision. Therefore, the objective function is expressed as:

[0119]

[0120] 3. Constraints

[0121] The total amount of each sub-batch is equal to J i Total number of jobs:

[0122]

[0123] Sub-batch completion time constraints:

[0124]

[0125] Batch partitioning is only available when the machine is selected:

[0126]

[0127] Process O i,j The maximum completion time C i,j Greater than the completion time of any of its subbatches:

[0128]

[0129] 4. Solution process

[0130] When solving the process-level batch machine collaborative optimization problem, a mixed integer linear programming (MILP) model is constructed and solved using the Gurobi optimizer. Through optimization, the model calculates the number of workpieces σ assigned to each machine. i,j,k and the processing time σ of each machine i,j,k ·P i,j,k These two optimization results are then passed to the environment, which updates its state based on this information.

[0131] 4. Scheduling Method

[0132] Based on the designs in the above one to three, when the embodiment of the present invention performs inconsistent batch batch flow flexible job shop scheduling, such as Figure 1 As shown, the following steps are included:

[0133] Use the proximal policy optimization algorithm to perform reinforcement learning training in a simulated environment:

[0134] S1, extracting environmental state features based on the current heterogeneous graph;

[0135] S2, the actor network makes decisions based on the characteristics of the environment state and determines the current processing steps and batch division coefficients;

[0136] S3. Solve the mixed integer linear programming model for the current process to be processed to obtain a scheduling solution, i.e., batch division and machine selection for the process to be processed;

[0137] S4, update the heterogeneous graph based on the scheduling scheme and optimize the decision network parameters according to the reward function;

[0138] S5. Repeat S1 to S4 until all decisions on pending processes are completed, and a flexible job shop scheduling method is obtained.

[0139] 5. Numerical Experiments and Analysis

[0140] The effectiveness and advancement of the proposed method are verified through ablation experiments and numerical comparisons with multiple baseline algorithms. Several popular baseline algorithms are selected, including the traditional process sorting PDR (Priority Dispatching Rules), the exact Gurobi algorithm, and the genetic algorithm using advanced batching strategy. The main goal of the experiment is to evaluate the advantages of the proposed algorithm in scheduling performance, generalization ability, and computational efficiency.

[0141] 1. Test Case

[0142] The example is derived from the extension of the classic Flexible Job Shop Scheduling Problem (FJSP) example. The benchmark example is set up as follows:

[0143] Number of jobs D i ∈[200,400], machine initial setup time S i,j ∈[50,100], machine setup time S i,j,i′,j′ ∈[100,200], processing time P (i,j,k) ∈[3,7].

[0144] The scales of the examples are specifically 10×5, 15×5, 15×10, 20×5, 20×10, 30×10, 40×20, and 60×20. The present invention randomly generates training examples from benchmark examples of corresponding scales, and test examples and verification examples are pre-generated from each benchmark example.

[0145] 2. Training Configuration

[0146] The configuration of the training process is as follows: training cycle (Nep) = 3000, training set size |Etr| = 5, reset frequency (Nr) = 3, validation set size (Nval) = 20, and weight factor dimension n = 21.

[0147] When the candidate job size is not greater than 20, the job size is selected as 10, otherwise it is 20. The output dimension of the first layer of the attention network is set to 64, and the MLP θ and All are set to 3 layers, and the hidden layer dimension is 64.

[0148] For the parameter configuration based on genetic algorithm: termination condition: number of workpiece types * number of machines * 10s, population size: 600, crossover probability: 0.9, mutation probability: 0.2, number of elite individuals: 6, scale of tournament selection: 2.

[0149] For the parameter configuration of the PPO algorithm: strategy coefficient α policy =1, value coefficient α value =0.5, entropy coefficient α entropy = 0.01, pruning parameter ∈ = 0.2, GAE parameter λ = 0.95, discount factor γ = 0.99, learning rate lr = 5 × 10 4 .

[0150] All experiments were performed on a machine equipped with an Intel(R) Core(TM) i5-12600KF CPU and an NVIDIA GeForce RTX4060Ti 16G GPU.

[0151] 3. Baseline Algorithms and Performance Indicators

[0152] In order to compare the performance of the proposed algorithm, four PDRs (Priority Dispatching Rules) for process sorting and three PDRs for batch partitioning were selected, with a total of 12 composite scheduling rules:

[0153] Process sorting PDRs: FIFO (First In, First Out): select the earliest ready candidate process and the earliest ready compatible machine; MOPNR (Most Operations Remaining): select the candidate process with the most subsequent operations and select the machine that can process it immediately; SPT (Shortest Processing Time): select the compatible process-machine pair with the shortest processing time; MWKR (Most Work Remaining): select the process with the most average processing time of subsequent operations and select the machine that can process it immediately.

[0154] Batch-machine collaborative division PDRs: Two-batch: divide the task into two batches with as even a quantity as possible; Equal-inconsistent batching: minimize the completion time of each process while ensuring that the processing quantity on each machine is distributed as evenly as possible; Unequal-inconsistent batching: minimize the completion time of each process and perform unequal batching according to the machine load.

[0155] Each composite scheduling rule was run 20 times independently in the experiment, and the average result of each experiment was recorded to evaluate the performance of the algorithm.

[0156] 4. Experimental results and analysis

[0157] Through the above experiments, the proposed algorithm is compared with the baseline algorithm in detail. Table 2 shows the scheduling effect of the algorithm and the composite rule. The algorithm has a strong solving ability. Table 3 shows the performance of the test algorithm on examples of different scales and characteristics. The algorithm has a strong generalization ability. Table 4 compares the calculation time of different intelligent optimization algorithms. The proposed algorithm has high computational efficiency when solving large-scale examples. Through these experiments, it is verified that the proposed algorithm can provide more efficient and accurate solutions than the baseline algorithm when dealing with complex scheduling problems.

[0158] Table 2 Comparison test table with composite scheduling rules

[0159]

[0160] Table 3 Generalization test table compared with composite scheduling rules

[0161]

[0162] Table 4 Comparison test table with intelligent optimization algorithm

[0163]

[0164] It will be easily understood by those skilled in the art that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A flexible job shop scheduling method for inconsistent batch flow integrating mathematical model and reinforcement learning, characterized in that: The steps include: The flexible job shop scheduling problem is represented as a heterogeneous graph, where processes and machines are different types of nodes, and process-machine processing relationships are edges. The flexible job shop scheduling problem is transformed into a Markov decision process, and the action space, reward function and decision network are designed. The action space is a combination of the processing steps to be processed and the batch partition coefficient γ. The batch partition coefficient γ is used to allocate the processing steps to the machines with lower processing time cost. The decision network is constructed based on the actor-critic network. A mixed integer linear programming model for process scheduling is established, and the objective function of the mixed integer linear programming model is constructed based on the batch partition coefficient; Use the proximal policy optimization algorithm to perform reinforcement learning training in a simulated environment, including: S1, extracting environmental state features based on the current heterogeneous graph; S2, the actor network makes decisions based on the environmental state characteristics and determines the current processing steps and batch division coefficients; S3. Solve the mixed integer linear programming model for the current process to be processed to obtain a scheduling solution, i.e., batch division and machine selection for the process to be processed; S4, update the heterogeneous graph based on the scheduling scheme and optimize the decision network parameters according to the reward function; S5. Repeat S1 to S4 until all decisions on pending processes are completed, and a flexible job shop scheduling method is obtained.

2. The inconsistent batch flow flexible job shop scheduling method integrating mathematical model and reinforcement learning as claimed in claim 1 is characterized in that: The action space A(t) is expressed as: A(t)={(O i,j ,γ)|O i,j ∈J c (t),γ∈Γ} Among them, O i,j represents the jth process of the i-th type of operation, i.e., the process to be processed, J c (t) represents the set of candidate processes at time t; Γ represents the discretized set of batch partition coefficients γ, Γ={γ1,γ2,…,γ l }, Where l represents the number of discretizations, γ v ∈[0,1].

3. The inconsistent batch flow flexible job shop scheduling method integrating mathematical model and reinforcement learning as claimed in claim 1 is characterized in that: The mixed integer linear programming model uses batch partitioning variables, machine selection variables and completion time variables in coordination, and takes minimizing the maximum completion time and processing preparation time as the optimization goal.

4. The inconsistent batch flow flexible job shop scheduling method integrating mathematical model and reinforcement learning as claimed in claim 3 is characterized in that: For the jth process O of the i-th type of operation i,j , the objective function of the mixed integer linear programming model is: Among them, C i,j Indicates O i,j The maximum completion time of the sub-batch, B i,j Indicates O i,j Total number of sub-batches, S i,j,k Indicates O i,j In M i,j,k Preparation time for processing, x i,j,k Indicates O i,j Is it in M i,j,k Upper processing, M i,j,k Indicates O i,j The kth processable machine.

5. The inconsistent batch flow flexible job shop scheduling method integrating mathematical model and reinforcement learning as claimed in claim 1, characterized in that: When the proximal policy optimization algorithm is used for reinforcement learning training, the objective function is to minimize the total completion time of the workpiece.

6. The inconsistent batch flow flexible job shop scheduling method integrating mathematical model and reinforcement learning as claimed in claim 1, characterized in that: The reward function r t It is expressed as: Among them, C((O i,j ,γ),s t )、C((O i,j ,γ),s t+1 ) represent process O respectively i,j The estimated completion time at time step t and time step t+1, s t 、s t+1 Represent the states at time step t and time step t+1 respectively.

7. The inconsistent batch flow flexible job shop scheduling method integrating mathematical model and reinforcement learning as described in any one of claims 1 to 6, characterized in that: The environmental state characteristics include process characteristics Machine Features Sub-batch characteristics Process Machine Pair Features and the global feature h G .

8. The inconsistent batch flow flexible job shop scheduling method integrating mathematical model and reinforcement learning as claimed in claim 7 is characterized in that: The environmental state characteristics are determined as follows: Based on the current heterogeneous graph, obtain the initial process characteristics Initial Machine Characteristics Sub-batch characteristics and process machine pair features Initial Operation Features Initial Machine Characteristics Input graph attention network to get process characteristics Machine Features Operation Features Machine Features and sub-batch characteristics After average pooling, the global feature h is obtained. G .

9. A flexible job shop scheduling system for inconsistent batch flow integrating mathematical model and reinforcement learning, characterized in that: It includes a processor, which is used to execute the inconsistent batch batch flow flexible job shop scheduling method integrating mathematical model and reinforcement learning as described in any one of claims 1-8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the inconsistent batch flow flexible job shop scheduling method integrating mathematical model and reinforcement learning as described in any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Flexible job shop scheduling method considering preparation time and workpiece batching

    CN112541694A

  • Flexible job shop scheduling method and device based on sub-batch division

    CN116880410A

  • Inconsistent batch flow flexible job shop scheduling method and system

    CN116974245A

  • Flexible workshop scheduling method and system based on deep reinforcement learning and heterogeneous graph network

    CN119129930A

  • Production task scheduling method, system and device for flexible assembly job shop

    US20240177251A1

Cited By

  • Two-stage batch flow flexible job shop dynamic scheduling method and device

    CN121903309A