CNN-Transform-based dynamic flexible job shop scheduler and scheduling method
By combining the characteristics of CNN and Transformer in dynamic flexible operation workshop scheduling, local and global features are extracted, and the problem that the existing technology is difficult to meet real-time and global in large-scale and dynamic environments is solved, and efficient scheduling scheme generation is achieved.
Patent Information
- Application Number
- CN202510337647.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-27
AI Technical Summary
When the existing technology faces the problem of dynamic flexible work workshop scheduling, heuristic algorithms lack globality, and metaheuristic algorithms are difficult to meet real-time requirements in large-scale problems and rapidly changing environments.
A dynamic flexible work workshop scheduler based on CNN-Transformer is adopted to extract local features and Transformer through CNN to supplement global information, and combine token mixer and dynamic feature fusion to improve the generalization and interpretability of the model.
It realizes efficient solution to large-scale FJSP problems, improves the quality and operation efficiency of the scheduling scheme, and adapts to changes in the dynamic environment.
Smart Images

Figure CN120215438A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a dynamic flexible job shop scheduler and scheduling method based on CNN-Transformer. Background Art
[0002] Intelligent manufacturing penetrates all aspects of manufacturing activities such as design, production, management, and service, and is a new production method with functions such as self-perception, self-learning, self-decision-making, self-execution, and self-adaptation. As an important part of intelligent manufacturing, the progress of production scheduling technology is an important force to promote the digital transformation of manufacturing enterprises and support the development of intelligent manufacturing.
[0003] Flexible Job Shop Scheduling Problem (FJSP) is widely applied in fields such as manufacturing and logistics, and involves scheduling different tasks in one or more workshops to ensure the efficiency of the production process. In a flexible job shop, there are multiple processing operations, and each operation is completed by multiple different machines. The scheduling problem lies in how to optimize the processing sequence of tasks and the allocation of resources under limited resources and time constraints.
[0004] In actual production, the flexible job shop scheduling problem has high uncertainty and dynamics. Equipment failures, changes in task priorities, and shortages of raw materials may occur during the production process, requiring the scheduling system to be able to flexibly adjust to handle emergencies. Therefore, the Dynamic Flexible Job Shop Scheduling Problem (DFJSP) has also become an important research topic in academia and industry.
[0005] Traditional solution methods for the dynamic flexible job shop scheduling problem mainly include heuristic algorithms and meta-heuristic algorithms.
[0006] (1) Heuristic algorithms
[0007] Heuristic algorithms are rule-based solution algorithms set according to workshop information or experience. The priority scheduling rule is the most commonly used method in actual production. By evaluating the priorities of the operations or machines to be scheduled, scheduling is performed according to the priority relationship to obtain a scheduling solution. Commonly used scheduling rules such as SPT (Shortest Processing Time), EDD (Earliest Due Date), etc.
[0008] Scheduling rules have the advantages of low computational complexity and easy implementation. They can make local decisions when dynamic events occur and have thus been used to solve a large number of complex dynamic scheduling problems. Panwalkar and Iskander summarized and analyzed the priority criteria and applicable optimization objectives of 113 scheduling rules. Chandrasekharan and Oliver studied the relative performance of 13 scheduling rules considering different optimization objectives for the dynamic flow shop and job shop problems with missing operations. The results show that the performance of scheduling rules is greatly affected by the problem scale and shop configuration. Yi et al. considered the priority numbers, regular pauses, and workpiece weights of general priority rules and proposed a scheduling algorithm based on multi-level priority rules, achieving good optimization results when solving the scheduling problem of the production workshop of die enterprises with workpiece insertion and processing pauses. Peng et al. proposed a simple combinatorial scheduling algorithm design framework and generated 72 combinatorial rule algorithms by embedding scheduling rules in the framework to solve the task-tool joint dynamic scheduling problem in a complex surface automation unit with dynamically arriving workpieces. The results show that no single algorithm can guarantee an optimal solution under all preset environmental parameters, and no algorithm is found to guarantee a better solution in all evaluation indicators.
[0009] Although scheduling rules meet the real-time requirements for solving dynamic scheduling problems, different scheduling rules are applicable to different optimization objectives. When facing different dynamic events and changing problem scales, the solution quality of the same rule varies greatly. Selecting the most effective scheduling rule at the scheduling decision point is the difficulty of the problem. To address this issue, some studies combine deep reinforcement learning methods to train agents to select the most appropriate scheduling rule according to the production status at each decision point. Li et al. developed a hybrid deep Q-network for the dynamic FJSP problem with new workpiece insertion and machine failures, enabling the agent to select a suitable scheduling rule according to the production status at each decision point. However, the selected scheduling rules are still limited to the scope of manual design. Other studies have attempted to develop new scheduling rules to improve scheduling performance. Wang et al. proposed a framework combining job due date prediction, agent-based simulation, and evolutionary algorithms to generate new scheduling rules to solve the job shop scheduling problem with machine failures, dynamically arriving workpieces, and uncertain processing times. However, the heuristic algorithms using scheduling rules to solve dynamic scheduling only consider local information and usually have poor solution quality and lack of globality.
[0010] (2) Metaheuristic algorithms
[0011] Metaheuristic algorithms are optimization algorithms based on heuristic search. They do not depend on the specific characteristics of the problem and search the large solution space iteratively by imitating natural or social phenomena, constantly seeking approximate optimal solutions. Common metaheuristic algorithms include genetic algorithms, particle swarm algorithms, ant colony algorithms, simulated annealing algorithms, etc.
[0012] Meta-heuristic algorithms are flexible, efficient, and highly scalable, making them suitable for solving complex optimization problems of different scales. There have been many studies and applications in solving dynamic job shop scheduling problems. Nouiri et al. proposed a two-stage particle swarm optimization method for generating a pre-scheduling plan for the flexible job shop scheduling problem under machine failures, and the obtained solution results have a certain degree of robustness. Zhu et al. proposed a restructured memetic algorithm to solve the distributed FJSP considering order cancellations, and designed a method for rescheduling according to the different processing states of workpieces when orders are cancelled. Wei et al. proposed a multi-objective migratory bird optimization algorithm based on game theory to solve the bi-objective dynamic flexible job shop scheduling problem with machine failures, and adopted a hybrid rescheduling strategy combining periodic rescheduling and event-driven to obtain better solution quality. You et al. adopted a hybrid approach driven by events and cycles for the FJSP problem with stochastic machine failures, and combined the elite selection genetic algorithm for rescheduling. Duan et al. considered the reusability of the system and the repeatability of processing tasks, and proposed a multi-objective particle swarm optimization algorithm to solve the flexible job shop scheduling problem with machine failures and the arrival of new workpieces. Different rescheduling strategies were adopted according to the nature of dynamic events, and comparative experiments showed that their method could effectively improve the robustness of the solution results. Li et al. used the non-dominated sorting genetic algorithm II to solve the flexible job shop problem under dynamic workpiece arrivals and machine failures, and updated the states of in-progress and unfinished workpieces and machines when dynamic events occurred. Shen Chunya et al. proposed an improved genetic algorithm and designed a rescheduling mechanism based on the evaluation of the dominance relationship of optimization objectives to solve the dynamic scheduling problem of a weaving workshop with complex production scenarios such as order insertion and sample making. Jiang Fei et al. designed an improved grey wolf optimization algorithm with dynamically changing convergence factors and leader wolf weights to cope with the disturbance events of machine failures for solving the fuzzy flexible job shop dynamic scheduling problem with makespan and customer satisfaction as optimization objectives. An et al. proposed a hybrid multi-objective evolutionary algorithm for dealing with the flexible job shop rescheduling problem with real-time order acceptance and state-based preventive maintenance, and adopted different rescheduling strategies according to the remaining processing operations and the scale of newly arrived workpieces. Ge Yan et al. introduced the population and mutation concepts of the genetic algorithm into the simulated annealing algorithm to solve the flexible job shop scheduling problem considering the rework of individual processes and workpiece rework caused by workpiece quality inspection.
[0013] In summary, although meta-heuristic algorithms have been widely used in the research of dynamic job shop scheduling problems and achieved good solution results, when facing the growth of problem scale and the changes caused by dynamic events, the running time of the algorithm to search the solution space is relatively long, making it difficult to meet the real-time requirements of solving job shop scheduling under dynamic events.
[0014] In recent years, with the continuous development of artificial intelligence technology, especially machine learning algorithms, the application of machine learning in job shop scheduling has attracted wide attention. Machine learning, especially deep learning and reinforcement learning, provides new solutions for job shop scheduling, especially showing unique advantages when dealing with dynamic scheduling problems. Traditional scheduling methods mostly rely on heuristic algorithms and optimization models, while machine learning can automatically learn scheduling strategies in a data-driven manner to adapt to changes in dynamic environments. Aiming at the deficiencies of the traditional heuristic algorithm with poor global solution quality and the meta-heuristic algorithm with long solution time, the method of automatic learning and exploration by agents is used, and certain scheduling performance has been obtained in the solution of most scheduling problems through training.
[0015] Deep learning is a special machine learning method that can abstract higher-dimensional features of input features through a deep neural network and obtain the connection between input features and the desired output. In recent years, with the development of deep learning methods, deep learning has been widely used to solve classification problems, regression problems, and sequence problems, which provides a new idea for solving dynamic scheduling problems.
[0016] In DFJSP, a suitable workpiece processing sequence and the corresponding machine sequence need to be obtained. Therefore, DFJSP can be regarded as a type of sequence problem. Deep learning methods such as CNN and Recurrent Neural Network (RNN) can be used to solve sequence problems. Most scholars regard job shop scheduling or flow shop scheduling as sequence problems and use deep learning methods to solve them. For example, Qiu Sichen et al. adopted an encoder-decoder structure based on deep learning methods to solve the scheduling problem. The high-dimensional features of the input data were extracted through the encoder, and then the processing order of the workpieces was obtained through the decoder. Ren Jianfeng et al. proposed an RNN model based on the pointer network. After training the RNN, the characteristics that the RNN can output sequence data were used to solve the job shop scheduling problem, and the effectiveness of the model was verified by comparing it with the meta-heuristic algorithm. Cxw et al. proposed a deep learning model for solving the dynamic scheduling problem of flow shop. The processing sequence of the workpieces was output through this model, so as to obtain a solution with better quality in a relatively short time. Liu proposed a deep learning model based on CNN for solving the dynamic flexible job shop scheduling problem, and the processing sequence of the workpieces was output through this model.
[0017] Deep learning is good at extracting local features and can capture the spatial structure information in the job shop scheduling problem. When solving the job shop scheduling problem, it can take both the solution quality and the solution efficiency into account. The extraction of the job shop scheduling state, as the input of the scheduling strategy, has an important impact on the generalization effect of the model. Some studies use defined matrices, feature vectors, or sets of feature vectors to represent the current state of the job shop. However, suitable features need to be manually designed in combination with the specific problem characteristics and optimization objectives. Deep learning is difficult to model long-distance dependencies, has relatively insufficient generalization ability, often requires more computing resources to capture global information, and is difficult to effectively and comprehensively express job shop information.
[0018] Transformer was initially used in natural language processing and endows each element in the sequence with a global receptive field through the self-attention mechanism to capture context information in parallel. With the success of models such as ViT, DETR, and Swin Transformer in the field of computer vision, Transformer, which can capture long-distance dependencies, has also attracted much attention in medical image segmentation tasks. TransUNet is the first model to apply Transformer to the U-Net architecture, regarding the image encoding block from the shallow CNN as a semantic sequence, and then using Transformer to extract global context information. In the decoder part, the current features are fused with the high-resolution, low-semantic features from the shallow CNN. Subsequently, Wang et al. proposed TransBTS, which embeds Transformer into the neck of the three-dimensional U-Net, reducing the computational cost while learning global information. nnFormer alternately uses convolution and self-attention in the encoder to extract features and applies spatial attention to the skip connections in the decoder. UNETR++ further proposed an Efficient Paired Attention (EPA) block composed of spatial and channel attention, reducing the computational complexity of spatial attention from square to linear and improving the computational efficiency through parameter sharing. FCT embeds convolution in the Transformer architecture, replacing the linear projection layers in the self-attention module and the feed-forward neural network with dilated convolution and depthwise separable convolution respectively to improve performance by leveraging the strong spatial information of convolution.
[0019] In the field of scheduling, CNN is widely used for its excellent local feature extraction ability, but its global feature extraction ability is limited by the size of the convolution kernel. In contrast, Transformer with a global receptive field can effectively extract global information, but it has poor adaptability on small-scale datasets and a high computational complexity. One of the future development trends is to more comprehensively and deeply integrate deep learning models with different structures such as CNN and Transformer to improve the performance of the model. Although many studies have explored the combination of convolution and self-attention mechanisms, there is still room for improvement in the degree of integration, and these studies often ignore the problem of increased model parameters and computational volume. Summary of the Invention
[0020] The token mixer of the Transformer module allows the model to parallelly focus on different subspaces of the input features, such as simultaneously focusing on multi-dimensional features such as the processing time of processes and machine loads, accelerating the training process, having strong generalization ability, and being applicable to large-scale FJSP problems. Through the token mixer and dynamic feature fusion, CNN-Transformer can more efficiently solve large-scale FJSP problems. Therefore, the present invention provides a dynamic flexible job shop scheduler and scheduling method based on CNN-Transformer, which complements the advantages of CNN being good at extracting local features and Transformer being good at supplementing global information, enhances the generalization and interpretability of the deep learning model, and is used to solve large-scale FJSP. Scheduling knowledge is extracted from the scheduling scheme obtained by the hybrid particle swarm tabu search algorithm as the training data of CNN-Transformer, scheduling attributes that can reflect the scheduling scheme information are selected, and a method for converting scheduling attributes into training data is given. The effectiveness of the dynamic flexible job shop scheduler based on CNN-Transformer of the present invention is verified through experiments.
[0021] The technical solution adopted by the present invention is as follows:
[0022] A dynamic flexible job shop scheduler based on CNN-Transformer, characterized in that it includes:
[0023] An input layer, configured to select the scheduling attribute values of each workpiece on each machine in each process from the scheduling scheme and convert the scheduling attribute values into input features in the training data of the deep learning model;
[0024] The CNN-Transformer model includes a CNN encoder and a Transformer module, where: the CNN encoder extracts deeper features from the input features through convolutional layers; the Transformer module uses the self-attention mechanism to map the deeper features extracted by the CNN encoder from a high-dimensional space to a low-dimensional space, and decodes and outputs the process sequences and machine sequences of each workpiece in the rescheduling plan through a channel multi-layer perceptron;
[0025] An output layer for outputting the process sequences and machine sequences of each workpiece in the rescheduling plan.
[0026] The scheduling attribute values include: workpiece number, the processing status of the process at the rescheduling moment, the processing time of the process on each machine, the processing machine corresponding to the process, the order of the process on the machine, the remaining processing time of the process being processed, the available start time of the machine where the process is located, the cumulative processing duration of the machine where the process is located, and the utilization rate of the machine where the process is located.
[0027] The CNN-Transformer model inserts a Transformer module into the CNN encoder, and the Transformer module serves as a decoder. The CNN encoder consists of a convolutional layer, a pooling layer, a normalization layer, an activation function layer, an Fc layer, and a Dropout layer. The Transformer module consists of six Transformer_Blocks, and each Transformer_Block is correspondingly connected to one layer in the CNN encoder. Each Transformer_Block consists of a token mixer, a channel multi-layer perceptron (Multilayer Perceptron, MLP), a learnable scaling factor (scale), and two layer normalizations (LayerNorm). The token mixer is used to enrich the feature representation of the model; the channel multi-layer perceptron is responsible for feature processing in the channel dimension; scale is used to balance the influence of the outputs of each part, and the two layer normalizations are respectively used for the normalization of the input and output.
[0028] The token mixer adopts the self-attention mechanism to capture the relationship between any two positions in the input sequence, so as to extract the complex dependencies within the sequence.
[0029] The channel multi-layer perceptron (MLP) includes two fully connected networks for non-linearly transforming and dimensionally adjusting the features.
[0030] The learnable scaling factor (scale) is dynamically adjusted through training to optimize the weights of the outputs of each part in the Transformer_Block.
[0031] Scheduling method for a dynamic flexible job shop scheduler based on CNN-Transformer, characterized by the following steps:
[0032] (1) Select scheduling attribute values: The input layer selects the scheduling attribute values of each workpiece on each machine in each process from the scheduling plan; the scheduling attribute values include workpiece number, processing status of the process at the rescheduling moment, processing time of the process on each machine, processing machine corresponding to the process, processing order of the process on the machine, remaining processing time of the process being processed, available start time of the machine where the process is located, cumulative processing duration of the machine where the process is located, and utilization rate of the machine where the process is located;
[0033] (2) Convert scheduling attribute values: Convert the scheduling attribute values obtained in step (1) into input features in the training data of the deep learning model; where the columns of the input features represent each process, and the rows of the input features represent the scheduling attribute values of each process;
[0034] (3) Input the input features into the CNN-Transformer model, extract the deep features of the input features through the convolutional layer, and perform dimension expansion, depth convolution, and dimension reduction operations through the token mixer of the Transformer module to capture local features and complex dependencies within the sequence;
[0035] (4) Output the workpiece sequence and machine sequence of the rescheduling plan calculated by the CNN-Transformer model.
[0036] The label of the input features is a one-dimensional matrix formed by splicing the workpiece sequence and machine sequence of the scheduling plan after rescheduling.
[0037] The calculation process of the input features and output features of the convolutional layer is shown in formula (1-1):
[0038]
[0039] Among them, is the data block at coordinates (row, col) in the th feature of the lay-th convolutional layer, f represents the activation function, Cha represents the number of channels of the feature, is the data block at coordinates (row, col) in the th feature of the (lay-1)-th convolutional layer, wk lay represents the convolutional kernel of the lay-th convolutional layer, bias lay represents the bias of the lay-th convolutional layer.
[0040] The token mixer of the Transformer module adopts the self-attention mechanism to capture complex dependencies within the sequence by calculating the relationship between any two positions in the input sequence.
[0041] The dynamic flexible job shop scheduler based on CNN-Transformer of the present invention combines the advantages of CNN's ability to extract local features and Transformer's ability to supplement global information, enhances the generalization and interpretability of the deep learning model, and is used to solve larger-scale FJSP. Scheduling knowledge is extracted from the scheduling scheme obtained by the hybrid particle swarm tabu search algorithm as the training data of CNN-Transformer, scheduling attributes that can reflect the scheduling scheme information are selected, and a method for converting scheduling attributes into training data is given. The effectiveness of the dynamic flexible job shop scheduler based on CNN-Transformer of the present invention is verified through experiments. Brief Description of the Drawings
[0042] Figure 1 It is a schematic structural diagram of the present invention.
[0043] Figure 2 It is a schematic diagram of the Transformer model of the present invention.
[0044] Figure 3 It is a schematic diagram of the initial scheduling scheme of the present invention.
[0045] Figure 4 It is a schematic diagram of the rescheduling scheme of the present invention.
[0046] Figure 5 It is a schematic diagram of the self-attention mechanism structure of the present invention. Detailed Embodiment
[0047] The present invention will be further described in conjunction with the accompanying drawings and embodiments.
[0048] As Figure 1 shown, a dynamic flexible job shop scheduler based on CNN-Transformer of the present invention mainly can be divided into three parts: an input layer, a CNN-Transformer model, and an output layer;
[0049] The input layer is used to select the scheduling attribute values of each workpiece on each machine in each process from the scheduling plan and convert the scheduling attribute values into input features in the training data of the deep learning model. The scheduling attribute values include: workpiece number, processing status of the process at the rescheduling moment, processing time of the process on each machine, processing machine corresponding to the process, processing order of the process on the machine, remaining processing time of the ongoing process, available start time of the machine where the process is located, cumulative processing duration of the machine where the process is located, and utilization rate of the machine where the process is located. Among them, the columns in the input features represent each process, and the rows of the input features represent the scheduling attribute values of each process, and the encoded features are transmitted to the Transformer module. The label of the input features is a one-dimensional matrix formed by splicing the workpiece sequence and the machine sequence of the scheduling plan after rescheduling.
[0050] The CNN-Transformer model inserts a CNN encoder into the Transformer module, and the Transformer module serves as the decoder. The CNN encoder extracts deeper features from the input features through convolutional layers. The Transformer module uses the self-attention mechanism to map the deeper features extracted by the CNN encoder from a high-dimensional space to a low-dimensional space, and decodes and outputs the process sequence and machine sequence of each workpiece in the rescheduling plan through a channel multi-layer perceptron.
[0051] Such as Figure 1As shown, Conv_Block_1 - Conv_Block_6 in the CNN encoder are all used to represent CNN. The CNN encoder consists of a convolutional layer, a pooling layer, a normalization layer, an activation function layer, an Fc layer, and a Dropout layer. The Transformer module consists of six Transformer_Blocks. Each Transformer_Block is correspondingly connected to one layer in the CNN encoder. Each Transformer_Block consists of a token mixer, a channel Multilayer Perceptron (MLP), a learnable scale factor, and two Layer Norms. The token mixer is used to enrich the feature representation of the model. The token mixer adopts a self-attention mechanism to capture the relationship between any two positions in the input sequence, thereby extracting complex dependencies within the sequence. The channel Multilayer Perceptron is responsible for feature processing in the channel dimension. The channel Multilayer Perceptron (MLP) includes two fully connected networks for non-linearly transforming and adjusting the dimensions of the features. The scale is used to balance the influence of the outputs of each part. The learnable scale factor is dynamically adjusted through training to optimize the weights of the outputs of each part in the Transformer_Block. The two Layer Norms are respectively used for the normalization of the input and output.
[0052] The output layer is used to output the operation sequence and machine sequence of each workpiece in the rescheduling plan.
[0053] The scheduling method of the dynamic flexible job shop scheduler based on CNN-Transformer is characterized by the following steps:
[0054] (1) Select scheduling attribute values: The input layer selects the scheduling attribute values of each workpiece on each machine in each operation from the scheduling plan. The scheduling attribute values include workpiece number, the processing status of the operation at the rescheduling moment, the processing time of the operation on each machine, the processing machine corresponding to the operation, the processing order of the operation on the machine, the remaining processing time of the operation being processed, the available start time of the machine where the operation is located, the cumulative processing duration of the machine where the operation is located, and the utilization rate of the machine where the operation is located.
[0055] (2) Conversion of scheduling attribute values: Convert the scheduling attribute values obtained in step (1) into input features in the training data of the deep learning model. Among them, the columns of the input features represent each operation, and the rows of the input features represent the scheduling attribute values of each operation.
[0056] (3) Input the input features into the CNN-Transformer model, extract the deep features of the input features through the convolutional layer, and perform dimension expansion, depth convolution, and dimension reduction operations through the token mixer of the Transformer module to capture local features and complex dependencies within the sequence;
[0057] (4) Output the workpiece sequence and machine sequence of the rescheduling plan calculated by the CNN-Transformer model.
[0058] The calculation process of the input features and output features of the convolutional layer is shown in formula (1-1):
[0059]
[0060] Among them, is the data block at coordinates (row, col) in the th feature of the lay-th convolutional layer, f represents the activation function, Cha represents the number of channels of the feature, is the data block at coordinates (row, col) in the th feature of the (lay-1)-th convolutional layer, wk lay represents the convolutional kernel of the lay-th convolutional layer, and bias lay represents the bias of the lay-th convolutional layer.
[0061] There are mainly 6 types of network layers in the CNN encoder built in the present invention, namely convolutional layer, pooling layer, normalization layer, activation function layer, Fc layer, and Dropout layer. The parameter settings for the CNN encoder are shown in Table 1:
[0062] Table 1 Hyperparameters of the CNN model
[0063]
[0064] Due to the limitation of the receptive field of the convolutional kernel in the CNN encoder, it is difficult to capture global information. Therefore, a Transformer module is inserted as a decoder in the CNN-based dynamic flexible job shop scheduler. As Figure 2 shown. After receiving the data processed by the CNN encoder, the Transformer module uses the self-attention mechanism to map the features encoded by the CNN encoder from the high-dimensional space to the low-dimensional space to learn semantic information at different levels, and then calculates and outputs the workpiece sequence and machine sequence of the rescheduling plan.
[0065] Next, the structure of the Transformer module is introduced. Each Transformer_Block consists of: a token mixer for enriching the feature representation ability of the model; a channel Multilayer Perceptron (MLP) responsible for feature processing in the channel dimension; a scale which is a learnable scaling factor, providing a flexible mechanism for the Transformer Block to balance the influence of the outputs of each part, enabling the model to learn useful features more efficiently and improving the training stability and overall performance of the model; and two Layer Norms for normalizing the input and output to ensure the stability and training efficiency of the model. Among them, the self-attention mechanism aims to map features from a high-dimensional space to a low-dimensional space to learn semantic information at different levels and is commonly used in the Transformer module structure.
[0066] When training the deep learning based on the CNN-Transformer model, the result of solving the MK04 example by the hybrid particle swarm tabu search algorithm is used as the initial scheduling plan. Based on the initial scheduling plan, eleven groups of examples are designed to generate training data. In these eleven groups of examples, in addition to the jobs in the initial scheduling plan there will also be jobs arriving at the production workshop dynamically. The processing information of the dynamically arriving jobs in one group of examples comes from the MK04 example, and the processing information of the dynamically arriving jobs in the other ten groups of examples comes from randomly generated examples. The random examples are generated referring to the method given by Chambers. When considering the DFJSP as a sequence problem to solve, the problem of combinatorial explosion of sequence data needs to be taken into account, that is, as the number of dynamically arriving jobs increases, the combination of sequence data will grow exponentially, but the structure and parameters of the model often cannot adapt to this change quickly, resulting in the model not being able to converge to a good result and thus unable to obtain the job sequence and its corresponding processing machine sequence. When designing the examples, considering the problem of combinatorial explosion of sequence data, the number of dynamically arriving jobs is set to 1. In addition, although the data used to train the CNN-Transformer model are all feasible solutions generated by the TSPSO algorithm, during the training process, it is also necessary to ensure that a certain process is processed on the alternative machines of this process. Therefore, when the model calculates that the processing machine of a certain process cannot process this process, the machine closest to the calculation result will be selected from the set of alternative machines of this machine as the calculation result.
[0067] The specific processing is as follows:
[0068] The method process of the CNN-Transformer model for processing data is as follows:
[0069] (1) Select scheduling attributes
[0070] The scheduling plan in actual production cannot be directly input into the neural network. Therefore, some scheduling attributes need to be selected from the scheduling plan, and the scheduling attribute values need to be converted into the data required for training the neural network. The selected scheduling attributes should reflect the characteristics of the actual scheduling plan as comprehensively as possible, which is conducive to the neural network learning scheduling-related knowledge and thus better solving the scheduling problem. The selected scheduling attributes are as follows:
[0071] ● Job number: The operations belonging to the same job are linked through the job number. The job number of the newly arrived job is recorded as 0.
[0072] ● Processing status of the operation at the rescheduling moment: Indicates whether a certain operation needs to be rescheduled. The completed operation is recorded as 0, the unprocessed operation is recorded as 1, and the operation in progress is recorded as 2.
[0073] ● Processing time of the operation on each machine: It can describe the time required for the operation to be processed. The number of lines occupied by this scheduling attribute is the same as the number of machines. If a machine cannot process a certain operation, it is recorded as 0; otherwise, it is recorded as the time for the machine to process the operation.
[0074] ● Processing machine corresponding to the operation: Represents the correspondence between the operation and the machine in the original scheduling plan. If the current operation has not been processed yet, it is recorded as 0.
[0075] ● Processing sequence of the operation on the machine: Represents the processing situation of the operation on the machine in the original scheduling plan,
[0076] If the current operation has not been processed yet, it is recorded as 0.
[0077] ● Remaining processing time of the operation in progress: Indicates whether a certain machine is occupied and the remaining time it needs to be occupied. If the current operation has not been processed, it is recorded as 0.
[0078] ● Available start time of the machine where the operation is located: Represents the available start time of the machine processing this operation at the rescheduling moment. Since the machine may be processing a certain operation at the rescheduling moment, resulting in the inability to immediately process other operations, this attribute can provide information for arranging subsequent operations on this machine. If the current operation has not been processed, it is recorded as 0.
[0079] ● Cumulative processing duration of the machine where the operation is located: Represents the duration of the processed operations on the machine. If the current operation has not been processed, it is recorded as 0.
[0080] ● Utilization rate of the machine where the operation is located: Represents the busy degree of the machine and can provide information for subsequent jobs to select processing machines. If the current operation has not been processed, it is recorded as 0.
[0081] (2) Conversion of scheduling attributes
[0082] After extracting the scheduling attributes, it is also necessary to convert the scheduling attribute values into input features in the training data of the deep learning model. Taking the operation as an example:
[0083]
[0084] After converting the scheduling attribute values into input features in the training data, it is also necessary to determine the labels of the input features; considering that the labels should reflect the processing machines and processing sequences arranged for the original workpieces and dynamically arriving workpieces after rescheduling, the workpiece sequence {2, 1, 3, 4, 1, 2, 4, 3, 4} and the machine sequence {3, 2, 1, 2, 3, 3, 2, 1, 3} of the scheduling plan after rescheduling are concatenated as the labels of the input features, and this label is a one-dimensional matrix.
[0085] (3) CNN-Transformer processes data
[0086] The input data of the CNN-Transformer model are the input features converted from the scheduling attribute values extracted from the scheduling plan, where the columns in the input features represent each operation, and the rows of the input features represent the scheduling attribute values of each operation. The output is the workpiece sequence and machine sequence of the rescheduling plan calculated by the CNN-Transformer model.
[0087] ① Convolutional layer
[0088] CNN extracts deeper features from the input features through the convolutional layer as the input for subsequent network layers. The convolutional layer uses a convolutional kernel to scan the input features. During the scanning process, the feature values covered by the convolutional kernel will perform a convolutional calculation with the convolutional kernel, and then the result of the convolutional operation is input into a new two-dimensional matrix, and this two-dimensional matrix is the output feature obtained after the input features pass through the convolutional layer. The calculation processes of the input features and output features of the convolutional layer are shown in formula (1-1);
[0089]
[0090] Among them, is the data block at the coordinate (row, col) in the th feature of the lay-th convolutional layer, f represents the activation function, Cha represents the number of channels of the feature, is the data block at the coordinate (row, col) in the th feature of the (lay-1)-th convolutional layer, wk lay represents the convolutional kernel of the lay-th convolutional layer, and bias lay represents the bias of the lay-th convolutional layer.
[0091] After the convolution operation in the CNN, the features after convolution are dimensionally expanded, deeply convolved, and dimensionally reduced again through the token mixer of the Transformer, aiming to efficiently capture local features while reducing computational complexity. The token mixer adopts the self-attention mechanism (SA), which is a mechanism for calculating the relationship between any two positions in the input sequence. By calculating the attention of each element in the input sequence with other elements in the sequence, the model focuses on the features of important positions, thereby capturing complex dependencies within the sequence. In the Transformer, SA is used to assign attention weights to data features from the perspective of time steps, enabling the model to focus on the feature information of important time steps during feature extraction, thereby enhancing the feature extraction ability and compression representation ability. The structure of the self-attention mechanism is as Figure 5 shown. First, the input features and three 1×1 convolutional kernels are used to construct the query Q, key K, and value V. Among them, Q represents the current element, and K represents other elements in the time-series data. The attention weights are calculated by computing the similarity between Q and K where is the number of channels of Q and K. The feature distribution of the inner product result of Q and K is related to the number of channels and is decoupled by dividing by to make the training process gradient stable. Then, the obtained time-step attention weights are applied to V to output the new feature A. Finally, the output feature O′ of the SA mechanism is obtained through the residual connection between A and the input feature , as shown in Equation (1-2).
[0092]
[0093] where is the input feature, Q represents the current element, and K represents other elements in the time-series data; A score is the attention weight; d k is the number of channels of Q and K; A is the residual connection; O′ is the output feature of the SA mechanism.
[0094] Introducing instance normalization and ReLU activation functions after each convolution operation can effectively adjust the feature distribution and enhance the model's expressive ability by introducing non-linearity. This is not just a simple feature stacking, but through a hierarchical and recursive way, the features of each layer are iteratively fused from shallow to deep, extracting useful information layer by layer to ensure that the important information of each layer is fully reflected in the final feature representation.
[0095] ② Normalization layer
[0096] In a CNN, even a slight change in the parameters of the network layers in the lower levels can have a significant impact on the distribution of the input features for each subsequent network layer. This leads to the need for the network layers in the lower levels to adjust their parameters very frequently to adapt to the changes in the input features, and ultimately, the situation of gradient disappearance or even non - convergence of the network may occur. To solve this problem, by normalizing the input features, the distribution of the input features in different batches will not deviate too much, so that the network layers farther from the input end no longer frequently adjust their own parameters according to the changes in the input features, accelerating the convergence of the network. Therefore, a normalization layer is added after the convolutional layer to achieve the normalization of features. After the calculation of the normalization layer, some data in the input features can fall into the non - saturation region, which is beneficial to solving the problem of gradient disappearance.
[0097] The process of normalizing each data in the input features is shown in formula (1 - 3), where x′ represents the normalized data, x represents the data before normalization, μ is the mean of all data in the input features, and σ 2 is the variance of all data in the input features, and ξ is a number very close to 0, used to prevent the denominator from being 0.
[0098]
[0099] ③ Pooling layer
[0100] After the input features are calculated by the convolutional layer, a two - dimensional matrix representing the output features will be obtained. However, not all information in this two - dimensional matrix is useful. The useless information will bring useless parameters and affect the calculation efficiency of the CNN. Therefore, it is necessary to use the pooling layer to eliminate some useless information, which will not only improve the calculation efficiency of the CNN, but also reduce the risk of overfitting of the CNN to a certain extent. Max - pooling is one of the most commonly used pooling methods. Max - pooling takes the maximum value of the feature values in the feature block as the pooling result of the feature block and inputs it into the output features.
[0101] After passing through the pooling layer, the size of the features will be further reduced. Assuming that the size of the input features of the pooling layer is H input ×W input , the size of the pooling kernel is E h ×E w , then the size of the output features after pooling can be calculated by formula (1 - 4). Where H output represents the height of the output features, W output represents the width of the output features, P represents the size of the boundary padding for the input features, S represents the step size of the pooling kernel's movement each time, represents rounding up.
[0102]
[0103] ④Activation function layer
[0104] Whether in the convolutional layer or the pooling layer, the finally obtained output features are still a linear combination of the input features. However, in practical problems, the input and output are often not simple linear combinations. Therefore, a non-linear activation function is needed to perform non-linear calculations on the input features to improve the expressive power of the neural network. Commonly used activation functions include the Sigmoid function, the Tanh function, and the ReLU function. The mathematical expressions of the three activation functions are shown in formula (1-5).
[0105]
[0106] The Sigmoid function and the Tanh function are relatively similar. The Sigmoid function can non-linearly transform any data in the input features to between 0 and 1, and the Tanh function can non-linearly transform any data in the input features to between -1 and 1. However, as the variable x becomes larger and larger, the gradients of the Sigmoid function and the Tanh function will gradually approach 0, resulting in the phenomenon of gradient disappearance, which is not conducive to the training of the neural network. The ReLU function can non-linearly transform any data in the input features to 0 or itself (depending on the size of the data itself), and there will be no phenomenon of gradient disappearance. Therefore, the ReLU function is used as the non-linear activation function.
[0107] ⑤Fc layer
[0108] After the input features are calculated by the Conv_Block in the CNN model, an output feature in the form of a two-dimensional matrix will be output. In the output feature, each feature value contains rich local feature information. In order to make full use of this information, an Fc layer is often added near the end of the neural network model. The Fc layer uses a fully connected method to connect the nodes in the previous network layer with the nodes in the current network layer, and then converts the input features into a one-dimensional matrix through weighted summation.
[0109] ⑥Dropout layer
[0110] After adding the Fc layer to the CNN model, a large number of parameters are introduced, which increases the risk of overfitting and also slows down the convergence speed of the neural network. To solve the problems brought by the Fc layer, a Dropout layer is often added between two Fc layers. The Dropout layer temporarily removes the nodes of the original network layer with a certain probability. In this way, when updating the parameters, it can be considered that this node does not exist, thereby reducing the influence of this node on the weight update. When updating the parameters next time, the nodes deleted last time will be restored, and at the same time other nodes will be temporarily removed. Repeating this operation is beneficial to improving the convergence speed of the model and alleviating the overfitting problem.
[0111] ⑦ Loss function layer
[0112] To make the output value of the CNN model closer to the label, a loss function needs to be introduced. The loss function is used to represent the difference between the final output value of the CNN model and the label value. During the process of updating the parameters of the CNN model, the value of the loss function will also keep decreasing. When the value of the loss function no longer decreases, it can be considered that the CNN model training is completed. The MSE is used as the loss function in the built CNN model, as shown in Equation (1-6). In the equation, N represents the number of samples in the dataset, y′ n is the output value of the CNN model, and y″ n is the label value.
[0113]
[0114] ⑧ Output layer normalization
[0115] After a series of feature processing operations of the CNN, the output layer normalization of the Transformer takes effect to ensure the stability and training efficiency of the CNN-Transformer model. Finally, the output is the workpiece sequence and machine sequence of the rescheduling scheme calculated by the CNN-Transformer model.
[0116] Case verification
[0117] The test case verification is divided into two parts. The first part is the verification of the effectiveness of the scheduler. A case adapted from a standard case is designed, and different methods are used to solve this case to verify the effectiveness of the scheduler. The second part is the verification of the generality of the scheduler. Ten random cases are designed, and different methods are used to solve these random cases to verify the generality of the scheduler.
[0118] (1) Verification of scheduler effectiveness
[0119] To verify the effect of the dynamic flexible job shop scheduler based on the convolutional transformation network, a case is designed. The initial scheduling scheme in this case is as Figure 3As shown, the dynamically arriving workpiece J 16 The processing time information comes from the 10th workpiece in the MK04 example. Solve the above example using a scheduler. The input of the scheduler is the scheduling attribute values of each process at the rescheduling moment, and the output is the workpiece processing sequence calculated by the scheduler and its corresponding machine selection sequence. The rescheduling plan obtained by this scheduler is as shown in Figure 4 shown. Figure 4 The black dotted line in it indicates the rescheduling moment 25.
[0120] From Figure 4 it can be seen that the makespan of the scheduling plan generated by the dynamic flexible job shop scheduler based on the convolutional transformation network is 72. The dynamically arriving J 16 and the unprocessed processes of the original J1-J 15 are rescheduled for processing after the rescheduling moment, so this scheduling plan is feasible.
[0121] To compare the advantages and disadvantages of this scheduler with other methods, use the dynamic flexible job shop scheduler based on the convolutional transformation network, the dynamic flexible job shop scheduler based on a single CNN, and scheduling rules to solve the above example respectively. There are many types of scheduling rules, and the scheduling rules adopted are "LWT×FIFO", "LWT×SPT", and "LWT×FDPNR". The results are shown in Table 2.
[0122] Table 2 Verification of Scheduler Effectiveness
[0123]
[0124] The makespan of the scheduling plan obtained by using the dynamic flexible job shop scheduler based on CNN-Transformer is shorter than that of the dynamic flexible job shop scheduler based on a single CNN, and the running time is faster, indicating that the scheduler can better solve the trained problems. For the scheduling problem of dynamically arriving workpieces, the dynamic flexible job shop scheduler based on CNN-Transformer can quickly obtain a scheduling plan with better results and is more adaptable to the actual production environment.
[0125] (2) Verification of Scheduler Generalization
[0126] To further verify the generality of the proposed method, a dynamic flexible job shop scheduler based on CNN-Transformer, a dynamic flexible job shop scheduler based on CNN, and scheduling rules were used to solve ten random instances. Different methods were used to solve the random instances. The input of the dynamic flexible job shop scheduler based on CNN-Transformer was the scheduling attribute values of each operation at the rescheduling moment, and the output was the workpiece processing sequence calculated by the scheduler and its corresponding machine selection sequence. The makespan obtained by each method is shown in Table 3, where C1 in Table 3 represents the first random instance.
[0127] Table 3 Verification of Scheduler Generality
[0128]
[0129] The result obtained by using the dynamic flexible job shop scheduler based on CNN-Transformer was 2.8% shorter than the average of the results obtained by the dynamic flexible job shop scheduler based on CNN. That is, when solving randomly generated instances, this scheduler could effectively utilize the learned scheduling knowledge to solve the DFJSP.
[0130] The makespan of the scheduling plan obtained by using the dynamic flexible job shop scheduler based on CNN-Transformer was 17.2% shorter than the average of the makespans of the scheduling plans obtained by using the scheduling rules. This shows that by using the knowledge obtained through training with this scheduler, a better scheduling plan than the ordinary scheduling rules could be obtained. Therefore, this scheduler could solve random instances more effectively than the scheduling rules.
Claims
1. A dynamic flexible job shop scheduler based on CNN-Transformer, characterized by: include: The input layer is used to select the scheduling attribute values of each workpiece on each machine in each process from the scheduling plan, and convert the scheduling attribute values into input features in the deep learning model training data; The CNN-Transformer model includes a CNN encoder and a Transformer module, wherein: the CNN encoder extracts deeper features from input features through a convolutional layer; the Transformer module uses a self-attention mechanism to map the deeper features extracted by the CNN encoder from a high-dimensional space to a low-dimensional space, and decodes and outputs the process sequence and machine sequence of each workpiece of the rescheduling plan through a channel multi-layer perceptron; The output layer is used to output the process sequence and machine sequence of each workpiece in the rescheduling plan.
2. The dynamic flexible job shop scheduler based on CNN-Transformer according to claim 1, characterized in that: The scheduling attribute values include: workpiece number, processing status of the process at the time of rescheduling, processing time of the process on each machine, processing machine corresponding to the process, order of processing of the process on the machine, remaining processing time of the process being processed, start time of the machine where the process is located, cumulative processing time of the machine where the process is located, and utilization rate of the machine where the process is located.
3. The dynamic flexible job shop scheduler based on CNN-Transformer according to claim 1, characterized in that: The CNN-Transformer model is a CNN encoder with a Transformer module inserted, and the Transformer module serves as a decoder. The CNN encoder consists of a convolutional layer, a pooling layer, a normalization layer, an activation function layer, an Fc layer, and a Dropout layer. The Transformer module consists of six Transformer_Blocks, each of which corresponds to a layer in the CNN encoder. Each Transformer_Block consists of a token mixer, a channel multilayer perceptron (MLP), a learnable scaling factor (scale), and two layer normalizations (Layer Norm). The token mixer is used to enrich the feature representation of the model; the channel multilayer perceptron is responsible for feature processing in the channel dimension; scale is used to balance the influence of the output of each part, and the two layer normalizations are used for input and output normalization, respectively.
4. The dynamic flexible job shop scheduler based on CNN-Transformer according to claim 3, characterized in that: The token mixer adopts a self-attention mechanism to capture the relationship between any two positions in the input sequence, thereby extracting complex dependencies within the sequence.
5. The dynamic flexible job shop scheduler based on CNN-Transformer according to claim 3, characterized in that: The channel multi-layer perceptron (MLP) includes two layers of fully connected networks for performing nonlinear transformation and dimensionality adjustment on features.
6. The CNN-Transformer-based dynamic flexible job shop scheduler according to claim 3, characterized in that: The learnable scaling factor (scale) is dynamically adjusted through training to optimize the weights of the outputs of each part in Transformer_Block.
7. The scheduling method of the dynamic flexible job shop scheduler based on CNN-Transformer according to any one of claims 1 to 6, characterized in that Follow these steps: (1) Selecting scheduling attribute values: The input layer selects the scheduling attribute values of each workpiece on each machine in each process from the scheduling plan; the scheduling attribute values include the workpiece number, the processing status of the process at the time of rescheduling, the processing time of the process on each machine, the processing machine corresponding to the process, the order of processing the process on the machine, the remaining processing time of the process being processed, the available start time of the machine where the process is located, the cumulative processing time of the machine where the process is located, and the utilization rate of the machine where the process is located; (2) Scheduling attribute value conversion: The scheduling attribute value obtained in step (1) is converted into input features in the deep learning model training data; wherein the columns of the input features represent each process, and the rows of the input features represent the scheduling attribute values of each process; (3) inputting the input features into the CNN-Transformer model, extracting deep features of the input features through the convolution layer, and performing dimension expansion, deep convolution and dimension reduction operations through the token mixer of the Transformer module to capture local features and complex dependencies within the sequence; (4) Output the workpiece sequence and machine sequence of the rescheduling plan calculated by the CNN-Transformer model.
8. The scheduling method of the dynamic flexible job shop scheduler based on CNN-Transformer according to claim 7, characterized in that: The label of the input feature is a one-dimensional matrix formed by concatenating the workpiece sequence and the machine sequence of the rescheduling plan.
9. The scheduling method of the dynamic flexible job shop scheduler based on CNN-Transformer according to claim 7, characterized in that: The calculation process of the input features and output features of the convolutional layer is shown in formula (1-1): in, is the data block with coordinates (row, col) in the th feature of the lay-th convolutional layer, f represents the activation function, Cha represents the number of channels of the feature, is the data block with coordinates (row, col) in the th feature of the lay-1th convolutional layer, wk lay Represents the convolution kernel of the lay-th convolution layer, bias lay Represents the bias of the lay-th convolutional layer.
10. The scheduling method of the dynamic flexible job shop scheduler based on CNN-Transformer according to claim 7, characterized in that: The token mixer of the Transformer module adopts a self-attention mechanism to capture complex dependencies within the sequence by computing the relationship between any two positions in the input sequence.
Citation Information
Patent Citations
Prediction model establishment method and prediction method for residual life of engineering mechanical part
CN114169091A
Large-scale fuzzy flexible job shop scheduling method and related equipment
CN118760100A
Medical image segmentation method and equipment based on multi-scale feature fusion
CN118898773A
Estimation method and system for daily order completion amount of online car-hailing platform
CN119006037A
Monocular depth estimation method based on CNN-Transform hybrid architecture
CN119478000A