Flexible job shop scheduling method and system based on evolutionary strategy and meta reinforcement learning

By adopting evolutionary strategies and meta-reinforcement learning methods in flexible work workshop scheduling, the problems of insufficient generalization capabilities and high computing costs in the existing technology are solved, and better scheduling performance and adaptability are achieved.

CN120044900AActive Publication Date: 2025-05-27SHANDONG UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510145958.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-27
Estimated Expiration
2045-02-10

AI Technical Summary

Technical Problem

The existing deep reinforcement learning methods have problems such as insufficient generalization capability and high computational cost in the scheduling problems of flexible work workshops. Especially when facing specific instances, the actual performance of the model may be greatly deviated from the ideal return.

Method used

A flexible work workshop scheduling method based on evolutionary strategies and meta-reinforcement learning is adopted to randomly generate training data sets of different instances, a meta-reinforcement learning framework is constructed, meta-reinforcement learning uses evolutionary strategies to optimize meta-model parameters, and each instance is fine-tuned for a limited number of times during the inference process to meet the needs of a specific instance.

Benefits of technology

The generalization ability and practical application effect of the scheduling system on different instances is improved, and the scheduling performance is optimized, so that the system can automatically adjust decision strategies under different production environments and machine configurations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120044900A_ABST
    Figure CN120044900A_ABST
Patent Text Reader

Abstract

The invention provides a flexible job shop scheduling method and system based on an evolutionary strategy and meta reinforcement learning, and the method comprises the steps: randomly generating a set number of flexible job shop scheduling problem instances, forming a training data set, and replacing the training data set every fixed update round; constructing a meta-reinforcement learning framework based on an evolutionary strategy for training a meta-model, and determining an optimal parameter of the meta-model by minimizing a verification set total completion average time to serve as an initialization model for adapting to a new task in a reasoning process; and obtaining the completion time of the test data by using the trained meta-model, and performing finite fine tuning on each instance in the test data to obtain an optimal result for each instance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of flexible job shop scheduling, and in particular relates to a flexible job shop scheduling method and system based on evolutionary strategy and meta-reinforcement learning. Background Art

[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.

[0003] The development of industrial technology is changing the way companies manufacture and distribute products, moving towards fast, intelligent and flexible manufacturing, leading to fundamental changes in the production capacity of companies. The flexible job shop scheduling problem (FJSP) is a classic problem that represents a typical scenario faced by flexible manufacturing. It allows each operation to be processed on multiple different machines. The machine allocation problem increases manufacturing flexibility, giving FJSP a more complex topology and a larger solution space. It has been proven that FJSP is a strong NP (Nondeterministic Polynomial time problems) hard problem. The combinatorial nature of FJSP makes it challenging to find (close to) optimal solutions using traditional operations research methods such as constraint programming. The computational cost of these methods is intractable and increases dramatically with the size of the problem, making them unsuitable for large-scale applications. In order to strike a balance between solution quality and computational cost, research in this field has gradually shifted from traditional heuristic and meta-heuristic methods to intelligent methods such as data-driven deep learning and deep reinforcement learning (DRL).

[0004] Metaheuristic algorithms, including genetic algorithms, particle swarm optimization, differential evolution, and artificial bee colonies, have been widely used in scheduling problems, and they usually find high-quality solutions through complex solution search procedures. In contrast, rule-based heuristic methods, such as priority scheduling rules (PDRs), are more practical due to their ease of implementation and high efficiency. PDRs repeatedly select the highest priority operations or machines according to some prescribed rules until a complete schedule is generated. Nevertheless, designing effective PDRs usually requires a lot of expertise and research efforts, and they may only perform well in specific tasks. Currently, DRL methods have become a promising approach to solving FJSP, which model the scheduling process as a Markov decision process (MDP). In these methods, a parameterized neural network model is designed to receive information about the production environment as a state and output the priority of each feasible scheduling operation, such as assigning operations to machines, forming an end-to-end learning method. By training on a set of production process data, the DRL model learns to adaptively select the best action in a certain state to maximize the total reward associated with the production goal.

[0005] However, these methods generally have some limitations. Usually, the optimal model obtained by training with a large number of training sets can achieve the best average performance under a certain size or distribution. However, this training method often ignores the particularity of a single instance of the scheduling problem. Therefore, although the model can achieve good global performance in a statistical sense, when faced with a specific instance, the actual performance of the model may deviate greatly from the ideal return. This phenomenon shows that the existing DRL methods rely too much on the extensiveness of the training set and fail to fully consider the unique needs and scheduling characteristics of individual instances, thereby limiting their generalization ability and practical application effects in complex and dynamic environments. To this end, the algorithm framework optimized for individual instances, especially the introduction of attention to instance particularity during model training, has become a key research direction to improve the application effect of DRL in FJSP.

[0006] Existing patents disclose solving the problem of flexible job shop scheduling through heuristic methods. However, heuristic methods require a large number of manual rules and domain expertise for each specific problem, and they may only perform well in specific tasks. Currently, methods based on deep reinforcement learning are more promising, however, these methods generally have some limitations. Usually, the optimal model obtained by training with a large number of training sets can achieve the best average performance under a certain size or distribution. However, this training method often ignores the particularity of a single instance of the scheduling problem. Summary of the invention

[0007] To overcome the deficiencies of the above-mentioned prior art, the present invention provides a flexible job shop scheduling method based on evolutionary strategy and meta-reinforcement learning, which fully considers the unique requirements and scheduling characteristics of different instances to further minimize the total completion time.

[0008] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions:

[0009] In the first aspect, a flexible job shop scheduling method based on evolutionary strategy and meta-reinforcement learning is disclosed, including:

[0010] A set number of flexible job shop scheduling problem instances are randomly generated to form a training data set, and the training data set is replaced every fixed update round;

[0011] A meta-reinforcement learning framework based on evolutionary strategies is constructed. The training dataset is used to train the meta-model. The optimal parameters of the meta-model are determined by minimizing the average completion time of the validation set. The meta-model is used as the initialization model to adapt to new tasks during the reasoning process.

[0012] Using the trained meta-model, we get the completion time of the test data, and perform a finite number of fine-tuning on each instance in the test data to get the optimal result for each instance.

[0013] The above-mentioned minimization of the average total completion time of the validation set is to minimize the average total completion time of all instances in the validation set.

[0014] As a further technical solution, the metamodel training process includes inner loop optimization and outer loop optimization in each iteration, the inner loop optimization iteratively optimizes the task-specific model, and the outer loop optimization updates the metamodel, with the aim of maximizing the few-sample generalization performance of the task-specific model.

[0015] As a further technical solution, the training data set includes:

[0016] The set consisting of n jobs and m machines is represented by J and M respectively;

[0017] Assume that all jobs are produced in the system at time T s = 0 arrives at the same time, each job J i ∈J contains n i operations that must be assembled in a specific order, consisting of Description: The set of all operations of all jobs is represented by O, and each operation O ij Processed by multiple machines, but only on the set M of available and compatible machines ij ∈M is processed on a machine;

[0018] Machine M k ∈M ij The relevant processing time on Given.

[0019] As a further technical solution, the flexible job shop scheduling problem is used to determine the appropriate processing machine and start time for each operation while complying with the following constraints:

[0020] 1.J i The operation must be in accordance with O i Sequential processing in

[0021] 2. Each operation must be assigned to exactly one compatible machine;

[0022] 3. Each machine can process at most one operation at a time;

[0023] The goal of FJSP is to minimize the maximum completion time of all jobs, that is, to minimize the total completion time.

[0024] As a further technical solution, the solution process of the FJSP instance G is formulated as an MDP, and the strategy is parameterized by a neural network constructed by an encoder-decoder to learn the node selection for constructing the solution;

[0025] The encoder in the policy network outputs a global representation of the instance, which together with the representation of the context captures the current state. The decoder takes the global and context representations as input to compute the probability of visiting nodes, i.e., actions, selecting them in turn until a complete journey τ is constructed.

[0026] As a further technical solution, the meta-goal is defined as follows:

[0027]

[0028] in, In the instance T i Upper θ 0 Fine-tuning the model after K gradient updates; is an instance of T i The loss function on , uses the same loss function for different instances.

[0029] Secondly, a flexible job shop scheduling system based on evolutionary strategy and meta-reinforcement learning is disclosed, including:

[0030] The training data set construction module is configured to: randomly generate a set number of flexible job shop scheduling problem instances to form a training data set, and replace the training data set every fixed update round;

[0031] The meta-model training module is configured to: construct a meta-reinforcement learning framework based on evolutionary strategies to train the meta-model, determine the optimal parameters of the meta-model by minimizing the average completion time of the verification set, and use it as an initialization model to adapt to new tasks during the reasoning process;

[0032] The scheduling module is configured to: use the trained meta-model to obtain the completion time of the test data, and perform a limited number of fine-tuning on each instance in the test data to obtain the optimal result for each instance.

[0033] One or more of the above technical solutions have the following beneficial effects:

[0034] The technical solution of the present invention improves its generalization ability and practical application effect for different instances by fully considering the unique needs and scheduling characteristics of individual instances. This meta-learning-based strategy enables the scheduling system to automatically adjust its decision-making strategy when facing different production environments, machine configurations, and job characteristics, thereby optimizing scheduling performance. At the same time, the present invention replaces the optimization algorithm based on policy gradients with an evolutionary strategy algorithm. As a gradient-free optimization method, evolutionary strategy has significant parallelization advantages and can efficiently and independently optimize multiple scheduling instances at the same time in a single-card GPU environment. And during the training process, the evolutionary strategy uses global search rather than local gradient updates, so that the optimization process is not easy to fall into the local optimal solution, thereby enhancing the global generalization ability of the model and further improving the training efficiency and fine-tuning performance of the metamodel.

[0035] Advantages of additional aspects of the present invention will be given in part in the following description, and in part will become obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0037] Figure 1 This is a meta-learning framework diagram for an embodiment of the present invention;

[0038] Figure 2 It is a parallel computation graph of the evolution strategy according to an embodiment of the present invention;

[0039] Figure 3 The figure is a flow chart of a method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0040] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.

[0041] It should be noted that the terms used herein are for describing specific embodiments only and are not intended to be limiting of exemplary embodiments according to the present invention.

[0042] In the absence of conflict, the embodiments of the present invention and the features of the embodiments may be combined with each other.

[0043] Embodiment 1

[0044] This embodiment discloses a flexible job shop scheduling method based on evolutionary strategy and meta-reinforcement learning, including:

[0045] Step 1: Randomly generate 20 FJSP instances to form a training data set, and replace them every fixed update round. That is, every fixed update round, 20 instances will be regenerated to replace the data in the previous training set.

[0046] In this implementation example, the training data set includes 20 FJSP (flexible shop scheduling) instances, each of which represents a specific shop situation, namely the number of workpieces, the number of machines, the number of operations for each workpiece, and the processing time of each operation on different machines.

[0047] Step 2: Generate 100 FJSP instances to form a verification data set and a test data set;

[0048] Step 3: Construct a meta-reinforcement learning framework based on evolutionary strategies, with the goal of training a meta-model as a well-initialized model that can effectively adapt to new tasks during reasoning. The meta-training process includes inner-loop and outer-loop optimizations for each iteration. The inner-loop optimization iteratively optimizes the task-specific model, which is similar to the fine-tuning phase during reasoning. The outer-loop optimization updates the meta-model with the goal of maximizing the few-shot generalization performance of the task-specific model.

[0049] Train the meta-model and determine the optimal parameters of the meta-model by minimizing the average completion time of the verification set;

[0050] The meta-model is a deep neural network with tens of thousands of parameters. These parameters are updated through back-propagation, so that the output of the model gradually improves. Finally, the model parameters corresponding to the best output result are taken as the optimal parameters.

[0051] Step 4: Use the trained meta-model to obtain the completion time of the test data, and perform a finite number of fine-tuning on each instance in the test data to obtain the optimal result for each instance, that is, the total time required to complete the workshop scheduling.

[0052] For each instance, the workshop information represented by this instance (number of workpieces, number of machines, number of operations for each workpiece, and processing time of each operation on the machine) is embedded into a vector, which is the input data of the model. After the vector is input into the model, a series of calculations are performed to finally obtain the probability of the next action, that is, matching the workpiece and the machine. After matching, the workshop information changes and is embedded into a new vector. This cycle is repeated to obtain a complete workshop scheduling process. The output result is the total time to complete the workshop scheduling.

[0053] The fine-tuning process is the same as the training process, where the model is back-propagated and updated for a single instance, but the update amplitude is reduced.

[0054] The flexible job shop scheduling method based on evolutionary strategy and meta-reinforcement learning in this example improves its generalization ability and practical application effect for different instances by fully considering the unique needs and scheduling characteristics of individual instances. This meta-learning-based strategy enables the scheduling system to automatically adjust its decision-making strategy when facing different production environments, machine configurations, and job characteristics, thereby optimizing scheduling performance. Therefore, the method of the present invention has a good application prospect.

[0055] When the technical solution of the present application is implemented, a limited number of fine-tuning is performed on the specific production environment, machine configuration and operating characteristics, so that the model can adapt to the specific workshop conditions and output a scheduling result that is better than the original model.

[0056] FJSP can be formally stated as follows. Consider a set of n jobs and m machines, denoted by J and M respectively. Assume that all jobs are produced in the system at time T s = 0 arrive at the same time. Each job J i ∈J contains n i operations, J represents the set of operations, J i and n i The i in represents the i-th operation in the set. These operations must be assembled in a specific order (i.e., priority constraints), as given by Description. The set of all operations of all jobs is represented by O. Each operation O ij Can be processed by multiple machines, but only on a set of available and compatible machines M ij ∈M. Machine M k ∈M ij The relevant processing time on Given, i is the job number, j is the operation number in the job, for example, Oij refers to the jth operation of the ith job. K is the machine number, p ij k It means the time required for the kth machine to process the jth operation of the ith job.

[0057] In this implementation, FJSP attempts to design a plan that determines the appropriate processing machine and start time for each operation while respecting the following constraints:

[0058] 1.J i The operation must be in accordance with O i Sequential processing in (i.e., precedence constraints);

[0059] 2. Each operation must be assigned to exactly one compatible machine;

[0060] 3. Each machine can process at most one operation at a time. The goal of FJSP is to minimize the maximum completion time of all jobs, that is, to minimize the total completion time.

[0061] The scheduling process can be understood as the dynamic allocation of ready operations to compatible idle machines. In this way, the decision point t is the production system time T S (t), at this time there is at least one compatible operating machine pair (O ij , M k ), so that O ij At time T S (t) On machine M k is processed.

[0062] The decision point is the node where the job and the machine need to be paired during the scheduling process. It is equivalent to step t. Because operations have processing time on the machine, each decision point has at least one node that is compatible with the machine operation.

[0063] At step t, for example, a scheduling process takes 50 steps to complete, then t is 1 to 50, and the DRL model receives the state s from the environment t , and take action to specify a compatible pair at time T S (t) starts processing immediately, Ts(t) refers to the time that the workshop has scheduled at step t, for which the FJSP environment returns a reward r related to the maximum completion time t By repeating this process |O| times for the entire set of operations in the task, we can obtain the path τ of the solution of FJSP.

[0064] Status S t Refers to: the number of workshop jobs, the number of operations for each job, the number of machines, the processing time of each operation on the machine, the jobs and machines that have been scheduled and completed in the workshop, and the jobs and machines that have not been scheduled in the workshop. The status of the workshop changes continuously with the scheduling. Therefore, the current status of the workshop must be received at each scheduling.

[0065] The cost function c(·) considered in this example is the total completion time after scheduling is completed. The goal of FJSP is to find the optimal path τ with the shortest total completion time. * .

[0066]

[0067] Where S is a discrete search space containing all feasible paths subject to the constraints of the specific problem, and G is a single instance.

[0068] The neural construction method formulates the solution process of the FJSP instance G as an MDP, parameterizing the policy through an encoder-decoder neural network to learn the node selection for constructing the solution. The encoder in the policy network outputs a global representation of the instance, which captures the current state together with the representation of the context (e.g., scheduled operations and machines). The decoder takes the global and contextual representations as input to calculate the probability of the node to be visited (i.e., action). The node is selected in turn until the complete journey τ is constructed. Therefore, the probability of the path is decomposed into:

[0069]

[0070] Among them, π θ (t) and π θ (<t) are the selected node and the current partial solution at time step t, respectively; T represents the total number of steps. The reward is defined as the negative cost of a trip, i.e.

[0071] The purpose of formula (2) is to calculate the probability of each possible scheduling under the condition of fixed model parameters. By knowing the probability of each path, we can obtain the expected return of the model and optimize the model according to the expectation.

[0072] To train the policy network, a reinforcement learning algorithm is usually used to estimate the expected reward Gradient

[0073] Define a task T{(n*m), D} as a class of instances of size (n*m) and distribution d∈D, where n and m are the number of operations and the number of machines respectively; is a collection of distributions; Ti is used to represent an instance. A general meta-learning framework is applied to improve the generalization ability of the VRP neural method, which is model-independent and compatible with any model trained with gradient updates. The framework is as follows Figure 1 shown.

[0074] In the meta-learning framework, the goal is to train a meta-model θ 0 , as a well-initialized model, it can effectively adapt to new tasks during reasoning. Formally, the meta-objective is defined as follows:

[0075]

[0076] in In the instance T i Upper θ 0 Fine-tuning the model after K gradient updates; is an instance of T i The loss function on , E represents the expectation.

[0077] In this implementation example, the same loss function (e.g., reinforcement loss) is used for different instances. In order to directly optimize this goal, the meta-training process includes inner-loop and outer-loop optimizations for each iteration, specifically including:

[0078] Step 1: Initialize the metamodel;

[0079] Step 2: Randomly generate instances;

[0080] Step 3: Inner loop simulation fine-tuning;

[0081] Step 4: The outer loop updates the metamodel;

[0082] Step 5: Repeat steps 3 and 4 until the training is completed.

[0083] The meta-training pseudo code is as follows.

[0084]

[0085] Inner loop optimization: It iteratively optimizes an instance-specific model, similar to the fine-tuning stage during inference. Specifically, given an instance T i ∈T, initialize the instance-specific model through the metamodel, that is, And adapt it to Ti by performing K gradient update steps on the training instance. The k-th step loss function The gradient of is calculated as follows:

[0086]

[0087] Formula (4) calculates the inner loop optimization gradient and simulates fine-tuning to update the model of the corresponding instance.

[0088] Outer loop optimization: It uses the objective in (3) to optimize the metamodel. Specifically, for each instance T i , evaluate task-specific models The meta-gradient is obtained as follows:

[0089]

[0090] Formula (5) updates the meta-model by calculating the outer loop optimization gradient.

[0091] In a batch of tasks After inner loop optimization, the metamodel θ 0 Intuitively, the inner-loop optimization acts as a task adaptation phase, mimicking the fine-tuning process during inference, while the outer-loop optimization updates the meta-model with the goal of maximizing the instance-specific model The few-shot generalization performance of

[0092] Therefore, after meta-training, a good initialization model can be obtained The model can effectively adapt to new instances using only limited data. Note that Equation (5) requires second-order derivatives because it is desired to obtain 0 The gradient direction of . Its first-order approximation method will be introduced.

[0093] The meta-gradient in the equation involves the gradient of the gradient (i.e., the second-order derivative), so it is computationally expensive to obtain due to the calculation of the Hessian vector product. To address this issue, a first-order approximation method is applied, which simply removes the second-order terms. Empirical evidence of its effectiveness has been verified in few-shot supervised learning. Specifically, the first-order approximation of the meta-model update can be expressed as:

[0094]

[0095] In the process of solving the flexible job shop scheduling problem, the optimization method based on policy gradient is usually used, which can update the policy parameters through local gradient information. However, under the meta-reinforcement learning framework, the optimization algorithm based on policy gradient faces high computational overhead when training multiple scheduling instances, especially in a single-card GPU environment. Specifically, the optimization algorithm based on policy gradient relies on back-propagation to calculate the gradient, and usually needs to update the policy of each instance within multiple training cycles. This gradient-dependent iterative optimization method is not only computationally inefficient, but also difficult to fully utilize parallel computing resources.

[0096] In order to solve the above problems, the policy gradient-based optimization algorithm is replaced by the evolution strategy algorithm (ES). As a non-gradient optimization method, the evolution strategy optimizes the parameter population by simulating the process of natural selection. Its core idea is to perturb multiple candidate solutions at the same time in each generation and evaluate their fitness, thereby guiding the optimization of model parameters through population-level updates. This method has significant parallelization advantages and can efficiently optimize multiple scheduling instances independently at the same time in a single-card GPU environment. The parallel computing framework diagram is shown in the figure below. Figure 2 shown.

[0097] During the training process, the evolutionary strategy uses global search rather than local gradient updates, making it less likely for the optimization process to fall into the local optimal solution, thereby enhancing the global generalization ability of the model and further improving the fine-tuning performance of the meta-model.

[0098] The evolution strategy gradient estimate is obtained as follows:

[0099]

[0100] Among them, ∈ iis the noise disturbance, which satisfies Gaussian distribution; σ is the standard deviation of the noise; n is the number of evolutionary populations.

[0101] The evolution strategy updates the model parameters in two repeated stages:

[0102] 1. Randomly perturb the parameters of the strategy to generate n populations and evaluate the resulting parameters by running one round in the environment;

[0103] 2. Combine the results of all populations, calculate the stochastic gradient estimate, and update the parameters.

[0104] The evolutionary strategy of this example is not to calculate the gradient directly, but to obtain the population by applying perturbations to the model and simulate the gradient calculation through the population.

[0105] The pseudo code is as follows:

[0106]

[0107] The sub-technical solution of this embodiment applies a meta-learning framework in the process of training the model. The purpose is to further optimize the scheduling effect by fine-tuning a single instance after the model training is completed. This reflects the "considering the unique needs and scheduling characteristics of individual instances". The specific approach is: during the training process, there are inner loops and outer loops, where the inner loop simulates fine-tuning for a single instance, and the outer loop updates the meta-model. The role of the inner loop is to consider the unique needs and scheduling characteristics of individual instances. The inner loop part of meta-learning is to introduce attention to the particularity of instances.

[0108] See attached Figure 3 As shown in the specific example, in the flexible job shop scheduling method based on evolutionary strategy and meta-reinforcement learning, it includes:

[0109] Step 1: Generate 100 instances as a validation set, and randomly generate 20 instances as a training set every fixed round.

[0110] Step 2: Initialize the metamodel.

[0111] Step 3: Assign the initial parameters of the meta-model to the model corresponding to each instance in the training set, and perform inner loop simulation fine-tuning on the model.

[0112] Step 4: The inner loop ends and the outer loop optimizes the metamodel.

[0113] Step 5: After training, select the best model through the validation set as the final result.

[0114] Step 6: Input the actual engineering examples into the meta-model.

[0115] Step 7: Fine-tune the instance a limited number of times to make the model better fit the specific instance.

[0116] Step 8: Select the best scheduling solution during fine-tuning.

[0117] Embodiment 2

[0118] The purpose of this embodiment is to provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the program.

[0119] Embodiment 3

[0120] The purpose of this embodiment is to provide a computer-readable storage medium.

[0121] A computer-readable storage medium stores a computer program, which executes the steps of the above method when executed by a processor.

[0122] Embodiment 4

[0123] The purpose of this embodiment is to provide a flexible job shop scheduling system based on evolutionary strategy and meta-reinforcement learning, including:

[0124] The training data set construction module is configured to: randomly generate a set number of flexible job shop scheduling problem instances to form a training data set, and replace the training data set every fixed update round;

[0125] The meta-model training module is configured to: construct a meta-reinforcement learning framework based on evolutionary strategies to train the meta-model, determine the optimal parameters of the meta-model by minimizing the average completion time of the verification set, and use it as an initialization model to adapt to new tasks during the reasoning process;

[0126] The scheduling module is configured to: use the trained meta-model to obtain the completion time of the test data, and perform a limited number of fine-tuning on each instance in the test data to obtain the optimal result for each instance.

[0127] Embodiment 5

[0128] The purpose of this embodiment is to provide a computer program product containing instructions, which, when running on a computer, enables the computer to execute the methods and functions involved in any of the above embodiments.

[0129] The steps involved in the apparatus of the above embodiment correspond to the method embodiment 1, and the specific implementation method can refer to the relevant description part of embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood to include any medium that can store, encode or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.

[0130] Those skilled in the art should understand that the modules or steps of the present invention described above can be implemented by a general-purpose computer device, or alternatively, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.

[0131] Although the above describes the specific implementation mode of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without creative work are still within the scope of protection of the present invention.

Claims

1. A flexible job shop scheduling method based on evolutionary strategy and meta-reinforcement learning, characterized by including: A set number of flexible job shop scheduling problem instances are randomly generated to form a training data set, and the training data set is replaced every fixed update round; A meta-reinforcement learning framework based on evolutionary strategies is constructed to train the meta-model. The optimal parameters of the meta-model are determined by minimizing the average completion time of the validation set. This is used as the initialization model to adapt to new tasks during the reasoning process. Using the trained meta-model, we get the completion time of the test data, and perform a finite number of fine-tuning on each instance in the test data to get the optimal result for each instance.

2. The flexible job shop scheduling method based on evolutionary strategy and meta-reinforcement learning as claimed in claim 1, characterized in that: The metamodel training process includes inner-loop optimization and outer-loop optimization in each iteration, wherein the inner-loop optimization iteratively optimizes the task-specific model, and the outer-loop optimization updates the metamodel, with the aim of maximizing the few-sample generalization performance of the task-specific model.

3. The flexible job shop scheduling method based on evolutionary strategy and meta-reinforcement learning as claimed in claim 1, characterized in that: The training dataset includes: The set consisting of n jobs and m machines is represented by J and M respectively; Assume that all jobs are produced in the system at time T s = 0 arrives at the same time, each job J i ∈J contains n i operations that must be assembled in a specific order, consisting of Description: The set of all operations of all jobs is represented by o, and each operation O ij Processed by multiple machines, but only on the set M of available and compatible machines ij ∈M is processed on a machine; Machine M k ∈M ij The relevant processing time on Given.

4. The flexible job shop scheduling method based on evolutionary strategy and meta-reinforcement learning as claimed in claim 1, characterized in that: The flexible job shop scheduling problem is used to determine the appropriate processing machine and start time for each operation while observing the following constraints: (1).J i The operation must be in accordance with O i Sequential processing in (2) Each operation must be assigned to exactly one compatible machine; (3) Each machine can process at most one operation at a time; The goal of FJSP is to minimize the maximum completion time of all jobs, that is, to minimize the total completion time.

5. The flexible job shop scheduling method based on evolutionary strategy and meta-reinforcement learning as claimed in claim 1, characterized in that: The solution process of the FJSP instance G is formulated as an MDP, and the policy is parameterized by a neural network constructed by an encoder-decoder to learn the node selection for constructing the solution; The encoder in the policy network outputs a global representation of the instance, which together with the representation of the context captures the current state. The decoder takes the global and context representations as input to compute the probability of visiting nodes, i.e., actions, selecting them in turn until a complete journey τ is constructed.

6. The flexible job shop scheduling method based on evolutionary strategy and meta-reinforcement learning as claimed in claim 1, characterized in that: The meta-goal is defined as follows: in, In the instance T i The fine-tuned model of θ0 after K gradient updates; is an instance of T i The loss function on , uses the same loss function for different instances.

7. A flexible job shop scheduling system based on evolutionary strategies and meta-reinforcement learning, characterized by: include: The training data set construction module is configured to: randomly generate a set number of flexible job shop scheduling problem instances to form a training data set, and replace the training data set every fixed update round; The meta-model training module is configured to: construct a meta-reinforcement learning framework based on evolutionary strategies to train the meta-model, determine the optimal parameters of the meta-model by minimizing the average completion time of the verification set, and use it as an initialization model to adapt to new tasks during the reasoning process; The scheduling module is configured to: use the trained meta-model to obtain the completion time of the test data, and perform a limited number of fine-tuning on each instance in the test data to obtain the optimal result for each instance.

8. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the method described in any one of claims 1 to 6 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method described in any one of claims 1 to 6 are performed.

Citation Information

Patent Citations

  • Electro-tricycle frame lightweight design method and system based on rigid-flexible coupling

    CN112035953A

  • Learning to schedule control fragments for physics-based character simulation and robots using deep q-learning

    US20180089553A1

  • Loss Function Optimization Using Taylor Series Expansion

    US20210089832A1

  • Meta-automated machine learning with improved multi-armed bandit algorithm for selecting and tuning a machine learning algorithm

    US20210224585A1

  • Systems and methods for black-box optimization

    WO2018222204A1