Single-machine scheduling method based on reinforcement learning enhanced genetic evolution

By constructing a deep reinforcement learning workpiece dispatch probability model and a genetic evolutionary algorithm, the initial solution set of the single-machine scheduling problem is optimized, which solves the problems of high computational complexity and local search falling into local optimality in the existing technology, and realizes efficient single-machine scheduling optimization.

CN120634075APending Publication Date: 2025-09-12NORTH CHINA ELECTRIC POWER UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510490376.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

When solving large-scale single-machine scheduling problems, existing technologies have high computational complexity and are difficult to guarantee the quality of solutions. Local search methods are prone to falling into local optimality and lack global search capabilities.

Method used

A reinforcement learning-enhanced genetic evolution method is adopted. By constructing a deep reinforcement learning job dispatch probability model and a genetic evolution algorithm, the initial solution set is optimized to improve the efficiency and quality of single-machine scheduling.

Benefits of technology

It reduces the search pressure of the optimization method, improves the search capability, takes into account both local and global search capabilities, and improves the efficiency and quality of solving single-machine scheduling problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120634075A_ABST
    Figure CN120634075A_ABST
Patent Text Reader

Abstract

The invention relates to a single-machine scheduling method based on reinforcement learning enhanced genetic evolution. The method comprises the following steps: setting problem parameters of single-machine total advance scheduling; constructing and optimizing a workpiece assignment probability model based on deep reinforcement learning, and inputting problem parameters of single machine total advance scheduling to obtain an initial solution set; and optimizing the initial solution set by using a genetic evolutionary algorithm to obtain an optimal single-machine scheduling scheme. Compared with a traditional method, the searching pressure of the optimization method can be relieved, the searching capability of the optimization method can be improved, and the solving efficiency and quality of the single-machine scheduling problem can be improved. Meanwhile, population hybrid initialization and population hybrid updating operation based on a workpiece assignment probability model are adopted, and through the hybrid population initialization operation, the population diversity can be guaranteed, and the initial population quality can be improved; and population mixed population updating operation is adopted, so that population evolution can be ensured, and gene diversity can be increased. Compared with a traditional method, it can be ensured that the algorithm considers both local search and global search capabilities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of production scheduling control, and in particular to a single-machine scheduling method based on reinforcement learning enhanced genetic evolution. Background Art

[0002] Scheduling problem, also known as sorting problem, refers to the process of processing several workpieces (jobs) on some machines and reasonably arranging the machines and workpieces to optimize the objective function.

[0003] The scheduling mentioned in the "single machine scheduling" problem refers to the scheduling of jobs and the order in which machines operate. It does not include the scheduling of vehicles, personnel, routes, and time.

[0004] The paper "Variable neighborhood search for the single machine scheduling problem to minimize the total early work. Optim Lett, 2023, 17, 2169-2184." proposes a heuristic algorithm based on a variable neighborhood search method. This method can solve larger-scale single-machine scheduling problems than the dynamic programming method. However, the method proposed in the article is still a local search method and is prone to falling into local optimality. When solving large-scale single-machine scheduling problems, this method not only has high computational complexity, but also has difficulty in ensuring the quality of the solution to the problem. To improve the solution quality of single-machine scheduling problems, an efficient optimization method that takes into account both global and local searches is needed. It has both good local search capabilities and excellent global search capabilities to escape local optimality. Summary of the Invention

[0005] The purpose of the present invention is to overcome the shortcomings of the existing technology and propose a single-machine scheduling method based on reinforcement learning enhanced genetic evolution, which can reduce the search pressure of the optimization method and improve the search ability of the optimization method, thereby improving the efficiency and quality of solving the single-machine scheduling problem.

[0006] The present invention solves the technical problem by adopting the following technical solutions:

[0007] A single-machine scheduling method based on reinforcement learning enhanced genetic evolution includes the following steps:

[0008] Step 1: Set the problem parameters for single-machine total lead time scheduling;

[0009] Step 2: Build and optimize a job dispatch probability model based on deep reinforcement learning, input the problem parameters of single-machine total lead time scheduling to obtain the initial solution set;

[0010] Step 3: Use the genetic evolutionary algorithm to optimize the initial solution set and obtain the optimal single-machine scheduling solution.

[0011] Moreover, the problem parameters of the single machine total lead time scheduling in step 1 include the number of workpieces n g , the release time r of the workpiece i ,i=1,2,...,n g , the processing time p of the workpiece i ,i=1,2,...,n g , the delivery time of the workpiece d i ,i=1,2,...,n g .

[0012] Furthermore, the step 2 includes the following steps:

[0013] Step 2.1: Set up a reinforcement learning environment for single-machine total lead time scheduling.

[0014] Step 2.2: Build a workpiece dispatch probability model based on deep reinforcement learning;

[0015] Step 2.3: Optimize the workpiece dispatch probability model.

[0016] Furthermore, the step 2.1 includes the following steps:

[0017] Step 2.1.1. Set the state S(t) of the deep reinforcement learning environment for single-machine total lead time scheduling at time step t:

[0018]

[0019] Among them, s i (t),i=1,2,...,n g represents the state of workpiece i at time step t, l i =0 means job i has not been scheduled yet, l i =1 means job i has been scheduled;

[0020] Step 2.1.2: Set the action space AS of the deep reinforcement learning environment for single-machine total lead time scheduling;

[0021]

[0022] Among them, action a i , i=1,2,...,n g Indicates the selection of workpiece i;

[0023] Step 2.1.3. Set the reward R(t) of the deep reinforcement learning environment for single-machine total lead time scheduling at time step t:

[0024] R(t)=f(S(t))-f(S(t-1))

[0025]

[0026] Among them, f(S(t)) represents the total lead time of the scheduling plan under state S(t), C i represents the completion time of job i.

[0027] Furthermore, the step 2.2 includes the following steps:

[0028] Step 2.2.1. Set the state data S(t) of the input time step t, and the number of nodes is n r =3*n g ;

[0029] Step 2.2.2, the number of layers is n c , the number of nodes in each layer is n d , the activation function is ReLU;

[0030] Step 2.2.3: Output the workpiece dispatch probability model P(S(t)) at time step t, with the number of nodes being n g , normalized using the Softmax function;

[0031]

[0032] Among them, p i ,i=1,2...,n g represents the dispatch probability of job i.

[0033] Furthermore, the step 2.3 includes the following steps:

[0034] Step 2.3.1. Set the Adam parameters: learning rate lr, exponential decay rate β1 of momentum, and exponential decay rate β2 of squared gradient;

[0035] Step 2.3.2. At time step t, sample and execute action a(t)∈AS according to the workpiece dispatch probability model P(S(t-1)), obtain state S(t) and reward R(t), and form experience E(t) = (S(t-1), a(t), S(t), R(t));

[0036] Step 2.3.3: Based on the experience E(t), use the Adam algorithm to optimize the workpiece dispatch probability model P(S(t)) and increase the learning round by 1;

[0037] Step 2.3.4: If the learning round reaches Ep, the optimization ends, otherwise return to step 2.3.2.

[0038] Furthermore, the step 3 includes the following steps:

[0039] Step 3.1, set the population size to N p , evolutionary algebra N g , the crossover probability is P c , the mutation probability is P m , the sampling individual proportion α of the workpiece assignment probability model;

[0040] Step 3.2: Initialize the mixed population based on the workpiece assignment probability model;

[0041] Step 3.3, calculate the individual ss in the population k ,k=1,2,....,N p The fitness value fit k ;

[0042] Step 3.4: Calculate the individual ss in the population k ,k=1,2,....,N p The selection probability ps k ;

[0043] Step 3.5: Use the roulette wheel method to select two individuals from the population as parent individuals, generate two offspring individuals through two-point crossover and single-point mutation operations, and increase the total number of generated offspring individuals by 2;

[0044] Step 3.6, if the total number of offspring individuals generated is less than N p , then return to step 3.5, otherwise proceed to step 3.7;

[0045] Step 3.7, update the population and increase the evolutionary generation by 1:

[0046] Step 3.8, if the evolutionary number is less than N g , return to step 3.3; otherwise, output the current optimal solution.

[0047] Furthermore, the step 3.7 includes the following steps:

[0048] Step 3.7.1: Sort all parent individuals and generated offspring individuals by fitness value from large to small, and select the first (1-α)×N p Individuals enter the population.

[0049] Step 3.7.2: Generate αN by random sampling using the workpiece dispatch probability model p Individuals enter the population.

[0050] The advantages and positive effects of the present invention are:

[0051] The present invention sets the problem parameters for the total lead time scheduling of a single machine; constructs and optimizes a workpiece dispatch probability model based on deep reinforcement learning, inputs the problem parameters for the total lead time scheduling of a single machine to obtain an initial solution set; and uses a genetic evolutionary algorithm to optimize the initial solution set to obtain the optimal solution for single-machine scheduling. Compared with traditional methods, the present invention can reduce the search pressure of the optimization method and improve the search ability of the optimization method, thereby improving the efficiency and quality of solving the single-machine scheduling problem. At the same time, the present invention adopts population mixing initialization and population mixing update operations based on the workpiece dispatch probability model: through the mixed population initialization operation, the population diversity can be guaranteed and the initial population quality can be improved; the population mixing population update operation can be adopted, which can both ensure the evolution of the population and increase the diversity of genes. Compared with traditional methods, it can ensure that the algorithm takes into account both local search and global search capabilities. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 It is the overall structural diagram of the method of the present invention;

[0053] Figure 2 A flowchart of the reinforcement learning steps for the artifact assignment probability model of the method of the present invention;

[0054] Figure 3 This is a flow chart of the genetic algorithm based on the workpiece assignment probability model of the method of the present invention. DETAILED DESCRIPTION

[0055] The present invention is further described below in conjunction with the accompanying drawings.

[0056] A single-machine scheduling method based on reinforcement learning enhanced genetic evolution, such as Figure 1 As shown, the following steps are included:

[0057] Step 1: Set the problem parameters for single-machine total lead time scheduling.

[0058] The problem parameters for single-machine total lead time scheduling include the number of jobs n g =6, the release time of the workpiece r i ,i=1,2,...,n g , the processing time p of the workpiece i ,i=1,2,...,n g , the delivery time of the workpiece d i ,i=1,2,...,n g , as shown in Table 1

[0059] Table 1 Problem parameters for single-machine total lead time scheduling

[0060] Workpiece number <![CDATA[Release time r i > <![CDATA[Processing time p i > <![CDATA[Delivery period d i > 1 3 4 10 2 2 7 15 3 3 4 24 4 4 6 30 5 5 6 14 6 7 3 26

[0061] Step 2: Build and optimize a workpiece dispatch probability model based on deep reinforcement learning, input the problem parameters of single-machine total lead time scheduling to obtain the initial solution set.

[0062] Step 2.1: Reinforcement learning environment setup for single-machine total lead time scheduling.

[0063] Step 2.1.1. Set the state S(t) of the deep reinforcement learning environment for single-machine total lead time scheduling at time step t.

[0064] The state S(t) of the deep reinforcement learning environment for single-machine total lead time scheduling at time step t=3 is as follows:

[0065]

[0066] Among them, s i (t),i=1,2,...,n g represents the state of workpiece i at time step t, l i =0 means job i has not been scheduled yet, l i =1 means job i has been scheduled.

[0067] Step 2.1.2: Set the action space AS of the deep reinforcement learning environment for single-machine total lead time scheduling;

[0068] The action space AS of the deep reinforcement learning environment for single-machine total lead time scheduling is set as follows:

[0069]

[0070] Among them, action a i , i=1,2,...,n g Indicates the selection of workpiece i.

[0071] Step 2.1.3. Set the reward R(t) of the deep reinforcement learning environment for single-machine total lead time scheduling at time step t.

[0072] The reward R(t) at time step t in the deep reinforcement learning environment for single-machine total lead time scheduling is set as follows:

[0073] R(t)=f(S(t))-f(S(t-1))

[0074] Where f(S(t)) represents the total lead time of the scheduling solution under state S(t):

[0075]

[0076] Among them, C i represents the completion time of job i.

[0077] Step 2.2: Build a workpiece dispatch probability model based on deep reinforcement learning.

[0078] Step 2.2.1, Input layer setting: Input state data S(t) at time step t, the number of nodes is n r =3*n g ;

[0079] Step 2.2.2, middle layer setting: the number of layers is n c =2, the number of nodes in each layer is n d =128, and the activation function is ReLU.

[0080] Step 2.2.3, output layer setting: output the workpiece dispatch probability model P(S(t)) at time step t, with the number of nodes being n g =6, and the Softmax function is used for normalization.

[0081]

[0082] Among them, p i ,i=1,2...,n g represents the dispatch probability of job i.

[0083] Step 2.3: Optimize the workpiece dispatch probability model:

[0084] Step 2.3.1. Adam algorithm parameter settings: learning rate lr = 0.001, exponential decay rate of momentum β1 = 0.9, exponential decay rate of squared gradient β2 = 0.999.

[0085] Step 2.3.2. At time step t, sample and execute action a(t)∈AS according to the workpiece dispatch probability model P(S(t-1)), obtain state S(t) and reward R(t), and form experience E(t) = (S(t-1), a(t), S(t), R(t));

[0086] Step 2.3.3: Based on the experience E(t), use the Adam algorithm to optimize the workpiece dispatch probability model P(S(t)) and increase the learning round by 1.

[0087] Step 2.3.4: If the learning round reaches Ep=20000, the optimization ends; otherwise, return to step 2.3.2.

[0088] Step 3: Use the genetic evolutionary algorithm to optimize the initial solution set and obtain the optimal single-machine scheduling solution.

[0089] Step 3.1, parameter setting: population size is N p =200, evolutionary algebra N g =100, the crossover probability is P c=0.95, the mutation probability is P m =0.05, and the sampling individual proportion of the workpiece assignment probability model α = 0.10.

[0090] Step 3.2, such as Figure 2 As shown, the mixed population is initialized based on the probability model of workpiece dispatching.

[0091] Step 3.2.1, let the number of generated individuals be n s = 0, random sampling is used to generate individuals using the workpiece assignment probability model:

[0092] Step 3.2.1.1, let the number of scheduled jobs be n J =0, time step t=0;

[0093] Step 3.2.1.2: Obtain the workpiece probability distribution model P(S(t)) based on the state S(t);

[0094] Step 3.2.1.3: Randomly sample and schedule job J according to P(S(t)) t , and let the number of scheduled jobs n J =n J +1, time step t=t+1, update state S(t);

[0095] Step 3.2.1.4: If the number of scheduled jobs is less than n g , return to step 3.2.1.2.

[0096] Step 3.2.1.5, let the number of generated individuals n s =n s +1 if n s <αN p , return to step 3.2.1.

[0097] Step 3.2.2: Generate (1-α)N using uniformly distributed random sampling p individual.

[0098] Step 3.3, population fitness evaluation: calculate the individual ss in the population k ,k=1,2,....,N p The fitness value fit k .

[0099]

[0100] Among them, w i (ss k )=max(C i (ss k )-d i ,0) indicates that workpiece i is in individual ss kMedium delay time, C i (ss k ) represents the workpiece i in individual ss k Completion time.

[0101] Step 3.4, selection probability calculation: calculate the individual ss in the population k ,k=1,2,....,N p The selection probability ps k .

[0102]

[0103] Step 3.5: Use the roulette wheel method to select two parent individuals from the population, generate two offspring individuals through two-point crossover and single-point mutation operations, and increase the total number of offspring individuals generated by 2.

[0104] Step 3.6: If the total number of offspring individuals generated is less than N p , return to step 3.5.

[0105] Step 3.7, update the next generation population:

[0106] Step 3.7.1. Sort all parent individuals and generated offspring individuals by fitness from large to small, and select the first (1-α)×N p Individuals enter the population.

[0107] Step 3.7.2, such as Figure 3 As shown, the workpiece dispatch probability model is used to randomly sample and generate αN p Individuals enter the generation population:

[0108] Step 3.7.2.1, let the number of generated individuals be n s =0;

[0109] Step 3.7.2.2, let the number of scheduled jobs be n J =0, time step t=0;

[0110] Step 3.7.2.3. Obtain the workpiece probability distribution model P(S(t)) based on the state S(t);

[0111] Step 3.7.2.4: Randomly sample and schedule job J based on P(S(t)) t , and let the number of scheduled jobs n J =n J +1, time step t=t+1, update state S(t);

[0112] Step 3.7.2.5: If the number of scheduled jobs is less than n g , return to step 3.7.2.3.

[0113] Step 3.7.2.6, let the number of generated individuals n s =n s +1 if n s <αN p , return to step 3.7.2.2.

[0114] Step 3.8, if the evolutionary number is less than N g , return to step 3.3; otherwise, output the current optimal solution.

[0115] It should be emphasized that the embodiments described in the present invention are illustrative rather than restrictive. Therefore, the present invention includes but is not limited to the embodiments described in the specific embodiments. Any other embodiments derived by those skilled in the art based on the technical solutions of the present invention also fall within the scope of protection of the present invention.

Claims

1. A single-machine scheduling method based on reinforcement learning enhanced genetic evolution, characterized by: The following steps are involved: Step 1: Set the problem parameters for single-machine total lead time scheduling; Step 2: Build and optimize a job dispatch probability model based on deep reinforcement learning, input the problem parameters of single-machine total lead time scheduling to obtain the initial solution set; Step 3: Use the genetic evolutionary algorithm to optimize the initial solution set and obtain the optimal single-machine scheduling solution.

2. The single-machine scheduling method based on reinforcement learning enhanced genetic evolution according to claim 1, characterized in that: The problem parameters of the single-machine total lead time scheduling in step 1 include the number of workpieces n g , the release time of the workpiece r i ,i=1,2,...,n g , the processing time p of the workpiece i ,i=1,2,...,n g , the delivery time of the workpiece d i ,i=1,2,...,n g .

3. The single-machine scheduling method based on reinforcement learning enhanced genetic evolution according to claim 1, characterized in that: The step 2 comprises the following steps: Step 2.1: Set up a reinforcement learning environment for single-machine total lead time scheduling. Step 2.2: Build a workpiece dispatch probability model based on deep reinforcement learning; Step 2.3: Optimize the workpiece dispatch probability model.

4. The single-machine scheduling method based on reinforcement learning enhanced genetic evolution according to claim 3, characterized in that: The step 2.1 includes the following steps: Step 2.1.

1. Set the state S(t) of the deep reinforcement learning environment for single-machine total lead time scheduling at time step t: Among them, s i (t),i=1,2,...,n g represents the state of workpiece i at time step t, l i =0 means job i has not been scheduled yet, l i =1 means job i has been scheduled; Step 2.1.2: Set the action space AS of the deep reinforcement learning environment for single-machine total lead time scheduling; Among them, action a i , i=1,2,...,n g Indicates the selection of workpiece i; Step 2.1.

3. Set the reward R(t) of the deep reinforcement learning environment for single-machine total lead time scheduling at time step t: R(t)=f(S(t))-f(S(t-1)) Among them, f(S(t)) represents the total lead time of the scheduling plan under state S(t), C i represents the completion time of workpiece i.

5. The single-machine scheduling method based on reinforcement learning enhanced genetic evolution according to claim 3, characterized in that: The step 2.2 includes the following steps: Step 2.2.

1. Set the state data S(t) of the input time step t, and the number of nodes is n r =3*n g ; Step 2.2.2, the number of layers is n c , the number of nodes in each layer is n d , the activation function is ReLU; Step 2.2.3: Output the workpiece dispatch probability model P(S(t)) at time step t, with the number of nodes being n g , normalized using the Softmax function; Among them, p i ,i=1,2...,n g represents the dispatch probability of job i.

6. The single-machine scheduling method based on reinforcement learning enhanced genetic evolution according to claim 3, characterized in that: The step 2.3 includes the following steps: Step 2.3.

1. Set the Adam parameters: learning rate lr, exponential decay rate β1 of momentum, and exponential decay rate β2 of squared gradient; Step 2.3.

2. At time step t, sample and execute action a(t)∈AS according to the workpiece dispatch probability model P(S(t-1)), obtain state S(t) and reward R(t), and form experience E(t) = (S(t-1), a(t), S(t), R(t)); Step 2.3.3: Based on the experience E(t), use the Adam algorithm to optimize the workpiece dispatch probability model P(S(t)) and increase the learning round by 1; Step 2.3.4: If the learning round reaches Ep, the optimization ends, otherwise return to step 2.3.

2.

7. The single-machine scheduling method based on reinforcement learning enhanced genetic evolution according to claim 1, characterized in that: The step 3 comprises the following steps: Step 3.1, set the population size to N p , evolutionary algebra N g , the crossover probability is P c , the mutation probability is P m , the sampling individual proportion α of the workpiece assignment probability model; Step 3.2: Initialize the mixed population based on the workpiece assignment probability model; Step 3.3, calculate the individual ss in the population k ,k=1,2,....,N p The fitness value fit k ; Step 3.4: Calculate the individual ss in the population k ,k=1,2,....,N p The selection probability ps k ; Step 3.5: Use the roulette wheel method to select two individuals from the population as parent individuals, generate two offspring individuals through two-point crossover and single-point mutation operations, and increase the total number of generated offspring individuals by 2; Step 3.6, if the total number of offspring individuals generated is less than N p , then return to step 3.5, otherwise proceed to step 3.7; Step 3.7, update the population and increase the evolutionary generation by 1: Step 3.8, if the evolutionary number is less than N g , return to step 3.3; otherwise, output the current optimal solution.

8. The single-machine scheduling method based on reinforcement learning enhanced genetic evolution according to claim 7, characterized in that: The step 3.7 includes the following steps: Step 3.7.1: Sort all parent individuals and generated offspring individuals by fitness value from large to small, and select the first (1-α)×N p Individuals enter the population. Step 3.7.2: Generate αN by random sampling using the workpiece dispatch probability model p Individuals enter the population.