Multi-target dual hyper-heuristic method for flexible job shop self-organizing scheduling
Through a multi-objective dual-hyper-heuristic method combining genetic programming and deep reinforcement learning, flexible job shop scheduling rules are generated and optimized, which solves the global optimization problem of scheduling in a dynamic environment and achieves efficient self-organizing scheduling and rapid response.
Patent Information
- Application Number
- CN202510705948.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-09-05
AI Technical Summary
Existing flexible job shop scheduling methods find it difficult to achieve efficient global optimization when faced with dynamic changes. Existing scheduling algorithms also find it difficult to balance the efficiency dilemma between local and global optimization when dealing with rapidly changing and multi-interference workshop environments, and lack real-time, high-quality global solutions.
A multi-objective dual-hyper-heuristic method for self-organizing scheduling of flexible job shops is adopted, combined with genetic programming rule generation technology and deep reinforcement learning method to generate process and interval selection rules. The multi-step action sequence is optimized through the Markov decision process to achieve autonomous generation and selection of dynamic decision strategies.
A high-quality self-organizing scheduling solution is implemented in a dynamic environment, significantly improving the self-healing capability and response speed of the production system. It can make autonomous decisions under disturbances such as equipment failures, processing fluctuations, and the insertion of new orders, providing an efficient scheduling solution for intelligent manufacturing scenarios.
Smart Images

Figure CN120595582A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of industrial scheduling, and in particular to a multi-objective dual-hyper heuristic method for self-organizing scheduling of flexible job shops. Background Art
[0002] The dynamic flexible job shop scheduling problem (DFJSP) has been a research hotspot for decades. The goal of dynamic scheduling is to maximize production efficiency and minimize deviations from the original schedule when job- or resource-related disturbances occur in the system. When new events occur, dynamic scheduling methods adjust the original schedule through various schedule repair or rescheduling techniques. Common rescheduling strategies include fully reactive, robust-proactive, and predictive-reactive. While traditional predictive scheduling methods can generate optimal plans under ideal conditions, they lack flexibility in the face of dynamic changes. Fully reactive scheduling allocates tasks and resources locally in real time to address emergencies, but lacks a global optimization perspective. In predictive-reactive scheduling, a pre-scheduling plan is always present to guide the organization of production resources such as energy, materials, and personnel. Once a rescheduling event is triggered, the plan is optimized and updated. Due to its high flexibility, robustness, and stability, predictive-reactive scheduling is the most widely used in practice. However, since it involves real-time, multi-objective global optimization, the problem of predictive reactive scheduling is highly complex, computationally intensive, and requires good generalization capabilities to adapt to different disturbance scenarios.
[0003] To ensure efficient production, it is necessary to reduce the time cost of self-organizing scheduling. Optimization methods for self-organizing scheduling are mainly divided into exact mathematical algorithms, metaheuristic methods based on solution search, and heuristic methods based on solution generation. Research shows that scheduling problems with more than two jobs are nondeterministic polynomial problems. Traditional mathematical methods struggle to obtain satisfactory results within a limited timeframe and are particularly limited in dynamic environments. Metaheuristics, including tabu search, genetic algorithms, particle swarm optimization, and hybrid algorithms, iterate and search the solution space to obtain near-optimal solutions. For flexible job shop problems, the high time complexity of metaheuristics means that they require several minutes of computation time. In dynamic applications, production conditions and order status may fluctuate unpredictably due to computational time constraints, resulting in limited or even inaccurate performance of self-organizing scheduling. Heuristic methods specify production tasks and resources through prioritization and selection rules. Their ease of interpretation and real-time responsiveness have led to their widespread use in real-world production. Many scheduling rules are widely used, such as LPT (Longest Processing Time), WSPT (Weighted Shortest Processing Time), and some combined rules.
[0004] Scheduling rules are typically manually designed for specific optimization objectives and problem scenarios, relying heavily on expert knowledge and experience. Most dispatching rules adhere to the same criteria at defined decision points. However, over time, changes in the environment can render the rules inapplicable, making them unreliable at both the local and global levels. To design more efficient heuristic rules, hyperheuristics have attracted research attention. Hyperheuristics seek to find heuristics that generate solutions without requiring extensive domain knowledge. Their efficient automatic learning and solution generation capabilities have greatly improved the application of hyperheuristics in dynamic scenarios. Furthermore, methods for selecting scheduling rules have been explored, such as the Analytic Hierarchy Process (AHP), semantic methods, and deep reinforcement learning based on deep Q-networks. However, most applications of hyperheuristics in dynamic scheduling fall within the realm of fully reactive scheduling, aiming to allocate and sequence processes in real time, such as in scenarios with dynamic workpiece arrival. In predictive reactive scheduling, high-quality, real-time global solutions are crucial for guiding production scheduling. Existing scheduling optimization algorithms still face challenges in coping with rapidly changing and multi-interference shop floor environments, struggling to strike a balance between local and global optimization efficiency. Furthermore, the differences between new and old solutions in predictive reactive scheduling affect production stability, requiring an efficient search mechanism that balances stability and performance objectives. Therefore, DFJSP urgently needs methods that can provide more accurate and robust global solutions in real time to adapt to different types of dynamic changes and uncertainties in the production environment.
[0005] Currently, no effective solutions have been proposed for the problems in related technologies. Summary of the Invention
[0006] In response to the problems in the related art, the present invention proposes a multi-objective dual-hyper heuristic method for flexible job shop self-organizing scheduling to overcome the above-mentioned technical problems existing in the existing related art.
[0007] To this end, the specific technical solutions adopted in the present invention are as follows:
[0008] The multi-objective dual-hyper heuristic method for flexible job shop self-organizing scheduling includes the following steps:
[0009] S1. Generate process selection rules and interval selection rules using genetic programming rule generation technology, and generate a rule set based on the process selection rules and interval selection rules;
[0010] S2. Based on deep reinforcement learning, a dynamic decision-making strategy is constructed in combination with a set of rules, and multi-step action sequence optimization is performed when self-organizing scheduling is triggered;
[0011] S3. Generate instances by randomly combining parameter indicators of predetermined dynamic events, and calculate performance indicators based on the generated instances to verify feasibility.
[0012] Furthermore, the method of using genetic programming rule generation technology to generate process selection rules and interval selection rules, and generating a rule set based on the process selection rules and interval selection rules includes the following steps:
[0013] S11. Use binary trees to represent feasible solutions in genetic programming. Generate an initial population randomly in the form of binary trees with random depths within the depth range of the binary trees. Each individual in the population is considered as a function.
[0014] S12. Determine the average maximum completion time after repair and the average maximum completion time after damage based on the random self-organizing scheduling instance, and use the difference between the average maximum completion time before and after repair as the fitness function;
[0015] S13, using a roulette wheel selection strategy to select individuals from the parent population, and using the fitness function to determine the roulette wheel probability;
[0016] S14. Based on the selected individuals, combined with the crossover rate and mutation rate, the process tree and the interval tree are crossovered and mutated accordingly, so that the subtree below the mutation point will be cut and replaced by a randomly generated tree; the new tree that exceeds the depth is pruned;
[0017] S15. According to the evolutionary results, output several individuals with the best fitness in the current disturbance environment, update the disturbance environment, and finally output individuals with the best fitness in all environments, which together constitute a set of rules.
[0018] Furthermore, the binary tree is composed of a function set and a terminal set, and the function set includes digital operators, mathematical functions and Boolean operators, and the terminal set is composed of a terminal set for selecting a key process and a terminal set for selecting a key interval;
[0019] The terminal set for selecting the critical process includes: the shortest processing time of the process, the longest processing time of the process, the completion time of the previous process of the process, the difference between the pre-scheduled process completion time and the process completion time of the disrupted scheduling plan, the subsequent total processing time on the processing machine of the process, the subsequent total processing time required for the workpiece corresponding to the process, the subsequent total slack time on the processing machine of the process, the subsequent total slack time required for the workpiece corresponding to the process, the subsequent total processing time required for the workpiece corresponding to the process on the critical path, the subsequent total processing time of the processing machine of the process on the critical path, and whether the next process of the machine is on the critical path;
[0020] The terminal set for selecting the critical interval includes: the duration of the interval, the total processing time of the machine after the interval, the total slack time of the machine after the interval, the start time of the interval, the equipment utilization rate of the interval machine, the total processing time of the machine on the critical path after the interval, the total number of operations of the machine on the critical path after the interval, and whether the next process of the machine is on the critical path.
[0021] Furthermore, each individual solution of genetic programming is encoded by a process tree and an interval tree, which respectively represent the selection rules of processes in the process set and the selection rules of available intervals.
[0022] Furthermore, the method of building a dynamic decision-making strategy based on a deep reinforcement learning method in combination with a rule set and performing multi-step action sequence optimization when self-organizing scheduling is triggered includes the following steps:
[0023] S21, converting the local optimization process obtained by decomposition after the self-organizing scheduling is triggered into a Markov decision process;
[0024] S22. Based on the Markov decision process, deep reinforcement learning is used to implement dynamic decision-making and update scheduling strategies within the rule set to achieve multi-step action sequence optimization.
[0025] Furthermore, converting the local optimization process obtained by decomposition after the self-organizing scheduling is triggered into a Markov decision process includes the following steps:
[0026] S211. Design state features of deep reinforcement learning using information based on slack time and critical path, and normalize each state feature;
[0027] S212. Based on the damaged schedule, the self-organizing scheduling agent selects a repair heuristic action at time point t based on the deep neural network trained by the PPO algorithm;
[0028] S213: Decode the selected repair heuristic action into a corresponding repair behavior according to the current state, insert the corresponding process into the specified position, and transfer the environment to the new state according to the transition probability;
[0029] S214. As actions are executed and schedules change, the agent receives rewards from the reward function and updates its policy based on the observed states, actions, rewards, and resulting states.
[0030] Among them, the reward is defined as C after local optimization max The difference is expressed as r(s t ,a t ,s t+1 )=C max (s t )-C max (st+1 ), where r t represents the reward at time t, s t represents the state at time point t, a t represents the action at time t, s t+1 Indicates the state at time t+1, C max (s t ) means s t The maximum completion time under the state, C max (s t+1 ) means s t+1 The maximum completion time under the state;
[0031] When the discount factor γ = 1, the cumulative reward of self-organizing scheduling is Where, Indicates the number of unfinished processes that need to be processed for rescheduling, C max Indicates the maximum completion time.
[0032] Furthermore, the state characteristics include global state characteristics and critical path state characteristics;
[0033] Among them, the global state characteristics include: the average total processing time on each machine at time step t, the average total slack time of each workpiece at time step t, the average total slack time of each machine at time step t, the minimum completion time between workpieces at time step t, the minimum start time of the remaining processes at time step t, the difference between the maximum completion time after repair and the maximum completion time after damage at time step t, the ratio of the total processing time of completed processes to the total processing time at time step t, the ratio of the number of completed processes to the total number of processes at time step t, and the ratio of the maximum completion time at time step t of self-organized scheduling to the maximum completion time after damage.
[0034] The critical path status characteristics include: the proportion of the number of processes on the critical path at time step t to the total number of processes, the maximum processing time of all processes on the critical path at time step t, the minimum processing time of all processes on the critical path at time step t, the average left slack time of each process on the critical path at time step t, and the average right slack time of each process on the critical path at time step t.
[0035] Furthermore, the Markov decision process-based, deep reinforcement learning-based dynamic decision-making and updating scheduling strategy within the rule set to achieve multi-step action sequence optimization includes:
[0036] The deep neural network is trained through the proximal policy optimization algorithm, and the trained deep neural network is used to realize the end-to-end mapping from the state of the Markov decision process to the action selection probability within the rule set, realizing multi-step action sequence optimization and obtaining the optimized scheduling strategy.
[0037] Furthermore, deep reinforcement learning uses a proximal policy optimization algorithm for training, and its objective function is:
[0038]
[0039] Where J(θ) represents the objective function, represents the expected value of the data at the tth time step, r t (θ) represents the action probability ratio of the new and old strategies, Indicates state s t Next select action a t The relative gain of r t (θ)>1+∈, it is truncated to 1+∈, when r t When (θ)<1-∈, it is truncated to 1-∈, otherwise it is retained r t (θ),∈ represents the cropping coefficient, π θ Indicates the current strategy, Indicates the old policy.
[0040] Furthermore, the generating of instances by randomly combining parameter indicators of predetermined dynamic events includes:
[0041] Taking the arrival of new jobs, machine failures, and changes in job processing time as dynamic events, we set several parameters for machine failures, several parameters for changes in job processing time, and several parameters for the arrival of new workpieces, and then randomly generate instances through random combinations.
[0042] The beneficial effects of the present invention are:
[0043] 1) This invention addresses the challenge of obtaining a high-quality global solution for self-organizing scheduling online, achieving an effective self-organizing scheduling solution that balances performance and stability objectives in near real time. Comparisons with widely used evolutionary algorithms on an extended benchmark demonstrate the effectiveness of the proposed method and its versatility across diverse scenarios.
[0044] 2) The present invention theoretically expands the scope of application of existing scheduling algorithms, especially in dealing with scheduling problems in multi-objective and dynamic environments. Compared with the local optimization based on rules or time windows in the past, the hyper-heuristic generation and selection strategy based on repair expands the search space and solution efficiency of the solution, and provides a new theoretical framework for achieving high-quality global solutions. Through the dynamic collaborative optimization mechanism of self-organizing scheduling, it is possible to achieve an autonomous decision-making closed loop under multiple disturbances such as equipment failure, processing fluctuations and new order insertion, significantly improving the self-healing ability, response speed and anti-interference resilience of the production system in a dynamic environment, and providing an intelligent decision-making solution with autonomous evolution capabilities for complex scheduling problems in intelligent manufacturing scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0046] Figure 1 is a flow chart of a multi-objective dual-hyper heuristic method for flexible job shop self-organizing scheduling according to an embodiment of the present invention;
[0047] Figure 2 1 is a general framework diagram of a multi-objective dual-hyper heuristic method for flexible job shop self-organizing scheduling according to an embodiment of the present invention;
[0048] Figure 3 2 is a schematic diagram of decoding based on dual-tree repair in a multi-objective dual-hyper heuristic method for flexible job shop self-organizing scheduling according to an embodiment of the present invention;
[0049] Figure 4 is a flowchart of an implementation of the GP method in the multi-objective dual-hyper heuristic method for flexible job shop self-organizing scheduling according to an embodiment of the present invention;
[0050] Figure 5 is a flowchart of an implementation of a DRL method in a multi-objective dual-hyper heuristic method for flexible job shop self-organizing scheduling according to an embodiment of the present invention;
[0051] Figure 6 is an average graph and confidence interval graph of IGD results under different disturbances according to an embodiment of the present invention;
[0052] Figure 7 is a Pareto frontier diagram under different disturbances according to an embodiment of the present invention, where C max is the maximum completion time. DETAILED DESCRIPTION
[0053] To further illustrate each embodiment, the present invention provides drawings, which are part of the disclosure of the present invention. They are mainly used to illustrate the embodiments and can be used in conjunction with the relevant descriptions in the specification to explain the operating principles of the embodiments. By referring to these contents, ordinary technicians in this field should be able to understand other possible implementation methods and advantages of the present invention. The components in the figures are not drawn to scale, and similar component symbols are generally used to represent similar components.
[0054] According to an embodiment of the present invention, a multi-objective dual-hyper heuristic method for flexible job shop self-organizing scheduling is provided.
[0055] The present invention will now be further described with reference to the accompanying drawings and specific embodiments. Figure 1-Figure 7 As shown, according to one embodiment of the present invention, a multi-objective dual-hyper heuristic method for flexible job shop self-organizing scheduling is provided, comprising the following steps:
[0056] S1. Generate process selection rules and interval selection rules using genetic programming rule generation technology, and generate a rule set based on the process selection rules and interval selection rules;
[0057] Specifically, if Figure 3 As shown in the figure, M1, M2, and M3 represent machines 1, 2, and 3 respectively. Each solution individual of Genetic Programming (GP) is encoded by two trees (i.e., process tree and interval tree), which represent the selection rules f for the process in the process set. α (O i,j ) and the selection rule for the available intervals f θ (itv g,h ), where O i,j represents the jth process of workpiece i, itv g,h Indicates process O g,h After the original scheduling plan is destroyed, the critical path process set CP rt and available interval sets ITV 3,2,rt According to f α (O i,j ) and f β (itv g,h ) sorted to obtain the sorted critical path process set CP' rt and available interval set ITV' 3,2,rt , where rt represents the fault start time, R rt,2 represents the duration of the fault on machine 2. The generated process selection rules and interval selection rules include:
[0058] S11. Use binary trees to represent feasible solutions in genetic programming. Generate an initial population randomly in the form of a binary tree with random depth within the depth range of the binary tree. Each individual in the population can be regarded as a function f, which is represented by the function f. s As the root node, the terminal as the leaf node, Figure 3 In the example, the process selection rule f α The root nodes are + and / , the leaf nodes are SPT, LPT, EST, and the selection rule of the available intervals is f β The root nodes are - and *, and the leaf nodes are IT, ST, and TRP;
[0059] Specifically, the binary tree is composed of a function set and a terminal set, and the function set includes operations on index values, including arithmetic operators, mathematical functions, Boolean operators and conditional expressions, etc. The function set is composed of f s ={+,-,*, / ,max,min,if}, where +,-,*, / are numeric operators, max and min are maximum and minimum values, and if is a Boolean operator. The terminal set consists of 11 terminal sets for selecting critical processes and 8 terminal sets for selecting critical intervals.
[0060] Among them, the 11 terminal sets for selecting critical processes include: the shortest processing time SPT of the process, the longest processing time LPT of the process, the completion time EST of the previous process of the process, the difference DET between the completion time of the pre-scheduled process and the completion time of the process in the disrupted scheduling plan, the subsequent total processing time TRPE on the processing machine of the process, the subsequent total processing time TRPJ required for the workpiece corresponding to the process, the subsequent total slack time TRSE on the processing machine of the process, the subsequent total slack time TRSJ required for the workpiece corresponding to the process, the subsequent total processing time JTC required for the workpiece corresponding to the process on the critical path, the subsequent total processing time ETC of the processing machine of the process on the critical path, and whether the next process of the machine is on the critical path (NICO), which is 1 when the next process of the machine is on the critical path, otherwise it is 0;
[0061] The 8 terminal sets for selecting critical intervals include: the duration of the interval IT, the total processing time of the machine after the interval TRP, the total slack time of the machine after the interval TRI, the start time of the interval ST, the equipment utilization rate of the interval machine EU, the total processing time of the machine on the critical path after the interval ECC, the total number of operations of the machine on the critical path after the interval ECT, and whether the next process of the machine is on the critical path NICI. When the next process of the machine is on the critical path, it is 1, otherwise it is 0.
[0062] S12, based on 10 random self-organizing scheduling instances, determine the average maximum completion time after repair and the average maximum completion time after damage, and use the difference between the average maximum completion time before and after repair as the average maximum completion time. is the fitness function to evaluate the fitness of the individual, is the maximum completion time after repair, is the maximum completion time after damage;
[0063] S13, using a roulette wheel selection strategy to select individuals from the parent population, and using the fitness function to determine the roulette wheel probability;
[0064] S14. In the evolution stage, based on the selected individuals, combined with the crossover rate and mutation rate, the process tree and the interval tree are crossovered and mutated respectively, so that the subtree below the mutation point will be cut and replaced by a randomly generated tree; the new tree that exceeds the depth is pruned;
[0065] S15. Based on the evolutionary results, output the p individuals with the best fitness in the current perturbation environment, update the perturbation environment, and finally output the individuals with the best fitness in all environments, which together constitute the action set of DRL (Deep Reinforcement Learning, deep reinforcement learning), that is, the rule set.
[0066] S2. Based on deep reinforcement learning, a dynamic decision-making strategy is constructed in combination with a set of rules, and multi-step action sequence optimization is performed when self-organizing scheduling is triggered;
[0067] Among them, such as Figure 4 As shown, the method based on deep reinforcement learning, combining rule sets to build a dynamic decision-making strategy, and optimizing multi-step action sequences when self-organizing scheduling is triggered includes the following steps:
[0068] S21, converting the local optimization process obtained by decomposition after the self-organizing scheduling is triggered into a Markov decision process;
[0069] Here, a Markov decision is represented by a five-element tuple:<S,a,P,γ,R> S represents the state space, which is the set of all possible situations or configurations that the environment can be in. A represents the action space, which is the set of all possible actions that the agent can take. P is the transition probability, which is the probability of taking a specific action a. t After that, from a state (s t )Transfer to another state(s t+1 ) related to the probability. It reflects the dynamic characteristics of the environment. R represents the reward function, which specifies the immediate reward (r t). γ represents the discount factor, which is a parameter between 0 and 1 that is used to discount future rewards to balance the trade-off between current and future rewards.
[0070] Specifically, converting the local optimization process obtained by decomposition after the self-organizing scheduling is triggered into a Markov decision process includes the following steps:
[0071] S211. Design state features of deep reinforcement learning using information based on relaxation time and critical path, and normalize each state feature to the range of [0, 1].
[0072] Specifically, the state features include 9 global state features and 5 critical path state features;
[0073] Among them, the nine global state features include: the average total processing time MRPT(t) on each machine at time step t, the average total slack time MSTJ(t) of each job at time step t, the average total slack time MSTM(t) of each machine at time step t, the minimum completion time minC(t) between jobs at time step t, the minimum start time minNST(t) of the remaining processes at time step t, the difference ΔC(t) between the maximum completion time after repair and the maximum completion time after damage at time step t, the ratio of the total processing time of completed processes to the total processing time CRO(t) at time step t, the ratio of the number of completed processes to the total number of processes ORO(t) at time step t, and the ratio of the maximum completion time at time step t to the maximum completion time after damage in self-organizing scheduling RRO(t).
[0074] The five critical path status characteristics include: the proportion of the number of processes on the critical path to the total number of processes at time step t CPO(t), the maximum processing time maxPCP(t) of all processes on the critical path at time step t, the minimum processing time minPCP(t) of all processes on the critical path at time step t, the average left slack time MLSCP(t) of each process on the critical path at time step t, and the average right slack time MRSCP(t) of each process on the critical path at time step t.
[0075] S212. Based on the damaged schedule, the self-organizing scheduling agent selects a repair heuristic action a at time point t based on strategy π t ;
[0076] S213, according to the current system status s t The selected repair heuristic action a t Decoded into the corresponding repair behavior (O i,j ,itv i,j,h,g ), where O i,jrepresents the jth process of workpiece i, itv g,h Indicates process O g,h The corresponding process is inserted into the specified position, and the environment is transferred to the new state (s t+1 );
[0077] S214. As actions are executed and schedules change, the agent receives rewards (r t ), and updates its policy π based on the observed state, action, reward, and resulting state;
[0078] Among them, the reward is defined as C after local optimization max The difference is expressed as r(s t ,a t ,s t+1 )=C max (s t )-C max (s t+1 ), where r t represents the reward at time t, s t represents the state at time point t, a t represents the action at time t, s t+1 Indicates the state at time t+1, C max (s t ) means s t The maximum completion time under the state, C max (s t+1 ) means s t+1 The maximum completion time under the state;
[0079] When the discount factor γ = 1, the cumulative reward of self-organizing scheduling is Where, Indicates the number of unfinished processes that need to be processed for rescheduling, C max represents the maximum completion time. For a specific instance of the self-organizing scheduling problem, C max (s0) is a constant, i.e. This means minimizing C max is equivalent to maximizing G.
[0080] S22. Based on the Markov decision process, deep reinforcement learning is used to implement dynamic decision-making and update scheduling strategies within the rule set to achieve multi-step action sequence optimization.
[0081] Specifically, the Markov decision process-based, deep reinforcement learning-based dynamic decision-making and updating scheduling strategy within a rule set to achieve multi-step action sequence optimization includes:
[0082] The deep neural network is trained through the proximal policy optimization algorithm, and the trained deep neural network is used to realize the end-to-end mapping from the state of the Markov decision process to the action selection probability within the rule set, realizing multi-step action sequence optimization and obtaining the optimized scheduling strategy.
[0083] Specifically, deep reinforcement learning uses the Proximal Policy Optimization (PPO) algorithm for training, and its objective function is:
[0084]
[0085]
[0086] Where J(θ) represents the objective function, represents the expected value of the data at the tth time step, r t (θ) represents the action probability ratio of the new and old strategies, Indicates state s t Next select action a t The relative gain of r t (θ)>1+∈, it is truncated to 1+∈, when r t When (θ)<1-∈, it is truncated to 1-∈, otherwise it is retained r t (θ),∈ represents the clipping coefficient, ∈∈(0,1), reflecting the similarity between the current strategy and the old strategy, π θ Indicates the current strategy, Indicates the old policy.
[0087] The above technical solution is divided into the GP layer and the DRL layer to describe the overall framework of the present invention. The following compares this method with some traditional dynamic scheduling heuristic algorithms. The MK series standard examples MK01-MK10 are used for simulation and experimental comparison.
[0088] S3. Generate instances by randomly combining parameter indicators of predetermined dynamic events, and calculate performance indicators based on the generated instances to verify feasibility.
[0089] In this embodiment, generating composite dynamic scene test data includes:
[0090] Three dynamic events are considered: new job arrival, machine failure, and job processing time change. Four parameters for machine failure, two parameters for job processing time change, and five parameters for new job arrival are set accordingly. Random instances are then generated through random combinations.
[0091] Among them, the parameters of the four machine failures are classified according to their characteristics:
[0092] (1) The mean time to repair (MTTR) is evenly distributed in [10, 20];
[0093] (2) Mean time between failures (MTBF) is evenly distributed in [50,70];
[0094] (3) The time to repair (TTR) is an exponential distribution of MTTR.
[0095] (4) The time between failures (TBF) is an exponential distribution of MTBF.
[0096] The parameters of the two job processing time changes are classified according to their characteristics:
[0097] (1) The variation of processing time is evenly distributed in [1,10];
[0098] (2) The probability of occurrence is 0.3.
[0099] The parameters of the 5 new workpieces arriving are classified according to their characteristics:
[0100] (1) The number of steps for each new job is uniformly distributed in [1, m + 2], where m is the number of machines;
[0101] (2) The number of machines available for each process is randomly selected from the set {1, 2, …, m}, where m is the number of machines;
[0102] (3) The processing time of each new process is evenly distributed in [1,10];
[0103] (4) The mean time between new workpiece arrivals (MTBJA) is 20;
[0104] (5) The time distribution of new workpiece arrival is the exponential distribution of MTBJA.
[0105] like Figure 6 and Figure 7 As shown in the figure, the two performance indicators of real-time calculation of inverted generational distance (IGD) and calculation time prove the feasibility of this method compared with existing methods, where DHH is the dual hyper heuristic method mentioned in the present invention, MOEA / D is a multi-objective evolutionary algorithm based on decomposition, NSGAII is a multi-objective genetic algorithm, NEFRL is a reinforcement learning algorithm combining neural network, evolutionary algorithm and fuzzy logic, IGA is an immune genetic algorithm, and MBO is a migratory bird optimization algorithm;
[0106] Specifically, IGD is used to evaluate the performance of different algorithms. The calculation formula of IGD is:
[0107]
[0108] Where P is the approximate Pareto solution set obtained by the algorithm, A is the true Pareto frontier, |P| is the number of solutions in the solution set P, and d i,p,A is the minimum Euclidean distance from the i-th solution in the solution set P to the reference set A;
[0109] IGD is a comprehensive indicator for measuring convergence and distribution. The smaller its value, the better the overall performance of the algorithm.
[0110] The CPU computation time is used to evaluate the performance of different algorithms.
[0111] The present invention proposes an efficient multi-objective dual-hyper-heuristic method for self-organizing scheduling of flexible job shops, the core of which is composed of a GP rule generation layer and a DRL strategy optimization layer. In the rule generation layer, a dual-tree encoding model is constructed through a multi-objective evolutionary mechanism: the process selection tree integrates 11 process status indicators such as SPT (shortest processing time) and JTC (total processing time on the subsequent critical path), and the interval selection tree integrates 8 dynamic interval features such as IT (interval duration) and TRP (subsequent processing time). Combined with the cross-mutation strategy, a diversified local repair rule set is generated. In the strategy optimization layer, a Markov decision process is constructed based on the PPO algorithm, and the GP rule base is mapped into a dynamic action space. Production disturbances are perceived in real time through 14-dimensional mixed state features (global load, critical path slack time, etc.), and a reward function with completion time difference as the core is designed to drive the intelligent agent to quickly select the optimal repair sequence when self-organizing scheduling is triggered, thereby realizing multi-objective collaborative optimization. The global scheduling solution space is decomposed into serialized local decisions, which significantly reduces the computational complexity.
[0112] The advantages of the present invention are reflected in three aspects: First, through the GP-DRL two-layer cascade mechanism, the autonomous generation of rules is decoupled from dynamic decision-making, which improves the computational efficiency by 98% compared with traditional metaheuristic algorithms (such as NSGA-II), and reduces the IGD value by 85% in the MK series benchmark case; second, it combines critical path characteristics with multi-objective reward design, and tests show that its stability is better than algorithms such as MOEA / D; third, for complex disturbance scenarios (machine failure, processing fluctuations, new job insertion), a dynamic event simulator is constructed and the generalization ability is enhanced through a hybrid training strategy. In randomly generated industrial-grade test cases, the average completion time is optimized by 42%, and the self-organizing scheduling response time is stabilized within 1 second, meeting real-time requirements. Experimental verification shows that the present invention provides an efficient and reliable solution for dynamic workshop scheduling and has significant industrial application value.
[0113] In the present invention, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A multi-objective dual-hyper-heuristic method for flexible job shop self-organizing scheduling, characterized by: The following steps are involved: S1. Generate process selection rules and interval selection rules using genetic programming rule generation technology, and generate a rule set based on the process selection rules and interval selection rules; S2. Based on deep reinforcement learning, a dynamic decision-making strategy is constructed in combination with a set of rules, and multi-step action sequence optimization is performed when self-organizing scheduling is triggered; S3. Generate instances by randomly combining parameter indicators of predetermined dynamic events, and calculate performance indicators based on the generated instances to verify feasibility.
2. The multi-objective dual-hyper heuristic method for flexible job shop self-organizing scheduling according to claim 1 is characterized in that: The method of using genetic programming rule generation technology to generate process selection rules and interval selection rules, and generating a rule set based on the process selection rules and interval selection rules includes the following steps: S11. Use binary trees to represent feasible solutions in genetic programming. Generate an initial population randomly in the form of binary trees with random depths within the depth range of the binary trees. Each individual in the population is considered as a function. S12. Determine the average maximum completion time after repair and the average maximum completion time after damage based on the random self-organizing scheduling instance, and use the difference between the average maximum completion time before and after repair as the fitness function; S13, using a roulette wheel selection strategy to select individuals from the parent population, and using the fitness function to determine the roulette wheel probability; S14. Based on the selected individuals, combined with the crossover rate and mutation rate, the process tree and the interval tree are crossovered and mutated accordingly, so that the subtree below the mutation point will be cut and replaced by a randomly generated tree; the new tree that exceeds the depth is pruned; S15. According to the evolutionary results, output several individuals with the best fitness in the current disturbance environment, update the disturbance environment, and finally output individuals with the best fitness in all environments, which together constitute a set of rules.
3. The multi-objective dual-hyper heuristic method for flexible job shop self-organizing scheduling according to claim 2 is characterized in that: The binary tree is composed of a function set and a terminal set, and the function set includes digital operators, mathematical functions and Boolean operators, and the terminal set is composed of a terminal set for selecting a key process and a terminal set for selecting a key interval; The terminal set for selecting the critical process includes: the shortest processing time of the process, the longest processing time of the process, the completion time of the previous process of the process, the difference between the pre-scheduled process completion time and the process completion time of the disrupted scheduling plan, the subsequent total processing time on the processing machine of the process, the subsequent total processing time required for the workpiece corresponding to the process, the subsequent total slack time on the processing machine of the process, the subsequent total slack time required for the workpiece corresponding to the process, the subsequent total processing time required for the workpiece corresponding to the process on the critical path, the subsequent total processing time of the processing machine of the process on the critical path, and whether the next process of the machine is on the critical path; The terminal set for selecting the critical interval includes: the duration of the interval, the total processing time of the machine after the interval, the total slack time of the machine after the interval, the start time of the interval, the equipment utilization rate of the interval machine, the total processing time of the machine on the critical path after the interval, the total number of operations of the machine on the critical path after the interval, and whether the next process of the machine is on the critical path.
4. The multi-objective dual-hyper heuristic method for flexible job shop self-organizing scheduling according to claim 2 is characterized in that: Each individual solution of genetic programming is encoded by a process tree and an interval tree, which represent the selection rules of processes in the process set and the selection rules of available intervals respectively.
5. The multi-objective dual-hyper heuristic method for flexible job shop self-organizing scheduling according to claim 1 is characterized in that: The method of building a dynamic decision-making strategy based on a deep reinforcement learning method and combining a rule set, and optimizing a multi-step action sequence when self-organizing scheduling is triggered, includes the following steps: S21, converting the local optimization process obtained by decomposition after the self-organizing scheduling is triggered into a Markov decision process; S22. Based on the Markov decision process, deep reinforcement learning is used to implement dynamic decision-making and update scheduling strategies within the rule set to achieve multi-step action sequence optimization.
6. The multi-objective dual-hyper heuristic method for flexible job shop self-organizing scheduling according to claim 5 is characterized in that: Converting the local optimization process obtained by decomposition after the self-organizing scheduling is triggered into a Markov decision process comprises the following steps: S211. Design state features of deep reinforcement learning using information based on slack time and critical path, and normalize each state feature; S212. Based on the damaged schedule, the self-organizing scheduling agent selects a repair heuristic action at time point t based on the deep neural network trained by the PPO algorithm; S213: Decode the selected repair heuristic action into a corresponding repair behavior according to the current state, insert the corresponding process into the specified position, and transfer the environment to the new state according to the transition probability; S214. As actions are executed and schedules change, the agent receives rewards from the reward function and updates its policy based on the observed states, actions, rewards, and resulting states. Among them, the reward is defined as C after local optimization max The difference is expressed as r(s t ,a t ,s t+1 )=C max (s t )-C max (s t+1 ), where r t represents the reward at time t, s t represents the state at time t, a t represents the action at time t, s t+1 Indicates the state at time t+1, C max (s t ) means s t The maximum completion time under the state, C max (s t+1 ) means s t+1 The maximum completion time under the state; When the discount factor γ = 1, the cumulative reward of self-organizing scheduling is Where, Indicates the number of unfinished processes that need to be processed for rescheduling, C max Indicates the maximum completion time.
7. The multi-objective dual-hyper heuristic method for flexible job shop self-organizing scheduling according to claim 6 is characterized in that: The state characteristics include global state characteristics and critical path state characteristics; Among them, the global state characteristics include: the average total processing time on each machine at time step t, the average total slack time of each workpiece at time step t, the average total slack time of each machine at time step t, the minimum completion time between workpieces at time step t, the minimum start time of the remaining processes at time step t, the difference between the maximum completion time after repair and the maximum completion time after damage at time step t, the ratio of the total processing time of completed processes to the total processing time at time step t, the ratio of the number of completed processes to the total number of processes at time step t, and the ratio of the maximum completion time at time step t of self-organized scheduling to the maximum completion time after damage. The critical path status characteristics include: the proportion of the number of processes on the critical path at time step t to the total number of processes, the maximum processing time of all processes on the critical path at time step t, the minimum processing time of all processes on the critical path at time step t, the average left slack time of each process on the critical path at time step t, and the average right slack time of each process on the critical path at time step t.
8. The multi-objective dual-hyper heuristic method for flexible job shop self-organizing scheduling according to claim 5 is characterized in that: The Markov decision process-based, deep reinforcement learning-based method implements dynamic decision-making and updating scheduling strategies within a rule set to achieve multi-step action sequence optimization, including: The deep neural network is trained through the proximal policy optimization algorithm, and the trained deep neural network is used to realize the end-to-end mapping from the state of the Markov decision process to the action selection probability within the rule set, realizing multi-step action sequence optimization and obtaining the optimized scheduling strategy.
9. The multi-objective dual-hyper heuristic method for flexible job shop self-organizing scheduling according to claim 8, characterized in that: Deep reinforcement learning uses the proximal policy optimization algorithm for training. The objective function of the proximal policy optimization algorithm is: Where J(θ) represents the objective function, represents the expected value of the data at the tth time step, r t (θ) represents the action probability ratio of the new and old strategies, Indicates state s t Next select action a t The relative gain of r t (θ)>1+∈, it is truncated to 1+∈, when r t When (θ)<1-∈, it is truncated to 1-∈, otherwise it is retained r t (θ),∈ represents the cropping coefficient, π θ Indicates the current strategy, Indicates the old policy.
10. The multi-objective dual-hyper heuristic method for flexible job shop self-organizing scheduling according to claim 1 is characterized in that: The example of randomly combining and generating parameters of predetermined dynamic events includes: Taking the arrival of new jobs, machine failures, and changes in job processing time as dynamic events, we set several parameters for machine failures, several parameters for changes in job processing time, and several parameters for the arrival of new workpieces, and then randomly generate instances through random combinations.