Hybrid flow shop scheduling method based on double-agent neighborhood search algorithm

Through the dual agent neighborhood search algorithm, Agent1 and Agent2 agents are built, which solves the real-time response problem of hybrid flow workshop scheduling in a dynamic environment, realizes efficient production scheduling and flexible scheduling strategy generation, and improves production efficiency.

CN120278448AActive Publication Date: 2025-07-08JINAN UNIVERSITY

Patent Information

Application Number
CN202510351673.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-08
Estimated Expiration
2045-03-24

AI Technical Summary

Technical Problem

The existing hybrid flow workshop scheduling algorithms are difficult to respond to environmental changes in real time in dynamic production environments, resulting in low production efficiency. Traditional precision algorithms have high computational complexity, and heuristic and metaheuristic algorithms lack flexibility, making it difficult to cope with dynamic changes in workshop environments.

Method used

Using the dual agent neighborhood search algorithm, two agents Agent1 and Agent2 are constructed. Agent1 determines the artifact sequence, and Agent2 selects the optimal neighborhood search strategy, combining adaptive exploration mechanisms and dynamic acceptance criteria, an efficient scheduling scheme is generated through real-time data interaction.

Benefits of technology

It realizes rapid response and dynamic adjustment of scheduling strategies in a dynamic environment, significantly improving production efficiency and scheduling flexibility, and is better than traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278448A_ABST
    Figure CN120278448A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of production scheduling, and discloses a hybrid flow shop scheduling method based on a double-agent neighborhood search algorithm, and the method comprises the following specific steps: 1, researching a real production scene of a hybrid flow shop, mining constraint conditions, and analyzing a main scheduling problem which mainly affects a production cycle; the invention integrates the advantages of a deep Q network (DQN) and a meta-heuristic algorithm, provides a double-agent neighborhood search algorithm, can quickly respond to the change of a workshop environment and dynamically adjust a scheduling strategy, and aims to provide a new efficient solution scheme for hybrid flow shop scheduling and derivative problems thereof. Two intelligent agents capable of interacting with the workshop environment in real time are constructed, wherein the first intelligent agent can determine the machining sequence of workpieces in each stage, and a promising initial solution is rapidly generated; and the second agent can select an optimal neighborhood search strategy, and dynamically adjust the search direction to accelerate convergence of the algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of production scheduling, and specifically relates to a hybrid flow shop scheduling method based on a double-agent neighborhood search algorithm. Background Art

[0002] The hybrid flow shop scheduling problem and its variants are common and complex combinatorial optimization problems in industrial production systems, involving multiple jobs, multiple stages, and multiple machines. Such problems combine the characteristics of traditional flow shop scheduling and parallel machine scheduling, and have significant NP-hard characteristics. HFSP widely exists in multiple industries such as chemical production, electronic assembly, biopharmaceuticals, and circuit board printing. By optimizing the scheduling of production tasks and resources, the production efficiency and operation effect of related manufacturing enterprises can be significantly improved.

[0003] Currently, most traditional scheduling algorithms rely on fixed rules or strategies to determine job sequences and machine allocations. These methods usually ignore the dynamic changes in the workshop environment as production tasks progress. As the production environment continues to change, existing scheduling rules cannot respond to these changes in real time, resulting in low production efficiency. Therefore, how to adjust the scheduling strategy in real time in a dynamic production environment and respond to environmental changes is the key challenge in solving HFSP and its derivative problems. Through the research of domestic and foreign scholars on HFSP, its solution methods can be divided into exact algorithms, heuristic algorithms, meta-heuristic algorithms, and algorithms based on deep or reinforcement learning. Exact algorithms such as branch and bound method, dynamic programming, and integer linear programming (ILP) calculate all possible scheduling schemes and use mathematical models and constraints to find the optimal solution. However, as the problem scale increases, the computational complexity of these exact algorithms grows exponentially, making it impossible for them to effectively solve large-scale problems. In this case, heuristic and meta-heuristic algorithms are widely used to balance the solution quality and computational time. Heuristic algorithms can quickly generate feasible solutions but usually cannot guarantee global optimality, while meta-heuristic algorithms can find solutions close to the global optimum in a shorter time through more complex search mechanisms and perturbation strategies. However, most existing meta-heuristic algorithms are based on fixed rules, lack flexibility, and are difficult to respond to the dynamic changes in the workshop environment in a timely manner, thereby having a negative impact on production efficiency. Summary of the Invention

[0004] The purpose of the present invention is to provide a hybrid flow shop scheduling method based on a double-agent neighborhood search algorithm to solve the problems raised in the above background art.

[0005] To achieve the above purpose, the present invention provides the following technical solution: A hybrid flow shop scheduling method based on a double-agent neighborhood search algorithm, and the specific steps are as follows:

[0006] Step 1: Study the actual production scenario of the hybrid flow shop, identify the constraints, and analyze the main scheduling problems that mainly affect the production cycle:

[0007] In the HFSP workshop, there are J jobs that need to pass through K processing stages in sequence. There are M k ≥ 1 selectable machines at each stage, and at least one stage satisfies M k > 1, k ∈ K. At each stage, job j can and can only select one machine for processing. After completing the processing of the current stage, the task immediately enters the subsequent processing stage, j ∈ J. All jobs and machines are ready at time 0. A machine can only process one job at the same time and the processing process is not allowed to be interrupted. The processing time of all jobs at each stage is known, and the transportation time between the front and back stages of the job and the setup time of the machine are not considered. The optimization goal is to minimize the production cycle, that is, to obtain the minimized makespan. The main scheduling task is to determine the job sequence and machine allocation at each stage;

[0008] Step 2: According to the problem characteristics, design appropriate encoding and decoding strategies to symbolically represent the solution:

[0009] Encoding is the representation method of the solution. The main task of hybrid flow shop scheduling is to determine the job sequence and machine allocation at each stage. A two-dimensional matrix Z with K rows and J columns is used K×J to represent the solution, where K is the number of stages and J is the number of jobs. For example, for a workshop with 3 stages and 5 jobs, its solution is represented as follows:

[0010]

[0011] Decoding is to convert the solution into an actual scheduling plan. For the job sequence at each stage, the position of job j in the k-th row of the matrix Z K×J represents its processing order at the k-th stage. For example, according to the matrix Z 3×5 , the job sequence in the first stage is {3, 2, 1, 5, 4}, that is, the processing order of the jobs in the first stage is job 3, job 2, job 1, job 5, and job 4 in turn. For machine allocation, the "earliest idle" rule is adopted, that is, the job is preferentially allocated to the machine that is in the idle state earliest for processing. If there are multiple idle machines at the same time, it is randomly allocated to one of them for processing;

[0012] Step 3: Construct an intelligent agent based on the Markov decision process, and design the state space, action space, state transition, reward function, and neural network architecture:

[0013] Two agents, Agent1 and Agent2, are constructed based on the Markov decision process. Among them, Agent1 determines the workpiece sequence at each stage according to the characteristics of the workpiece to be processed at each decision moment, so as to quickly generate an initial solution. Agent2 selects the optimal neighborhood perturbation strategy according to the workshop environment under the current solution, driving the solution to evolve in a potential direction;

[0014] Step 4: Introduce an adaptive exploration mechanism and a dynamic acceptance criterion to balance the exploitation and exploration capabilities of the algorithm:

[0015] During the training process of agents Agent1 and Agent2, to prevent the agents from always choosing the action with the highest Q value and over-relying on specific state-action combinations, thus restricting their exploration of other strategies and affecting the ability to identify potential optimal solutions, an adaptive exploration mechanism based on the ε-greedy strategy is introduced to balance the action selection probability of the agents;

[0016] At each decision point, a random number τ∈[0,1] is generated. If τ is greater than the threshold ε, the agent will choose the action with the highest Q value in the current state. On the contrary, if τ is less than or equal to ε, the agent will randomly select an action from the available actions. The threshold ε changes linearly during the training process, where E represents the total number of training iterations and e represents the current training number; as the training progresses, ε gradually decreases, while encouraging the algorithm to make more use of the learned knowledge, still retaining a certain exploration ability. This adaptive exploration mechanism ensures the balance of the learning process, promoting both the identification of the optimal strategy and the exploration of new possibilities;

[0017] Similar to the above adaptive exploration mechanism, to balance the local search ability and global search ability of the algorithm, a dynamic acceptance criterion is constructed. At the end of each iteration of the algorithm, the current solution is compared with the optimal solution. If the current solution is better than the optimal solution, the current solution is used to update the optimal solution and used as the initial solution for the next iteration of evolution. Otherwise, a random number δ∈[0,1] is generated;

[0018] Step 5: Train the agents to construct a double-agent neighborhood search algorithm:

[0019] Train agents Agent1 and Agent2. Their parameter settings are different, but the training steps are the same. First, initialize the experience replay buffer buffer_m and the reward buffer buffer_r, and construct the evaluation network Q and the target network Q * , Q = Q * , and then start the loop training. In each training episode, the agent first resets the environment and initializes the cumulative reward r c , at each decision point t, the agent is based on the current state st Select action a t , and execute action a t Make the agent enter the next state s t+1 , calculate the immediate reward r t , and transfer the state (s t , a t , r t , s t+1 ) and store it in the experience replay buffer. Once the buffer_m reaches the predefined threshold, randomly sample a batch of data (s t , a t , r t , s t+1 ), use the evaluation network Q and the target network Q * to calculate the Q value and the loss function Δloss. Then, use the Adam optimizer to minimize Δloss and update the evaluation network Q. At the end of each training, calculate the average reward r of the last 10 episodes ave , and compare it with the recorded best average reward r best If r ave > r best , then save the current target network Q * ;

[0020] Step 6: Generate an efficient scheduling scheme to achieve real-time scheduling and dynamic adjustment of production, and compare the proposed method with common methods to verify the efficiency of the proposed method:

[0021] During the workshop production process, sensors monitor the workshop equipment, workpieces, and environmental data in real time and feed this data back to the workshop scheduling system. The scheduling system analyzes the current real-time data and uses the double-agent neighborhood search algorithm to obtain the optimal operation in the current workshop environment, and then generates an optimal scheduling decision scheme. Subsequently, the workshop scheduling system executes this decision scheme to achieve efficient scheduling planning for workshop production and shorten the production cycle of workpieces;

[0022] To further verify the efficiency of the scheduling method based on the double-agent neighborhood search algorithm for solving the hybrid flow shop scheduling problem, compare it with the commonly used genetic algorithm (GA) and iterative greedy algorithm (IG) for solving such problems on 100 simulation instances.

[0023] As a preferred technical solution of the present invention, the state space design method of Agent1 and Agent2 described in step 3 is as follows: In HFSP, there are K stages and J workpieces, and there are a total of K×J decision time points. At each decision point, Agent1 can only schedule the workpieces that have reached the current stage. According to For the processing time, arrival time, and waiting time of the workpiece in the middle, 12 states are designed for Agent1; Agent2 selects the best neighborhood search strategy according to the workshop environment under the current solution, and 6 states are designed for it according to the workpiece and machine state characteristics.

[0024] As a preferred technical solution of the present invention, the action space design methods of Agent1 and Agent2 in step three are as follows: The action of Agent1 is to select a workpiece from the set The strategy of selecting a workpiece, according to the set For the processing time, arrival time, and waiting time of the workpiece in the middle, 9 workpiece selection rules (JSR) are constructed; Agent2 mainly selects the optimal neighborhood search strategy according to the workshop environment under the current solution to generate a new solution, and these neighborhood search strategies constitute the action space of Agent1. Based on the swap and insert operators, four neighborhood search strategies (NS) are designed for it.

[0025] As a preferred technical solution of the present invention, the reward function design methods of Agent1 and Agent2 in step three are as follows: The reward function is set to minimize the difference in the system completion time between two consecutive decisions. The reward r at time t t is the negative value of the difference between these two completion times, that is where represents the completion time at decision time t, represents the completion time at decision time t - 1.

[0026] As a preferred technical solution of the present invention, the state transition design methods of Agent1 and Agent2 in step three are as follows: At each decision point t, the agent, according to the current state s t , executes the action a t , and obtains the reward r t from the environment as feedback. This interaction causes the system to transfer to a new state s t+1 . In Agent1 and Agent2, a quadruple (s t , a t , r t , s t+1 ) is used to represent the state transition and is stored in the experience pool for subsequent training and decision-making.

[0027] As a preferred technical solution of the present invention, the network structure design methods of Agent1 and Agent2 in step three are as follows: Both Agent1 and Agent2 use a five-layer fully connected neural network. The input of the network is the state characteristics of the agent, and the output corresponds to the action space.

[0028] As a preferred technical solution of the present invention, in the dynamic acceptance criterion described in step four, where T represents the total running time of the algorithm, and t represents the currently elapsed running time.

[0029] As a preferred technical solution of the present invention, the specific comparison method for 100 simulation instances in step six is as follows: 100 simulation instances are collected, with stage scales of {3, 5, 8, 10} respectively and workpiece quantities of {40, 60, 80, 100, 120} respectively, resulting in a total of 4×5 = 20 combinations. Five different instances are generated for each combination, so there are a total of 100 instances. In each instance, the number of machines in each stage is randomly generated within the range of [1 - 5], and the processing time is randomly generated within the range of [10 - 50]. For each instance, the method proposed in the present invention, the genetic algorithm, and the iterative greedy algorithm are each solved 10 times. The average objective value of the 10 solutions of each algorithm is calculated, and the average value (Ave) of the 5 instances in each combination is calculated. At the same time, the relative percentage increase (RPI) index is introduced for measurement.

[0030] As a preferred technical solution of the present invention, the measurement formula for the relative percentage increase (RPI) index is:

[0031]

[0032] where C best is the best objective value obtained by all algorithms, and C avg is the average value of 10 solutions of each algorithm.

[0033] The beneficial effects of the present invention are as follows:

[0034] The present invention integrates the advantages of the deep Q - network (DQN) and meta - heuristic algorithms, and proposes a double - agent neighborhood search algorithm. This algorithm can quickly respond to changes in the workshop environment and dynamically adjust the scheduling strategy, aiming to provide a new efficient solution for the hybrid flow - shop scheduling and its derivative problems; the present invention constructs two agents that can interact with the workshop environment in real - time: the first agent can determine the processing sequence of workpieces in each stage and quickly generate a promising initial solution; the second agent can select the optimal neighborhood search strategy and dynamically adjust the search direction to accelerate the convergence of the algorithm. At the same time, in order to reduce the agent's dependence on specific state - actions, the present invention designs an adaptive exploration mechanism based on the ε - greedy strategy to enhance the agent's exploration ability in the entire state space. In addition, a dynamic acceptance criterion is introduced at the end of the iterative evolution process to balance the local and global exploration abilities of the algorithm. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 is the flow chart of the present invention;

[0036] Figure 2 Schematic diagram of the state space table of Agent1 of the present invention;

[0037] Figure 3 Schematic diagram of the state space table of Agent2 of the present invention;

[0038] Figure 4 Schematic diagram of the action space table of Agent1 of the present invention;

[0039] Figure 5 Schematic diagram of the algorithm result table of the present invention;

[0040] Figure 6 Flow chart of the double-agent neighborhood search algorithm of the present invention;

[0041] Figure 7 Confidence interval for the algorithm comparison of the present invention. Detailed implementation manners

[0042] Next, in combination with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0043] As Figures 1 to 7 shown, the embodiment of the present invention provides a hybrid flow shop scheduling method based on a double-agent neighborhood search algorithm, and the specific steps are as follows:

[0044] Step 1: Study the actual production scenario of the hybrid flow shop, dig out the constraint conditions, and analyze the main scheduling problems that mainly affect the production cycle:

[0045] In the HFSP workshop, there are J workpieces that need to pass through K processing stages in sequence. There are M k ≥1 selectable machines at each stage, and at least one stage satisfies M k >1, k ∈ K. At each stage, workpiece j can and can only select one machine for processing. After completing the processing of the current stage, the task immediately enters the subsequent processing stage, j ∈ J. All workpieces and machines are ready at time 0. A machine can only process one workpiece at the same time and the processing process is not allowed to be interrupted. The processing time of all workpieces at each stage is known, and the transportation time between the front and back two stages of the workpieces and the preparation time of the machines are not considered. The optimization goal is to minimize the production cycle, that is, to obtain the minimized makespan. The main scheduling task is to determine the workpiece sequence and machine allocation at each stage;

[0046] Step 2: According to the problem characteristics, design appropriate encoding and decoding strategies to symbolically represent the solution:

[0047] Encoding is a representation method of the solution. The main task of the hybrid flow shop scheduling is to determine the workpiece sequence and machine allocation at each stage. A two-dimensional matrix Z with K rows and J columns is used K×J to represent the solution, where K is the number of stages and J is the number of workpieces. For example, for a workshop with 3 stages and 5 workpieces, its solution is represented as follows:

[0048]

[0049] Decoding is to transform the solution into an actual scheduling plan. For the workpiece sequence at each stage, the position of workpiece j in the k-th row of the matrix Z K×J represents its processing order at the k-th stage. For example, according to the matrix Z 3×5 , the workpiece sequence in the first stage is {3, 2, 1, 5, 4}, that is, the processing order of the workpieces in the first stage is workpiece 3, workpiece 2, workpiece 1, workpiece 5, and workpiece 4 in sequence. For machine allocation, the "earliest idle" rule is adopted, that is, the workpiece is preferentially allocated to the machine that is in the idle state earliest for processing. If there are multiple idle machines at the same time, it is randomly allocated to one of them for processing;

[0050] Step 3: Construct an intelligent agent based on the Markov decision process, and design the state space, action space, state transition, reward function, and neural network architecture:

[0051] Two intelligent agents are constructed based on the Markov decision process: Agent1 and Agent2. Among them, Agent1 determines the workpiece sequence at each stage according to the characteristics of the workpiece to be processed at each decision moment, so as to quickly generate an initial solution. Agent2 selects the optimal neighborhood perturbation strategy according to the workshop environment under the current solution, and drives the solution to evolve in a potential direction;

[0052] Step 4: Introduce an adaptive exploration mechanism and a dynamic acceptance criterion to balance the exploitation and exploration capabilities of the algorithm:

[0053] During the training process of intelligent agents Agent1 and Agent2, to prevent the intelligent agents from always choosing the action with the highest Q value and overly relying on specific state-action combinations, thus restricting their exploration of other strategies and affecting the ability to identify potential optimal solutions, an adaptive exploration mechanism based on the ε-greedy strategy is introduced to balance the action selection probability of the intelligent agents;

[0054] At each decision point, a random number τ∈[0,1] is generated. If τ is greater than the threshold ε, the agent will choose the action with the highest Q value in the current state. On the contrary, if τ is less than or equal to ε, the agent will randomly select an action from the optional actions. The threshold ε changes linearly during the training process. Where E represents the total number of training iterations, and e represents the current number of training iterations. As training progresses, ε gradually decreases, which encourages the algorithm to make more use of the learned knowledge while retaining a certain exploration ability. This adaptive exploration mechanism ensures the balance of the learning process, which not only promotes the identification of the optimal strategy, but also encourages the exploration of new possibilities.

[0055] Similar to the above-mentioned adaptive exploration mechanism, in order to balance the local search ability and global search ability of the algorithm, a dynamic acceptance criterion is constructed. At the end of each iteration of the algorithm, the current solution is compared with the optimal solution. If the current solution is better than the optimal solution, the optimal solution is updated with the current solution and used as the initial solution for the next iterative evolution. Otherwise, a random number δ∈[0,1] is generated.

[0056] Step 5: Train the agent and construct a dual-agent neighborhood search algorithm:

[0057] Train Agent 1 and Agent 2. Their parameter settings are different, but the training steps are the same. First, initialize the experience playback buffer buffer_m and the reward buffer buffer_r, and build the evaluation network Q and the target network Q. * , Q=Q * Next, the training cycle begins. In each training round, the agent must first reset the environment and initialize the cumulative reward r c At each decision point t, the agent takes the current state s as t Select action a t , perform action a t Make the agent enter the next state s t+1 , calculate the immediate reward r t , and transfer the state (s t ,a t ,r t ,s t+1 ) is stored in the experience replay buffer. Once buffer_m reaches a predefined threshold, a batch of data (s) is randomly sampled from it. t ,a t ,r t ,s t+1 ), using the evaluation network Q and the target network Q * Calculate the Q value and loss function Δloss, then use the Adam optimizer to minimize Δloss and update the evaluation network Q. At the end of each training, calculate the average reward r of the last 10 roundsave and compare it with the recorded best average reward r best for comparison. If r ave > r best , then save the current target network Q * ;

[0058] Step 6: Generate an efficient scheduling scheme to achieve real-time scheduling and dynamic adjustment of production, and compare the proposed method with common methods to verify the efficiency of the proposed method:

[0059] During the workshop production process, sensors continuously monitor the data of workshop equipment, workpieces, and the environment, and feed this data back to the workshop scheduling system. The scheduling system analyzes the current real-time data and uses the double-agent neighborhood search algorithm to obtain the optimal operation under the current workshop environment, and then generates an optimal scheduling decision scheme. Subsequently, the workshop scheduling system executes this decision scheme to achieve efficient scheduling planning for workshop production and shorten the production cycle of workpieces;

[0060] To further verify the efficiency of the scheduling method based on the double-agent neighborhood search algorithm for solving the hybrid flow shop scheduling problem, it is compared with the commonly used genetic algorithm (GA) and iterative greedy algorithm (IG) for solving such problems on 100 simulation instances.

[0061] The present invention combines the real-time decision-making ability of DQN with the iterative evolution process of meta-heuristic algorithms, proposes a double-agent neighborhood search algorithm, and constructs an intelligent scheduling method based on this. Through the real-time interaction between two agents and the production environment, this method can quickly generate high-quality scheduling strategies in complex dynamic environments, significantly improve the flexibility and production efficiency of scheduling, provide a new technical path for solving HFSP and its derivative problems, and has important application value.

[0062] Among them, the state space design methods of Agent1 and Agent2 in Step 3 are as follows: In HFSP, there are K stages and J workpieces, with a total of K×J decision time points. At each decision point, Agent1 can only schedule the workpieces that have reached the current stage. Based on the processing time, arrival time, and waiting time of the workpieces in, 12 states are designed for Agent1; Agent2 selects the best neighborhood search strategy according to the workshop environment under the current solution, and 6 states are designed for it based on the workpiece and machine state characteristics.

[0063] Among them, represents the set of workpieces that have reached stage k and have not been processed.

[0064] Among them, the action space design methods of Agent1 and Agent2 in Step 3 are as follows: The action of Agent1 is to select from the set The strategy for selecting workpieces, according to the set of the processing time, arrival time, and waiting time of the workpieces in it, constructs 9 workpiece selection rules (JSRs); Agent2 mainly selects the optimal neighborhood search strategy according to the shop floor environment under the current solution to generate a new solution. These neighborhood search strategies constitute the action space of Agent1. Based on the swap and insert operators, four neighborhood search strategies (NSs) are designed for it.

[0065] The four neighborhood search strategies (NSs) are as follows: NS1 (random swap): In the random selection stage, randomly select two workpieces from the workpiece sequence and swap their positions in the sequence; NS2 (max-min swap): In the first stage, swap the positions of the workpiece with the longest total completion time and the workpiece with the shortest total completion time in the sequence; NS3 (max insert): Randomly select a workpiece on the machine with the largest total processing time, remove it from the sequence at the current stage, and then insert it into other positions in the sequence in a random manner; NS4 (min insert): Randomly select a workpiece on the machine with the smallest total processing time, remove it from the sequence at the current stage, and then insert it into other positions in the sequence in a random manner. It should be noted that when the workpiece sequence at a certain stage changes, the workpiece sequences in the subsequent stages must also be adjusted accordingly. If the workpiece sequence at stage k changes, the workpiece sequences in the subsequent stages will be redefined by Agent1.

[0066] Among them, the design method of the reward functions of Agent1 and Agent2 in step three is: The reward function is set to minimize the difference in the system completion time between two consecutive decisions. The reward r at time t t is the negative value of the difference between these two completion times, that is where represents the completion time at decision time t, represents the completion time at decision time t - 1.

[0067] Agent1 optimizes the scheduling process by reducing the difference in the completion time during consecutive system decisions. Therefore, based on this, the reward function is set to minimize the difference in the system completion time between two consecutive decisions.

[0068] Among them, the design method of the state transition of Agent1 and Agent2 in step three is: At each decision point t, the agent, according to the current state s t , executes the action a t , and obtains the reward r t from the environment as feedback. This interaction causes the system to transfer to a new state s t+1 . In Agent1 and Agent2, a quadruple (s t , at , r t , s t+1 ) is used to represent the state transition and store it in the experience pool for subsequent training and decision-making.

[0069] In reinforcement learning, state transition refers to the process by which an agent moves from one state to another through an action.

[0070] Among them, the network structure design methods of Agent1 and Agent2 in step three are as follows: both Agent1 and Agent2 use a five-layer fully connected neural network. The input of the network is the state features of the agent, and the output corresponds to the action space.

[0071] To enhance the network's ability to capture complex feature interactions, all hidden layers use the non-linear activation function ReLU, and the number of neurons in these hidden layers is set to 256, 128, 64, and 32 respectively.

[0072] Among them, in the dynamic acceptance criterion of step four, Among them, T represents the total running time of the algorithm, and t represents the current running time.

[0073] In the dynamic acceptance criterion, if δ≥ε, the current solution is used in the next iteration. If δ<ε, the agent generates a new solution and uses the new solution in the next iteration.

[0074] Among them, the specific comparison method for 100 simulation instances in step six is as follows: 100 simulation instances are collected, with stage scales of {3, 5, 8, 10} and workpiece quantities of {40, 60, 80, 100, 120} respectively, resulting in a total of 4×5 = 20 combinations. Five different instances are generated for each combination, so there are a total of 100 instances. In each instance, the number of machines in each stage is randomly generated within the range of [1 - 5], and the processing time is randomly generated within the range of [10 - 50]. For each instance, the method proposed in the present invention, the genetic algorithm, and the iterative greedy algorithm are each solved 10 times. The average objective value of the 10 solutions of each algorithm is calculated, and the average value (Ave) of the 5 instances in each combination is calculated. At the same time, the relative percentage increase (RPI) index is introduced for measurement.

[0075] According to the algorithm comparison results and the confidence intervals of the result data, it can be seen that the dual-agent neighborhood search algorithm proposed in the present invention can effectively solve the hybrid flow shop scheduling problem and is significantly superior to the genetic algorithm and the iterative greedy algorithm.

[0076] Among them, the measurement formula for the relative percentage increase (RPI) index is:

[0077]

[0078] Among them, C best is the best objective value obtained by all algorithms, and C avg is the average value of 10 solutions for each algorithm.

[0079] The Relative Percentage Increase (RPI) is a metric used to measure how much a value has increased as a percentage relative to its original or benchmark value. This type of measurement is widely used in various fields such as economics, business, education, and healthcare to evaluate data changes at different points in time or under different conditions.

[0080] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variation thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device.

[0081] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made in these embodiments without departing from the principles and spirit of the invention, and the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A hybrid flow shop scheduling method based on a double-agent neighborhood search algorithm, characterized in that The specific steps are as follows: Step 1: Study the actual production scenarios of the hybrid flow shop, identify the constraint conditions, and analyze the main scheduling problems that mainly affect the production cycle: In the HFSP workshop, there are J jobs that need to pass through K processing stages in sequence. At each stage, there are M k ≥ 1 selectable machines, and at least one stage satisfies M k > 1, k ∈ K. At each stage, job j can and can only select one machine for processing. After completing the current stage of processing, the task immediately enters the subsequent processing stage, j ∈ J. All jobs and machines are ready at time 0. A machine can only process one job at the same time and the processing process is not allowed to be interrupted. The processing time of all jobs at each stage is known, and the transportation time between two consecutive stages of a job and the setup time of the machine are not considered. The optimization goal is to minimize the production cycle, that is, to obtain the minimized makespan. The main scheduling task is to determine the job sequence and machine allocation at each stage; Step 2: According to the problem characteristics, design appropriate encoding and decoding strategies to symbolically represent the solution: Coding is a representation of the solution. The main task of the hybrid flow shop scheduling is to determine the workpiece sequence and machine allocation at each stage, and a two-dimensional matrix Z with K rows and J columns is used K×J to represent the solution, where K is the number of stages and J is the number of workpieces. For example, for a workshop with 3 stages and 5 workpieces, its solution is represented as follows: Decoding is to transform the solution into an actual scheduling plan. For the workpiece sequence in each stage, the position of workpiece j in matrix Z K×J in the k-th row indicates its processing order in the k-th stage. For example, according to matrix Z 3×5 , the workpiece sequence in the first stage is {3, 2, 1, 5, 4}, that is, the processing order of the workpieces in the first stage is workpiece 3, workpiece 2, workpiece 1, workpiece 5, and workpiece 4 in sequence. For machine allocation, the "earliest idle" rule is adopted, that is, the workpiece is preferentially allocated to the machine that is in the idle state earliest for processing. If there are multiple idle machines at the same time, it is randomly allocated to one of them for processing; Step 3: Construct an agent based on the Markov decision process, and design the state space, action space, state transition, reward function, and neural network architecture: Two agents are constructed based on the Markov decision process: Agent1 and Agent2. Among them, Agent1 determines the workpiece sequence at each stage according to the characteristics of the workpiece to be processed at each decision moment, so as to quickly generate an initial solution. Agent2 selects the optimal neighborhood perturbation strategy according to the shop floor environment under the current solution, driving the solution to evolve in a promising direction; Step 4: Introduce an adaptive exploration mechanism and a dynamic acceptance criterion to balance the exploitation and exploration capabilities of the algorithm: During the training process of agents Agent1 and Agent2, to prevent the agents from always choosing the action with the highest Q value and relying too much on a specific state-action combination, thereby restricting their exploration of other strategies and affecting their ability to identify potential optimal solutions, an adaptive exploration mechanism based on the ε-greedy strategy is introduced to balance the action selection probability of the agents; At each decision point, a random number τ ∈ [0, 1] is generated. If τ is greater than the threshold ε, the agent will choose the action with the highest Q-value in the current state. On the contrary, if τ is less than or equal to ε, the agent will randomly select an action from the available actions. The threshold ε changes linearly during the training process. where E represents the total number of training iterations and e represents the current training iteration; as the training progresses, ε gradually decreases, while encouraging the algorithm to make more use of the learned knowledge, still retaining a certain exploration ability. This adaptive exploration mechanism ensures the balance of the learning process, promoting both the identification of the optimal strategy and the exploration of new possibilities. Similar to the above adaptive exploration mechanism, to balance the local search ability and global search ability of the algorithm, a dynamic acceptance criterion is constructed. At the end of each iteration of the algorithm, the current solution is compared with the optimal solution. If the current solution is better than the optimal solution, the optimal solution is updated with the current solution and used as the initial solution for the next iteration of evolution. Otherwise, a random number δ∈[0,1] is generated; Step 5: Train the agents and construct a two-agent neighborhood search algorithm: Train agents Agent1 and Agent2. They have different parameter settings but the same training steps. First, initialize the experience replay buffer buffer_m and the reward buffer buffer_r, and construct the evaluation network Q and the target network Q * , Q = Q * , Next, start the loop training. In each training episode, the agent first resets the environment and initializes the cumulative reward r c , At each decision point t, the agent selects an action a based on the current state s t Select action a t , Execute action a t Make the agent enter the next state s t+1 , Calculate the immediate reward r t , and transfer the state (s t , a t , r t , s t+1 ) and store it in the experience replay buffer. Once buffet_m reaches the predefined threshold, randomly sample a batch of data (s t , a t , r t , s t+1 ), Use the evaluation network Q and the target network Q * Calculate the Q value and the loss function Δloss. Then, use the Adam optimizer to minimize Δloss and update the evaluation network Q. At the end of each training, calculate the average reward r of the last 10 episodes ave , and compare it with the recorded best average reward r best , If r ave > r best , then save the current target network Q * ; Step 6: Generate an efficient scheduling plan, realize real-time scheduling and dynamic adjustment of production, and compare the proposed method with common methods to verify the efficiency of the proposed method: During the shop floor production process, sensors monitor the shop floor equipment, workpieces, and environmental data in real time and feed this data back to the shop floor scheduling system. The scheduling system analyzes the current real-time data and uses the two-agent neighborhood search algorithm to obtain the optimal operation under the current shop floor environment, and then generates an optimal scheduling decision plan. Subsequently, the shop floor scheduling system executes this decision plan to achieve efficient scheduling planning of shop floor production to shorten the production cycle of workpieces; To further verify the efficiency of the scheduling method based on the two-agent neighborhood search algorithm for solving the hybrid flow shop scheduling problem, it is compared with the commonly used genetic algorithm (GA) and iterative greedy algorithm (IG) for solving such problems on 100 simulation instances.

2. A hybrid flow shop scheduling method based on a double-agent neighborhood search algorithm according to claim 1, characterized in that: The state space design methods of Agent1 and Agent2 described in Step 3 are as follows: In HFSP, there are K stages and J jobs, with a total of K×J decision time points. At each decision point, Agent1 can only schedule the jobs that have reached the current stage. Based on the processing time, arrival time, and waiting time of the jobs in it, 12 states are designed for Agent1; Agent2 selects the best neighborhood search strategy according to the workshop environment under the current solution, and 6 states are designed for it according to the job and machine state characteristics.

3. A hybrid flow shop scheduling method based on a double-agent neighborhood search algorithm according to claim 1, characterized in that: The design method of the action spaces of Agent1 and Agent2 in step three is as follows: The action of Agent1 is to select a workpiece from the set and construct nine job selection rules (JSRs) according to the processing time, arrival time, and waiting time of the workpieces in the set . Agent2 mainly selects the optimal neighborhood search strategy according to the shop floor environment under the current solution to generate a new solution. These neighborhood search strategies constitute the action space of Agent1. Based on the swap and insert operators, four neighborhood search strategies (NS) are designed for it.

4. A hybrid flow shop scheduling method based on a double-agent neighborhood search algorithm according to claim 1, characterized in that: The design methods of the reward functions for Agent1 and Agent2 in Step 3 are as follows: The reward function is set to minimize the difference in the system completion time between two consecutive decisions. The reward r at time t t is the negative value of this difference in the two completion times, that is where represents the completion time at decision time t, and represents the completion time at decision time t - 1.

5. A hybrid flow shop scheduling method based on a double-agent neighborhood search algorithm according to claim 1, characterized in that: The state transition design methods of Agent1 and Agent2 described in Step 3 are as follows: at each decision point t, the agent executes action a t according to the current state s t , and obtains a reward r t from the environment as feedback. This interaction causes the system to transition to a new state s t+1 . In Agent1 and Agent2, the quadruple (s t , a t , r t , s t+1 ) is used to represent the state transition and stored in the experience pool for subsequent training and decision-making.

6. A hybrid flow shop scheduling method based on a double-agent neighborhood search algorithm according to claim 1, characterized in that: The network structure design methods of Agent1 and Agent2 described in Step 3 are as follows: Both Agent1 and Agent2 use a five-layer fully connected neural network. The input of the network is the state characteristics of the agent, and the output corresponds to the action space.

7. A hybrid flow shop scheduling method based on a double-agent neighborhood search algorithm according to claim 1, characterized in that: In the dynamic acceptance criterion described in Step 4, where T represents the total running time of the algorithm, and t represents the current running time.

8. A hybrid flow shop scheduling method based on a double-agent neighborhood search algorithm according to claim 1, characterized in that: The specific comparison method for the 100 simulation examples in Step 6 is as follows: 100 simulation examples were collected, with stage scales of {3, 5, 8, 10} and workpiece quantities of {40, 60, 80, 100, 120}, resulting in a total of 4×5 = 20 combinations. Five different examples were generated for each combination, so there were a total of 100 examples. In each example, the number of machines in each stage was randomly generated within the range of [1 - 5], and the processing time was randomly generated within the range of [10 - 50]. For each example, the method proposed in the present invention, the genetic algorithm, and the iterative greedy algorithm were each solved 10 times. The average objective value of the 10 solutions of each algorithm was calculated, and the average value (Ave) of the 5 examples in each combination was also calculated. At the same time, the relative percentage increase (RPI) index was introduced for measurement.

9. A hybrid flow shop scheduling method based on a double-agent neighborhood search algorithm according to claim 8, characterized in that: The measurement formula for the relative percentage increase (RPI) index is as follows: Among them, C best is the best objective value obtained by all algorithms, and C avg is the average value of 10 solutions for each algorithm.

Citation Information

Patent Citations

  • Hybrid flow shop scheduling method based on time sequence difference

    CN112734172A

  • Hyper-heuristic reinforcement learning scheduling method for distributed manufacturing of mechanical equipment

    CN116300748A

  • Intelligent collaborative optimization method integrating flexible process, dynamic reconstruction and active scheduling

    CN118378747A

Cited By

  • Multi-agent batch processing machine and variable sub-batch hybrid flow shop scheduling method

    CN121258124A

  • GPU resource intelligent scheduling method and system based on reinforcement learning

    CN121807508A