Flexible workshop production resource scheduling method based on multi-strategy deep reinforcement learning
Through multi-strategy deep reinforcement learning methods, a policy network and value network for operations, machines and AGVs are built, which solves the resource scheduling problem under dynamic events in the flexible workshop, and realizes efficient collaboration and real-time optimization of the production process.
Patent Information
- Application Number
- CN202510856937.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-25
AI Technical Summary
In flexible workshops with multiple varieties and small batch production, dynamic events occur frequently, and existing scheduling methods are difficult to respond quickly and ensure production stability and efficiency. A single priority scheduling rule cannot guarantee global optimization, resulting in production delays and resource waste.
Multi-strategy deep reinforcement learning methods are adopted to build a policy network and value network for homework selection, machine selection and AGV selection, simulate strategic collaboration through Markov decision-making process, design reward sharing mechanisms, and optimize resource scheduling decisions.
The flexible workshop production resource coordinated scheduling in a dynamically changing environment is realized, which improves the resource allocation efficiency of the production process and the real-time scheduling decision-making, avoids local optimal solutions, and improves production stability and efficiency.
Smart Images

Figure CN120355202A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of intelligent manufacturing, and particularly relates to a flexible workshop production resource scheduling method based on multi-strategy deep reinforcement learning. Background Art
[0002] Multi-variety and small-batch production has become an important mode in modern manufacturing. Under this background, manufacturing enterprises have increasingly higher requirements for product quality and manufacturing process refinement to meet consumers' demands for high-quality and personalized products. With the continuous improvement of the intelligent level of manufacturing workshops, Automated Guided Vehicles (AGVs) play a key role in the production process. They are responsible for the task of circularly picking up and transporting raw materials and semi-finished products between machines, and together with various machines, they constitute the workshop production resource system.
[0003] In such a production scenario, the collaborative scheduling of production resources is particularly important. The collaborative scheduling of production resources aims to effectively integrate and rationally allocate various resources including machines, AGVs, etc., to ensure the efficient and smooth operation of the production process. However, in modern customized manufacturing, dynamic events occur frequently, such as the random arrival of jobs, equipment failures, order cancellations or modifications, etc. These dynamic interferences may cause the original static scheduling scheme to deviate from the expectation during the execution process, thus significantly reducing the stability and efficiency of production. Therefore, when dynamic events occur, there is an urgent need to comprehensively and collaboratively re-plan and schedule production resources including machines and AGVs. It is necessary to comprehensively consider various factors such as the status of production resources and task priorities to ensure that the scheduling system can quickly respond to and adapt to various emergencies, thereby maintaining the efficient and stable operation of the production resource collaborative system and ensuring the smooth completion of production tasks.
[0004] To address the above challenges, researchers have proposed meta-heuristic algorithms (such as genetic algorithms and particle swarm algorithms), and these methods can provide approximate optimal solutions within a limited time. However, these algorithms still face the challenge of excessive calculation time when dealing with large-scale problems, and when the problem scale changes, a large number of iterative calculations need to be carried out again. In contrast, Priority Dispatching Rules (PDR) make scheduling decisions quickly through pre-set rules (such as first-come-first-served, shortest processing time, shortest transportation time, etc.), and can effectively respond to the real-time changes of dynamic events. However, a single PDR can only generate feasible solutions and cannot guarantee global optimality, which is prone to problems such as production delays, resource idleness or over-utilization, thus reducing production efficiency, increasing production costs and affecting the on-time delivery rate of orders.
[0005] As a cutting-edge technology in the field of artificial intelligence, Deep Reinforcement Learning (DRL) has demonstrated great potential in dealing with complex decision-making problems. Through continuous interaction between the agent and the environment, DRL can autonomously learn and optimize decision-making strategies to maximize long-term goals. Multi-policy deep reinforcement learning is an advanced optimization method based on the DRL framework, which integrates multiple policies to handle complex decision-making problems. In traditional reinforcement learning, the agent usually relies on a single policy for decision-making, while the multi-policy method trains multiple policies in parallel, making it more adaptable and flexible in different situations. Specifically, multi-policy DRL can flexibly select the optimal policy or balance between multiple policies when facing a dynamically changing environment by introducing a policy collaboration mechanism, so as to cope with multi-dimensional constraints and uncertainty factors. Applying the multi-policy deep reinforcement learning method to the collaborative scheduling of flexible job shop production resources can not only effectively handle the complex and changeable resource allocation and task scheduling problems in the job shop, but also optimize the collaborative effect between different types of production resources (such as machines, AGVs, etc.) in the job shop. By comprehensively considering multiple production factors such as processing time, order priority, and delivery deadline, multi-policy deep reinforcement learning can adjust the scheduling strategy in real time under the dynamically changing production demands of the flexible job shop to achieve the optimal resource allocation and scheduling decision in the production process. Summary of the Invention
[0006] To overcome the above deficiencies, the present invention provides a flexible job shop production resource scheduling method based on multi-policy deep reinforcement learning to achieve the collaborative scheduling of machines and AGVs. This method is based on a flexible job shop scheduling model with random job arrivals and uses the Markov decision process to simulate the decision-making and collaboration relationships between multiple policies. In addition, a reward sharing mechanism is designed, and the reward function is accurately constructed to enhance the collaborative effect between policies, which helps to solve the problem of the solution speed and solution quality under dynamic events.
[0007] To achieve the above object, the technical solution adopted by the present invention is as follows: A flexible job shop production resource scheduling method based on multi-policy deep reinforcement learning adopts a multi-policy collaboration mechanism to construct a decision-making model including a job selection module, an AGV selection module, and a machine selection module to schedule the flexible job shop production resources with random job arrivals. The specific steps are as follows: S1. Establish a job selection module, a machine selection module, and an AGV selection module based on Markov decision-making. The job selection module, the machine selection module, and the AGV selection module all use an optimization algorithm including a policy network and a value network to train an action selection agent; the parameters of the policy network and the value network of the job selection module, the machine selection module, and the AGV selection module are independent of each other; design a composite scheduling rule for job selection, machine selection, and AGV selection. S2. Initialize the network parameters of the agents in the job selection module, the machine selection module, and the AGV selection module, and conduct training. S3. During the training process, if there is a new job, add the new job to the set of unfinished jobs and then calculate the states of the job, the machine, and the AGV. Otherwise, directly calculate the states. S4. After the states are determined, perform multi-strategy collaborative scheduling for job selection, machine selection, and AGV selection respectively at each scheduling moment. The outputs of the job selection policy network, the machine selection policy network, and the AGV selection policy network are fed into a function to obtain the probability distribution of multiple job selection rules, randomly sample a probability value from the probability distribution, and obtain its corresponding index. According to this index, select the corresponding rule from the composite scheduling rule in step S1, map it to a specific operation according to the job state , map it to a transportation operation according to the AGV state the AGV used, and map it to the machine for processing operation according to the machine state of the machine. S5. Through the multi-strategy collaboration in step S4, determine the operation by transported to the machine for processing, then enter the next state, calculate the start and end times of the operation , calculate the reward of the agent according to the average delay time, and store the states, actions, and rewards of job selection, machine selection, and AGV selection in the corresponding experience pool. S6. Repeat steps S3, S4, and S5 until the operations of the newly arrived jobs and the existing jobs in the workshop in a single scheduling example are all processed. Then sample a single round of data from the experience pool, calculate the loss, and update the network parameters. After the experience pool is emptied, enter the next round of iteration.
[0008] For further optimization, residual blocks are introduced into both the policy network and the value network. Each residual block includes two fully connected layers, uses the ReLU activation function, and has a skip connection. Multiple residual blocks in the network are stacked. The output of the last residual block is passed to the fully connected layer of the network and the result is output; the output of the policy network is: , where, It represents the output of the last fully connected layer of the policy network. , represents the weight matrix and bias term of the fully connected layer. is the output of the last residual block of the policy network; In the value network, the output of its last residual block is passed to the fully connected layer of the network to calculate the state value : Among them, is the output of the last residual block in the value network. , represents the weight matrix and bias term of the fully connected layer of the value network.
[0009] For further optimization, colored noise is added to the operations of the job selection policy network, machine selection policy network, and AGV selection policy network. The colored noise is the noise term corresponding to the output dimension of the policy network, and this noise term is added to the original output of the policy network to form a noisy output. In the policy network, the colored noise is added to : , through to calculate the probability distribution of actions Among them, is the number of actions.
[0010] For further optimization, the job selection policy network, machine selection policy network, and AGV selection policy network are updated using a shared immediate reward.
[0011] For further optimization, the action space of the composite scheduling rule in step S1 includes a job scheduling rule, a machine scheduling rule, and an AGV scheduling rule.
[0012] For further optimization, the job scheduling rule is as follows: (1), If is empty, select the job with the minimum average slack time. Otherwise, select the job with the longest overdue time and high priority for processing; (2), If is empty, select the job with a small slack time critical ratio and high priority. Otherwise, select the job with the longest overdue time and high priority for processing; (3), If is empty, select the job with a low operation completion rate and high priority. Otherwise, select the job with the longest overdue time and high priority for processing; (4) To avoid falling into local optimality, randomly select a job for processing from the set of unfinished jobs; Among them At the rescheduling point Set of overdue jobs.
[0013] For further optimization, the machine scheduling rules are as follows: (1) Select the earliest available machine; (2) Select the machine with the shortest processing time; (3) Select the machine with the lowest utilization rate; (4) Select the machine with the lowest load; (5) Randomly select a processable machine.
[0014] For further optimization, the AGV scheduling rules are as follows: (1) Select the earliest available AGV; (2) Select the AGV with the shortest transportation time; (3) Select the AGV with the lowest utilization rate; (4) Select the AGV with the lowest load; (5) Randomly select an available AGV.
[0015] For further optimization, in step S4, the job selection policy network maps according to the job scheduling rules to specific operations , and the AGV selection policy network maps according to the AGV scheduling rules to the transportation operation of the th AGV , and the machine selection policy network maps according to the machine status to the th machine for processing operations .
[0016] For further optimization, the method for generating the colored noise sequence is as follows: First, calculate the set of frequency components according to the input noise length , calculate the scaling factor for each frequency according to the color parameter of the noise, combine the scaling factors to obtain the spectral density , and calculate the standard deviation for normalization. Then, according to the standard deviation and the scaling factor Generate random numbers that conform to the normal distribution, which are used for the real and imaginary parts of the spectrum respectively. Combine the real and imaginary parts to generate a complex spectrum. Apply the inverse Fourier transform to the generated complex spectrum to obtain a noise sequence in the time domain. Normalize the generated noise sequence to obtain the final colored noise sequence.
[0017] The beneficial effects of the present invention are as follows: 1. Adopt a multi-strategy collaboration mechanism, including job selection, machine selection, and AGV selection. The decision-making process is based on the Markov decision model. In each selection module, use the proximal policy optimization algorithm with a policy-value system to train the action selection agent. These two networks are implemented through a multi-layer perceptron. The function of the policy network is to generate the probability distribution of different actions, which is used to guide the decision-making of the agent. The value network is responsible for estimating the value function of different states and providing support for the training of the policy network. The parameters of the policy network and the value network in each selection module are independent of each other and do not share with each other; 2. Introduce residual blocks into each policy network and value network, effectively improving the network performance, enabling the gradient to be transmitted more directly between different layers, thereby effectively alleviating the gradient disappearance problem. Add colored noise to the policy network to enhance the exploration ability of the agent and help break through the local optimal solution; In summary, through the introduction of a policy collaboration mechanism, multi-strategy DRL can flexibly select the optimal strategy or balance between multiple strategies when facing a dynamically changing environment to cope with multi-dimensional constraints and uncertainty factors. Applying the multi-strategy deep reinforcement learning method to the collaborative scheduling of flexible job shop production resources can not only effectively handle the complex and variable resource allocation and task scheduling problems in the job shop, but also optimize the collaborative effect between different types of production resources in the job shop. By comprehensively considering multiple production factors such as processing time, order priority, and delivery deadline, multi-strategy deep reinforcement learning can adjust the scheduling strategy in real time under the dynamically changing production requirements of the flexible job shop to achieve the optimal resource allocation and scheduling decision in the production process. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is a schematic diagram of the operation process of the present invention; Figure 2 is a schematic diagram of the structures of the policy network and the value network; Figure 3 is a schematic diagram of the collaborative scheduling scenario of production resources driven by dynamic events; Figure 4 is a schematic diagram of the multi-strategy collaboration framework; Figure 5 is a schematic diagram of the generation process of the colored noise sequence; Figure 6 is for schematic diagram of the calculation process; Figure 7 Schematic diagram of the training process of the MPPPO training algorithm for multi-strategy collaboration. Detailed implementation method
[0019] A flexible job shop production resource scheduling method based on multi-strategy deep reinforcement learning, as Figure 1 shown. This method adopts a multi-strategy collaboration mechanism, including job selection, machine selection, and automated guided vehicle selection. The decision-making process is based on the Markov Decision Process (MDP). In each selection module, the proximal policy optimization algorithm (PPO) with a policy-value (Actor-Critic) architecture is used to train the action selection agent. These two networks are implemented through a Multilayer Perceptron (MLP). The function of the Actor network is to generate the probability distribution of different actions to guide the decision-making of the agent; the Critic network is responsible for estimating the value function of different states to support the training of the policy network. The parameters of the policy network and the value network in each selection module are independent of each other and not shared.
[0020] Policy network: The network takes the current state as input. Each residual block consists of two fully connected layers, uses the ReLU activation function, and has a skip connection. The following is the formula representation of the residual block: where is the output of the first fully connected layer in the residual block, is the output of the second fully connected layer in the residual block, is the input of the residual block, that is, the state feature when it is the first layer of the first residual block, , are the weight matrix and bias term of the first fully connected layer, , are the weight matrix and bias term of the second fully connected layer. The formula for the skip connection is as follows: The output is used as the output of the residual block and is connected to the next layer. Multiple residual blocks can be stacked. Assuming there are residual blocks, the output of the th residual block is: where and are respectively the weight matrix and bias term of the first fully connected layer in the th residual block, and are respectively the weight matrix and bias term of the second fully connected layer in the th residual block.
[0021] Pass the output of the residual block to the fully connected layer of the policy network: wherein, represents the output after passing through the last fully connected layer of the policy network, , represents the weight matrix and bias term of the fully connected layer. Add colored noise to : Finally, calculate the probability distribution of the action through : wherein, is the number of actions, is the noisy output at the rescheduling time point t, is the noisy output corresponding to the i-th action.
[0022] Value network: Similar to the policy network, the value network also takes the state as input. The residual blocks of the value network are the same as those of the policy network, stacked residual blocks, and the formula is as follows: The calculation formula is the same as that of the policy network. Each residual block contains two fully connected layers and skip connections. Pass the output of the residual block to the fully connected layer to calculate the state value wherein, is the output of the last residual block in the value network, , represents the weight matrix and bias term of the fully connected layer.
[0023] Judge the pros and cons of a certain action under the current policy through the value network and the actual return, that is, the value of the advantage function (Advantage Function) is equal to the actual return minus the estimation of the current state value by the value network. If the value of the advantage function is positive, it means that this action is better than the average level, and the policy network should increase the probability of executing this action; otherwise, the probability should be reduced.
[0024] Flexible job shop production resource scheduling method based on multi-strategy deep reinforcement learning, the specific steps are as follows: S1. Construct a Markov decision model based on the agent models of job selection, machine selection, and AGV selection, and design the action space of the composite scheduling rules. The action space of the composite scheduling rules in step S1 includes job scheduling rules, machine scheduling rules, and AGV scheduling rules; S2. Initialize the network parameters of the job selection, machine selection, and AGV selection agents, and input the above model for training; S3. Based on the training in step S2, determine whether the training is over: If the training is not over, determine whether there is a new job. If there is a new job, add the new job to the set of unfinished jobs and then calculate the states of the jobs, machines, and AGVs. Otherwise, directly calculate the states; If the training is over, then perform the judgment of terminating the job; S4. After the state is determined, perform multi-strategy collaborative scheduling of job selection, machine selection, and AGV selection at each scheduling moment. Among them, colored noise is added to the operations of the job selection policy network, machine selection policy network, and AGV selection policy network. The colored noise is the noise term corresponding to the output dimension of the policy network. This noise term is added to the original output of the policy network to form a noisy output. The noisy output is sent into the function to obtain the probability distribution of multiple job scheduling rules, randomly sample a probability value from the probability distribution, and obtain its corresponding index. According to this index, select the corresponding rule from the composite scheduling rules in step S1 and map it to the specific operations, transportation operations, and processing operations of the target task; The generation method of the colored noise sequence is as follows: As Figure 5 shown, set the color parameter to 0.5, set the noise length to the output length of the policy network. In the first and second lines, first calculate the set of frequency components according to the input ; in the third and fourth lines, calculate the scaling factor for each frequency according to the of the noise; in the fifth and sixth lines, combine the scaling factors to obtain the spectral density , and calculate the standard deviation For normalization; in lines 7 - 8, random numbers conforming to the normal distribution are generated according to the standard deviation and the scaling factor, and are used for the real and imaginary parts of the spectrum respectively; in lines 9 - 12, special processing is required for the 0 frequency and the Nyquist frequency, which only exist when H is even, to ensure that the generated noise has the correct symmetry. H is the usage time of the AGV from the start to the rescheduling time point t. In line 13, a complex spectrum is generated by combining the real and imaginary parts; in line 14, the inverse Fourier transform is applied to the generated complex spectrum to obtain the noise sequence in the time domain; in line 15, the generated noise sequence is normalized to obtain the final colored noise sequence; S5. Execute the operation through multi - strategy collaboration in step S4 by the AGV transport to the machine for processing, calculate the operation start and end times, and store the states of job selection, machine selection, and AGV selection , actions and rewards into the corresponding experience pool (as Figure 4 shown), where t is the rescheduling time point, , , are the state, action, and reward of the job, machine, and AGV at the rescheduling time point t respectively. The job selection policy network, machine selection policy network, and AGV selection policy network of the present invention are updated using a shared immediate reward, that is where, is the average tardiness time corresponding to the state , is the average tardiness time corresponding to the state , is the job state corresponding to the time t + 1, is the machine corresponding to the time t + 1, is the AGV state corresponding to the time t + 1; S6. Repeat steps S3, S4, and S5 until the operations of the newly arrived jobs and the existing jobs in the workshop in a single scheduling instance are completed. Then sample a single - round of data from the experience pool, calculate the loss, and update the network parameters. After the experience pool is emptied, enter the next round of iteration.
[0025] Among them, steps S1, S2, S3, and S4 are the Markov decision process (MDP) of the job selection, machine selection, and AGV selection modules.
[0026] Job selection module MDP process Step 1. Construct job - scheduling rules and calculate at the rescheduling point (Job operation completed or new job arrived) Set of overdue jobs and set of unfinished jobs , where represents the operations completed by the job , represents the total number of operations of the job , represents the average completion time for the machine to finish the previous task represents the job completion time of the current operation.
[0027] Job scheduling rules : If is empty, select the job with the minimum average slack time , otherwise select the job with the longest overdue time and high priority for processing ; Job scheduling rules : If is empty, select the job with a small critical ratio of slack time and high priority , otherwise select the job with the longest overdue time and high priority for processing ; Job scheduling rules : If is empty, select the job with a low operation completion rate and high priority , otherwise select the job with the longest overdue time and high priority for processing ; Job scheduling rules : To avoid getting stuck in local optimality, randomly select a job in the set of unfinished jobs for processing ; Among them, represents the urgency level of the job , represents the estimated processing time for the remaining operations of the job , represents the estimated transportation time for the remaining operations of the job , represents the overdue time of the i-th job.
[0028] In summary, at the rescheduling point time Job selection action .
[0029] Step 2. Calculate the status characteristics of each job at each rescheduling point , where: Average completion rate of operations , Average completion rate of jobs , Estimate the proportion of overdue jobs in the total number of jobs , Proportion of actual overdue jobs in the total number of jobs , Indicates the job 's completion rate, Indicates the total number of jobs, 、 Calculate as follows: Among them, Indicates the job At the rescheduling point The set of estimated overdue operations, Indicates The number of overdue operations.
[0030] At the rescheduling time point The set of actual overdue jobs.
[0031] The present invention also takes the difference in the state between the rescheduling time point and as part of the state representation. The input of the policy network of the final job selection module is , where Indicates the state at the current rescheduling time point , Indicates the state at the current rescheduling time point And the difference in state from the previous rescheduling time point .
[0032] Step 3, The neural network structure of the job selection policy network is as Figure 2 shown. According to the input The original output generated, its dimension is 4, corresponding to 4 job scheduling rules. Generate a colored noise term with dimension 4 through Algorithm 1, and add it to the original output to form a noisy output . Then, the noisy output is sent into function to obtain the probability distribution of 4 job scheduling rules: Among them, Is the probability distribution of the job scheduling rule; Subsequently, a probability value is randomly sampled from the probability distribution, and its corresponding index is obtained. According to this index, the corresponding rule is selected from the job scheduling rules in Step 1 and mapped to the target job Specific operations .
[0033] MDP process of machine selection module Step 1: Construct the following machine scheduling rules Machine scheduling rules : Select the earliest available machine ; Machine scheduling rules : Select the machine with the shortest processing time ; Machine scheduling rules : Select the machine with the lowest utilization rate ; Machine scheduling rules : Select the machine with the lowest load ; Machine scheduling rules : Randomly select a processable machine ; Among them, represents the time when the th machine completes the previous processing task, represents the operation is 1 when processed on machine , otherwise 0, represents the operation on machine processing time on the machine, represents the makespan of the machine at the rescheduling time point . In summary, at the rescheduling time point Machine selection action , where is the set of optional actions for the machine
[0034] Step 2. Calculate the state characteristics of each machine at the rescheduling time point , among which, the average machine utilization rate , the standard deviation of machine utilization rate , the ratio of the maximum workload to the average workload of the machine , the ratio of the standard deviation of the machine workload to the average machine workload , , represents the utilization rate of machine , represents the number of machines, represents the The workload of a machine represents the average workload of the machine.
[0035] The present invention takes the difference in the states between the rescheduling time points and as part of the state representation. The input to the policy network of the final machine selection module is .
[0036] Step 3: The neural network structure of the machine selection policy network is as Figure 2 shown. According to the input to the policy network of the machine selection module, the original output generated is , whose dimension is 5, corresponding to 5 machine scheduling rules. A colored noise term with a dimension of 5 is generated through Algorithm 1 and added to the original output to form a noisy output . Then, the noisy output is fed into function to obtain the probability distribution of the 5 machine scheduling rules: Subsequently, a probability value is randomly sampled from the probability distribution, and its corresponding index is obtained. According to this index, the corresponding rule is selected from the machine scheduling rules in Step 1 and mapped to the machine .
[0037] AGV selection module MDP process Step 1: Construct the following AGV scheduling rules AGV scheduling rule : Select the earliest available AGV ; AGV scheduling rule : Select the AGV with the shortest transportation time ; AGV scheduling rule : Select the AGV with the lowest utilization rate ; AGV scheduling rule : Select the AGV with the lowest load ; AGV scheduling rule : Randomly select an available AGV ; Among them, represents the index of the AGV, represents the set of AGVs available for transportation operations , represents Transport the previous task The end time of transportation, Indicates the completion time of the operation The completion time, Indicates Using the th AGV for transportation is 1, otherwise 0, Indicates the use of Transportation operation The transportation time, Indicates the maximum transportation end time of the AGV at the rescheduling time point . In summary, at the rescheduling time point AGV selection action , where is the set of optional actions for the AGV.
[0038] Step 2. Calculate the state characteristics of each AGV at the rescheduling point , where the average utilization rate of the AGV , the standard deviation of the AGV utilization rate , the ratio of the maximum workload to the average workload of the AGV , the ratio of the standard deviation of the AGV workload to the average machine workload , Indicates the utilization rate of the AGV , Indicates the number of AGVs, Indicates the AGV Workload, Indicates the AGV Average workload.
[0039] The present invention also takes the difference in the state between the rescheduling time point and as part of the state representation. Finally, the input of the policy network of the AGV selection module is .
[0040] Step 3. The neural network structure of the AGV selection policy network is as Figure 2 shown. According to the input of the policy network of the AGV selection module being , the original output is generated as , whose dimension is 5, corresponding to 5 AGV scheduling rules. Generate a colored noise term with a dimension of 5 through Algorithm 1, and add it to the original output to form a noisy output ; then, the noisy output is fed into In the function, the probability distributions of 5 AGV scheduling rules are obtained: Subsequently, a probability value is randomly sampled from the probability distribution, and its corresponding index is obtained. According to this index, the corresponding rule is selected from the AGV scheduling rules in step one and mapped to the AGV .
[0041] In step S5, the reward is obtained after job selection, machine selection, and AGV selection, and the AGV transport operation to the machine for processing. After that, the environment transfers from the current state to the next state and feedbacks the reward to the agent to help it learn. The job selection policy network, machine selection policy network, and AGV selection policy network of the present invention are updated using a shared immediate reward, that is , where The calculation process of is Algorithm 2 (as Figure 6 shown), where represents the current rescheduling time point.
[0042] Algorithm 3 (as Figure 7 shown) provides the training process of the MPPPO training algorithm for multi-policy collaboration. In the algorithm, lines 2-13 represent the core scheduling process of a single scheduling instance. During this process, three agents will collect experience tuples. After the scheduling stage ends, the job rule selection agent, machine rule selection agent, and AGV rule selection agent respectively use the PPO algorithm (lines 14-24) with a Figure 7 neural network structure to independently calculate the error and update the parameters through gradient descent, where the value in line 20 is 0.2, represents the clipping function, which limits the change range of the probability ratio : .
[0043] The dynamic events considered in the present invention are new jobs with different urgencies , randomly arriving at the workshop. These jobs need to be processed on multiple machines , and each job contains a specific operation sequence , and each operation can be processed by a group of available machines , and the processing time of this operation on different machines is different. Represents the machine index. The arrival information of a new job is unknown before it enters the workshop, and its arrival time Following the exponential distribution, the overtime is set to , and the arrival time of the original job in the workshop is 0. When arriving at the workshop, the scheduling system will trigger the rescheduling process. According to the urgency and priority of the job, the system will re-plan the operations of all unfinished jobs and newly arrived jobs and generate a new production plan. Figure 3 As shown, according to the plan, the AGV transports the semi-finished work to the buffer of the designated machine (such as Figure 3 (a)). When two consecutive operations of a semi-finished product job need to be processed on different machines, the AGV is responsible for transporting the semi-finished product job from one machine to the target machine (such as Figure 3 (b)). If the AGV receives a second task after completing the first transport task, it can go directly to the next target machine without returning to the warehouse (e.g. Figure 3 (c)). When there is no new transport task, the AGV will return to the warehouse and wait for orders (as shown in Figure 3 (d)). Whenever an operation is completed, the relevant machine will select the next operation to be processed from the buffer and continue the production process.
[0044] By introducing a strategy collaboration mechanism, multi-strategy DRL can flexibly select the optimal strategy or balance between multiple strategies in the face of a dynamically changing environment to cope with multi-dimensional constraints and uncertainties, that is, to solve the problem of allocating each job to the appropriate machine at the same time, sorting the jobs on the machine, selecting the appropriate AGV, and determining the start processing time of each operation. The mean tardiness (MT) is used as the objective function, which should be minimized, and is mathematically described as follows: in, Indicates the job The actual completion time, Indicates the number of jobs.
[0045] The present invention can not only effectively handle the complex and changeable resource configuration and task scheduling problems in the workshop, but also optimize the synergy between different types of production resources in the workshop. By comprehensively considering multiple production factors such as processing time, order priority, and delivery deadline, multi-strategy deep reinforcement learning can adjust the scheduling strategy in real time under the dynamically changing production needs of the flexible workshop, so as to achieve the optimal resource allocation and scheduling decision-making in the production process.
[0046] The main features, usage methods, basic principles, and advantages of the present invention have been shown and described above. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will also have various changes and improvements according to actual situations, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A flexible job shop production resource scheduling method based on multi-strategy deep reinforcement learning, characterized in that Adopt a multi-strategy collaboration mechanism to construct a decision-making model including a job selection module, an AGV selection module, and a machine selection module, and schedule the flexible job-shop production resources that arrive randomly. The specific steps are as follows: S1. Establish a job selection module, a machine selection module, and an AGV selection module based on Markov decision-making. The job selection module, the machine selection module, and the AGV selection module all use an optimization algorithm including a policy network and a value network to train an action selection agent; the parameters of the policy network and the value network of the job selection module, the machine selection module, and the AGV selection module are independent of each other; design a composite scheduling rule for job selection, machine selection, and AGV selection; S2. Initialize the network parameters of the agents of the job selection module, the machine selection module, and the AGV selection module, and perform training; S3. During the training process, if there is a new job, add the new job to the set of unfinished jobs and then calculate the states of the job, the machine, and the AGV. Otherwise, directly calculate the states; S4. After the status is determined, multi-strategy collaborative scheduling of job selection, machine selection, and AGV selection is performed separately at each scheduling moment. The outputs of the job selection policy network, machine selection policy network, and AGV selection policy network are sent into the function to obtain the probability distribution of multiple job selection rules, randomly sample a probability value from the probability distribution, and obtain its corresponding index. According to this index, select the corresponding rule from the composite scheduling rules in step S1, and map it to the specific operation according to the job status , map it to the transportation operation according to the AGV status the AGV, and map it to the machine used for the processing operation according to the machine status the machine; S5. Determine the operation through the multi-strategy collaboration in step S4 by transported to the machine After processing, it enters the next state, calculates the operation start and end times, calculates the reward of the agent according to the average delay time, and stores the states, actions, and rewards of job selection, machine selection, and AGV selection in the corresponding experience pool; S6. Repeat steps S3, S4, and S5 until the operations of the newly arrived jobs and the existing jobs in the workshop in a single scheduling example are completed. Sample single-round data from the experience pool, calculate the loss, and update the network parameters. After the experience pool is emptied, enter the next round of iteration.
2. The flexible job shop production resource scheduling method based on multi-strategy deep reinforcement learning according to claim 1, wherein Residual blocks are introduced into both the policy network and the value network. Each residual block includes two fully connected layers, uses the ReLU activation function, and has a skip connection. Multiple residual blocks in the network are stacked. The output of the last residual block is passed to the fully connected layer of the network and the result is output; the output of the policy network is: , Among them, represents the output of the last fully connected layer of the policy network, , represents the weight matrix and bias term of the fully connected layer, is the output of the last residual block of the policy network; In the value network, the last residual block is passed to the fully connected layer of the network to calculate the state value : Among them, is the output of the last residual block in the value network, , represents the weight matrix and bias term of the fully connected layer of the value network.
3. The flexible job shop production resource scheduling method based on multi-strategy deep reinforcement learning according to claim 1 or 2, characterized in that During the operations of the above-mentioned task selection policy network, machine selection policy network, and AGV selection policy network, colored noise is added. The colored noise is the noise term corresponding to the output dimension of the policy network, and this noise term is added to the original output of the policy network to form a noisy output. In the policy network, the colored noise is added to : , By calculating the probability distribution of the action Among them, is the number of actions.
4. The flexible job shop production resource scheduling method based on multi-strategy deep reinforcement learning according to claim 1, wherein The job selection policy network, the machine selection policy network, and the AGV selection policy network are updated using a shared immediate reward.
5. The flexible job shop production resource scheduling method based on multi-strategy deep reinforcement learning according to claim 1, characterized in that, The action space of the composite scheduling rule in step S1 includes a job scheduling rule, a machine scheduling rule, and an AGV scheduling rule.
6. The flexible job shop production resource scheduling method based on multi-strategy deep reinforcement learning according to claim 5, wherein The job scheduling rule is: (1), If is empty, select the job with the minimum average relaxation time, otherwise select the job with the longest overdue time and high priority for processing; (2) If is empty, select the job with a small critical ratio of relaxation time and high priority; otherwise, select the job with the longest overdue time and high priority for processing. (3) If is empty, select the job with a low operation completion rate and a high priority. Otherwise, select the job with the longest overdue time and a high priority for processing; (4). To avoid falling into a local optimum, randomly select a job for processing in the set of unfinished jobs; Among them At the rescheduling point Set of overdue jobs 7. The flexible job shop production resource scheduling method based on multi-strategy deep reinforcement learning according to claim 5, characterized in that, The machine scheduling rule is: (1). Select the earliest available machine; (2). Select the machine with the shortest processing time; (3). Select the machine with the lowest utilization rate; (4). Select the machine with the lowest load; (5). Randomly select a processable machine.
8. The flexible job shop production resource scheduling method based on multi-strategy deep reinforcement learning according to claim 5, wherein, The AGV scheduling rule is: (1). Select the earliest available AGV; (2). Select the AGV with the shortest transportation time; (3). Select the AGV with the lowest utilization rate; (4). Select the AGV with the lowest load; (5). Randomly select an available AGV.
9. The flexible job shop production resource scheduling method based on multi-strategy deep reinforcement learning according to claim 1, wherein In step S4, the job selection policy network selects a scheduling rule based on the job status and maps it to a specific operation , the AGV selection policy network selects a scheduling rule based on the AGV status and maps it to a transportation operation the th AGV , the machine selection policy network maps based on the machine status to the th machine for processing operations .
10. The flexible job shop production resource scheduling method based on multi-strategy deep reinforcement learning according to claim 3, wherein The generation method of the colored noise sequence is as follows: First, according to the input noise length calculate the set of frequency components , and according to the color parameter of the noise, calculate the scaling factor for each frequency, combine the scaling factors to obtain the spectral density , and calculate the standard deviation for normalization. Then, according to the standard deviation and the scaling factor , generate random numbers that conform to the normal distribution, which are used for the real and imaginary parts of the spectrum respectively. Combine the real and imaginary parts to generate a complex spectrum, apply the inverse Fourier transform to the generated complex spectrum to obtain the noise sequence in the time domain, and normalize the generated noise sequence to obtain the final colored noise sequence.
Citation Information
Patent Citations
Job shop machine and AGV combined scheduling method based on deep reinforcement learning
CN116483075A
Equipment manufacturing workshop intelligent scheduling method and system based on deep reinforcement learning
CN116542445A
Dynamic scheduling method and device for flexible job shop based on deep reinforcement learning
CN118153896A
Dynamic flexible job shop scheduling method considering multiple types of events
CN119026832A
Workshop scheduling method considering AGV transportation time based on near-end strategy optimization algorithm
CN119151227A
Cited By
Control method and system based on intelligent collaborative production
CN120578145A
Dynamic scheduling method and device and storage medium
CN121638826A
Multi-agent flexible workshop production scheduling method and system based on condition attribution mechanism
CN122284552A
Multi-agent flexible workshop production scheduling method and system based on conditional attribution mechanism
CN122284552B